MTU Data Decoupling Synthesis and Model Training Method, System, Device and Medium
By decoupling table image rendering and question-and-answer pair generation, multi-modal table understanding data is formed, and mixed multi-resolution visual encoder and language model optimization model are used to solve the performance limitation problem caused by the limited data scale of the existing MTU model, achieving significant performance improvement.
Patent Information
- Application Number
- CN202510311103.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-17
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-03-17
AI Technical Summary
The existing multimodal table understanding (MTU) model has limited performance due to its limited data scale, and the existing methods cannot fundamentally solve the problem of data missing, resulting in the model performance being unable to achieve optimal performance.
Through a MTU data decoupling and synthesis method, table data is collected and diverse table images are generated through data augmentation technology, prompt words are generated based on question categories, and the language model is guided to output questions and answers, forming multimodal table understanding data. Then, a hybrid multi-resolution vision encoder and language model is built, and the model is optimized to improve performance through resolution transformation and feature fusion.
By synthesizing accurate MTU data, the problems caused by hallucinations are avoided, the performance of the MTU model is significantly improved, and the most advanced performance in multimodal table understanding tasks are achieved.
Smart Images

Figure CN119832579B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of multimodal table understanding, and in particular, to a method, system, device and medium for MTU data decoupling synthesis and model training. Background Art
[0002] Multimodal Table Understanding (MTU) is a process where, after a user inputs a table image and a corresponding question, the model retrieves and reasons about the table image for the question posed by the user and finally outputs the correct answer. Due to the limited scale of multimodal table understanding data, the performance of existing models is restricted.
[0003] Although a large number of methods have emerged in the multimodal community to improve the performance of MTU by designing different network structures, such methods cannot fundamentally solve the problem of data shortage, thus affecting the model performance from reaching the optimal. Another method is to use a Multimodal Large Language Model (MLLM) to synthesize more samples, but this will lead to the generation of hallucinations, thereby generating incorrect question-and-answer sample pairs, affecting the parameter optimization during the subsequent training of the MTU model, and the cost of labeled data is high.
[0004] In view of this, the present invention is specifically proposed. Summary of the Invention
[0005] The object of the present invention is to provide a method, system, device and medium for MTU data decoupling synthesis and model training, which can synthesize MTU data required for MTU model training and train an MTU model with greatly improved performance.
[0006] The object of the present invention is achieved by the following technical solutions:
[0007] An MTU data decoupling synthesis method includes:
[0008] Collect table data;
[0009] Perform diverse processing on different table data through data augmentation techniques to obtain visually diverse table images, where each table data corresponds to a table image;
[0010] Generate a prompt word corresponding to each table data in combination with a set question category, and guide a first large language model to output questions and answers;
[0011] Integrate the table image, questions and answers corresponding to each table data to form MTU data, where MTU is multimodal table understanding.
[0012] An MTU model training method includes:
[0013] Construct an MTU model, including: a hybrid multi-resolution visual encoder, a projection layer, and a second large language model;
[0014] Generate MTU data using the aforementioned MTU data decoupling and synthesis method, and train the MTU model. The training steps include: for each MTU data, perform operations of increasing and decreasing the resolution of the table image respectively to obtain images called high-resolution images and low-resolution images; the hybrid multi-resolution visual encoder extracts local features from the text regions in the high-resolution images, extracts the spatial context relationship between cells from the low-resolution images to obtain global features, and fuses the local features and global features to obtain hybrid-resolution visual features; the projection layer converts the hybrid-resolution visual features into the embedding space of the second large language model and outputs the converted visual features; the second large language model outputs answers according to the converted visual features and the corresponding questions; construct a loss function by combining the answers output by the second large language model and the answers included in the MTU data, and use the loss function to optimize the MTU model.
[0015] An MTU data decoupling and synthesis system for implementing the aforementioned MTU data decoupling and synthesis method. The system includes:
[0016] A data collection unit for collecting table data;
[0017] A table image rendering unit for diversifying different table data through data augmentation techniques to obtain visually diverse table images, where each table data corresponds to a table image;
[0018] A table question-answer pair generation unit for generating a prompt word corresponding to each table data in combination with a set question category and guiding the first large language model to output questions and answers;
[0019] An MTU data generation unit for synthesizing the table image, questions, and answers corresponding to each table data to form MTU data, where MTU is multi-modal table understanding.
[0020] An MTU model training system, including:
[0021] A model construction unit for constructing an MTU model, including: a hybrid multi-resolution visual encoder, a projection layer, and a second large language model;
[0022] A model training unit is used to generate MTU data by using the aforementioned MTU data decoupling and synthesis method and train the MTU model. The training steps include: for each MTU data, perform resolution increase and decrease operations on the resolution of the table image respectively, and the obtained images are called high-resolution images and low-resolution images; the hybrid multi-resolution vision encoder extracts local features from the text regions in the high-resolution images, extracts the spatial context relationship between cells from the low-resolution images, obtains global features, and fuses the local features and global features to obtain hybrid resolution vision features; the projection layer converts the hybrid resolution vision features into the embedding space of the second large language model and outputs the converted vision features; the second large language model outputs an answer according to the converted vision features and the corresponding questions; construct a loss function by combining the answer output by the second large language model and the answer included in the MTU data, and use the loss function to optimize the MTU model.
[0023] A processing device includes: one or more processors; a memory for storing one or more programs;
[0024] Wherein, when the one or more programs are executed by the one or more processors, the one or more processors implement the aforementioned MTU data decoupling and synthesis method and MTU model training method.
[0025] A readable storage medium stores a computer program, which implements the aforementioned MTU data decoupling and synthesis method and MTU model training method when executed by a processor.
[0026] As can be seen from the technical solutions provided by the present invention above, the MTU data synthesis process is decoupled into two independent steps: table image rendering and table question-answer pair generation. Accurate MTU data can be synthesized by combining the collected table data, avoiding the problem of hallucinations that are likely to occur when using MLLM annotations. Moreover, an MTU model can be trained by combining MTU data and a significant performance improvement has been achieved. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for description in the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained according to these drawings without creative efforts.
[0028] Figure 1 It is a flowchart of an MTU data decoupling and synthesis method provided by an embodiment of the present invention;
[0029] Figure 2Schematic diagram of an MTU data decoupling and synthesis method provided by an embodiment of the present invention;
[0030] Figure 3 Schematic diagram of an MTU model training method provided by an embodiment of the present invention;
[0031] Figure 4 Schematic diagram of an MTU data decoupling and synthesis system provided by an embodiment of the present invention;
[0032] Figure 5 Schematic diagram of an MTU model training system provided by an embodiment of the present invention;
[0033] Figure 6 Schematic diagram of a processing device provided by an embodiment of the present invention. Detailed implementation manners
[0034] Next, in combination with the accompanying drawings in the embodiments of the present invention, the technical solutions in the embodiments of the present invention are clearly and completely described. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0035] First, the following explanations are made for the terms that may be used in this article:
[0036] The description of terms such as "including", "comprising", "containing", "having" or other similar semantics should be interpreted as non-exclusive inclusion. For example: including a certain technical feature element (such as raw material, component, ingredient, carrier, dosage form, material, size, part, component, mechanism, device, step, process, method, reaction condition, processing condition, parameter, algorithm, signal, data, product or article, etc.) should be interpreted as not only including the clearly listed certain technical feature element, but also including other technical feature elements well-known in the art that are not clearly listed.
[0037] The term "consisting of" means excluding any technical feature element that is not clearly listed. If this term is used in a claim, this term will make the claim a closed type, making it not include technical feature elements other than the clearly listed technical feature elements, except for related conventional impurities. If this term only appears in a sub-clause of a claim, then it only limits the elements clearly listed in that sub-clause, and the elements recorded in other sub-clauses are not excluded from the overall claim.
[0038] The following provides a detailed description of a method, system, device, and medium for MTU data decoupling synthesis and model training provided by the present invention. The content not described in detail in the embodiments of the present invention belongs to the prior art well-known to those skilled in the art. In the embodiments of the present invention, those not specified in specific conditions are carried out according to the conventional conditions in the art or the conditions recommended by the manufacturer. The instruments used in the embodiments of the present invention not specified in the manufacturer are all conventional products that can be obtained through commercial purchase.
[0039] Embodiment 1
[0040] An embodiment of the present invention provides an MTU data decoupling synthesis method, as Figure 1 shown, which mainly includes the following steps:
[0041] Step 1: Collect tabular data.
[0042] In the embodiments of the present invention, tabular data can be collected from open-source datasets.
[0043] Preferably, the collected tabular data can be preprocessed, and the preprocessed tabular data is used in the subsequent steps for tabular image rendering and tabular question-answer pair generation.
[0044] Specifically, the preprocessing steps may include: converting the tabular data into tabular codes containing tabular content and structure respectively; filtering the tabular codes to retain the tabular codes whose number of rows and columns meet the set requirements; then, using the tabular codes obtained by preprocessing to generate corresponding tabular images, questions, and answers.
[0045] Step 2: Tabular image rendering.
[0046] In the embodiments of the present invention, different tabular data is diversified through data augmentation techniques to obtain visually diverse tabular images, where each tabular data corresponds to a tabular image.
[0047] Preferably, the tabular attributes and tabular layouts in different tabular data can be diversified.
[0048] Step 3: Tabular question-answer pair generation.
[0049] Combined with the set question categories, generate prompt words corresponding to each tabular data, and guide the first large language model to output questions and answers.
[0050] Preferably, the set question categories are MTU question types, and the question categories are subdivided into specific subcategories; for each tabular data, a plurality of subcategories are randomly selected by a prompt word generator to generate corresponding prompt words.
[0051] Preferably, the prompt words include a specified output format to guide the first large language model to output questions and answers in the corresponding format. The answers output by the first large language model include two answers, one is the answer containing the reasoning process, and the other is the direct result.
[0052] Step 4: Integrate the table images, questions, and answers corresponding to each table data to form MTU data.
[0053] It should be noted that the above step numbers are mainly used to identify different steps and do not represent the execution order of the steps. The specific execution order of the steps can be determined by those skilled in the art in combination with the content of the steps. For example, there is no strict execution order between Step 2 and Step 3, and they can be executed in parallel or in any order successively.
[0054] To more clearly present the technical solutions provided by the present invention and the resulting technical effects, the method provided by the embodiments of the present invention will be described in detail below with specific embodiments.
[0055] As Figure 2 shown, it is a schematic diagram of the MTU data decoupling and synthesis method provided by the embodiments of the present invention, mainly including two independent parts: table image rendering and table question and answer pair generation. These two parts are independent of each other and can work simultaneously. Before performing the above two parts of work, it is necessary to collect necessary table data as a medium to synthesize the required MTU data. The following will introduce each part of the work in detail. In addition, the present invention can support multiple types of languages. For example, Figure 2 Examples of English and Chinese language types are provided.
[0056] 1. Data collection.
[0057] In the embodiments of the present invention, table data can be collected from open source data sets. The table data can be in the text sequence format of tables, such as HTML (HyperText Markup Language), LaTeX (a high-quality typesetting system format), and Markdown (a lightweight markup language), etc.
[0058] Exemplarily, for example, 113K (K represents thousand) and 652K text sequence tables can be collected from the open source data sets FinTabNet and TableLLama respectively, mainly including table data in HTML and Markdown formats.
[0059] After that, preprocessing is required: (1) Convert the collected table data into table codes that only contain the table content and structure. (2) Filter the table codes and retain the table codes whose number of rows and columns meet the set requirements. For example, retain the table codes with the number of rows between 3 and 60 and at least 3 columns.
[0060] Exemplarily, 636K tables were finally selected for subsequent synthesis.
[0061] 2. Table image rendering.
[0062] In the embodiments of the present invention, a series of data augmentation techniques are adopted to enhance the visual diversity of table images. These augmentation techniques are applied to various table attributes, such as background color, font size, and font style, to generate a series of visually distinct table images. In addition, different table layouts are randomly selected, including pure rows, pure columns, or fully grid-separated ones, to achieve more diverse visual effects. Through these augmentation techniques, visually diverse table images are finally synthesized, which are very similar to the table images in the existing MTU samples.
[0063] Exemplarily, through the above process, 636K <code, image> samples were synthesized.
[0064] 3. Table question-answer pair generation.
[0065] Compared with directly using large language models to generate table question-answer pairs, the present invention defines 6 main question categories and 11 more detailed sub-categories to ensure that the synthesized data can cover a wide range of MTU question types. The question types in MTU are mainly divided into 6 major categories, including a total of 11 sub-categories: retrieval (table retrieval), data operation (counting, sorting, determining range, filtering), numerical calculation (simple numerical calculation, complex calculation), free response (free answering questions), selection (multiple choice, table fact verification), and summary (table summary).
[0066] Retrieval involves directly identifying and extracting specific information or multiple table cells from a table. Data operation processes table data according to questions to retrieve answers, and these operations include: counting, sorting, determining range, and filtering. Numerical calculation analyzes the quantitative data in the table, which includes simple arithmetic operations of addition, subtraction, multiplication, and division of numbers (i.e., simple numerical calculation) and complex mixed operations (i.e., complex calculation). Free response integrates facts and reasoning into a coherent sentence when answering questions. Selection involves selecting 1 correct option from N answers (i.e., multiple choice), usually N is 4. Table fact verification is a special case where N is 2, and it requires a choice between an affirmative or negative answer. Summary includes a concise and coherent description or title of the key information in the table.
[0067] In the embodiments of the present invention, a prompt generator is designed to generate specific prompt templates for standardizing the question types and output formats generated by the first large language model. Such as Figure 2As shown in the prompt generator in , first randomly select M different subcategories from the given task pool, where . The task pool represents the 6 question types defined above. In addition, the output format is specified in the prompt, and a structured template is provided for the generated Q&A pairs <question, answer>. Each question corresponds to a detailed answer and a short answer to reflect the problem-solving process. Through the above process, a variety of prompt templates are generated, significantly enriching the diversity of question types generated by the first large language model. In addition, introducing detailed steps in the reasoning process to reason and calculate the answers to questions is crucial for the MTU large model to achieve the MTU task. Like humans solving problems, large language models are good at solving problems step by step rather than directly generating the answers to questions, which will increase the learning difficulty of the MTU model.
[0068] Exemplarily, the first large language model can select the Doubao-Pro model, which is the commercial Doubao large model.
[0069] Exemplarily, use 636K codes to generate 1.8M (M represents million) Q&A sample pairs <code, question, answer>. Subsequently, use the codes as an intermediate medium to combine the generated images with the Q&A pairs to create MTU data <image, question, answer>.
[0070] Compared with the existing multi-modal table understanding synthesis methods, the present invention has the following advantages: 1) Lower cost. Since the price of the large language model is low and the number of input tokens of the table code is less than that of images, the present invention achieves lower cost in the synthesis process. 2) Higher efficiency. The commercial large language model Doubao-Pro can process parallel requests, which greatly improves the efficiency of data annotation. 3) Stronger robustness. The present invention directly processes structured table codes, enabling the large language model to achieve higher accuracy and reliability in row-column logic, context reasoning, and content consistency. In addition, the present invention can also reduce the visual noise and information loss of the table codes. Compared with the existing MLLM synthesis methods, the Q&A pairs generated by the large language model are more accurate, more coherent, and the hallucinations are significantly reduced.
[0071] Embodiment 2
[0072] The embodiment of the present invention provides an MTU model training method, which mainly includes:
[0073] (1) Construct an MTU model, including: a hybrid multi-resolution visual encoder, a projection layer, and a second large language model.
[0074] (2) Use the method provided in the foregoing embodiment to generate MTU data and train the MTU model.
[0075] The training steps include: for each MTU data, perform operations of increasing and decreasing the resolution of the table image respectively to obtain images called high-resolution images and low-resolution images; the hybrid multi-resolution vision encoder extracts local features from the text regions in the high-resolution images, extracts the spatial context relationships between cells from the low-resolution images, obtains global features, and fuses the local features and the global features to obtain hybrid-resolution vision features; the projection layer converts the hybrid-resolution vision features into the embedding space of the second large language model and outputs the converted vision features; the second large language model outputs an answer (the inference process and the final result) according to the converted vision features and the corresponding questions; construct a loss function by combining the answer output by the second large language model and the answer included in the MTU data, and use the loss function to optimize the MTU model.
[0076] Exemplarily, the second large language model can be the Vicuna-1.5 7B model. Vicuna-1.5 7B is an open-source large language model based on the self-attention mechanism, and 7B (B represents billion) is the number of parameters of the model.
[0077] As Figure 3 shown, it is a schematic diagram of the MTU model training method. The hybrid multi-resolution vision encoder receives high-resolution and low-resolution table images as inputs, uses the corresponding vision encoders to extract relevant local and global information, and finally concatenates them in the feature dimension. Among them, the high-resolution image can enable the second large language model to obtain more visual information, which is also important for table images. For example, a five-stage ConvNeXt can be used to encode the high-resolution image input to extract local features. Even though ConvNeXt can extract local features from the text regions of the table image, it cannot capture the spatial context relationships between cells. Therefore, the low-resolution image is also used to model the global spatial relationships in the table image. For example, the ViT-L / 14 model can be used, with the low-resolution image as the input, to model the global spatial relationships in the table image. Here, ConvNeXt is a convolutional neural network, and ViT-L / 14 is a self-attention neural network.
[0078] Specifically, given an input table image, perform an operation of increasing its resolution, adjust it to a fixed resolution size H×H to obtain a high-resolution image , where H is the height and width of the resolution. For example, H can be set to 1536. Input the high-resolution image into ConvNeXt to obtain a feature map , ConvNeXt processes the high-resolution image Downsampled by 64 times, resulting in 24×24 visual features, with each feature having a dimension of 3072. At the same time, the resolution of the input table image is decreased to obtain a low-resolution image. For example, the resolution can be 336×336. After being processed by ViT-L / 14, 576 visual features are output , and the size of each visual feature is 1024. Finally, merge along the feature dimension and to form hybrid-resolution visual features.
[0079] In the embodiments of the present invention, a two-layer multi-layer perceptron can be used as the projection layer to transform the visual features into the embedding space of the second large language model. Finally, the obtained visual features after transformation are merged with the embedded text features, and they are input into the second large language model to generate answers. Here, the embedded text features refer to the features obtained after embedding the question; similarly, the answers here include: answers containing the reasoning process and direct results.
[0080] For each MTU data, the answer contained therein and the answer generated by the second large language model can be used to construct a loss function, and then the MTU model is optimized. The optimization process involved in this part can refer to conventional techniques, and the present invention will not elaborate.
[0081] Experiments show that the present invention uses synthetic MTU data to train the MTU model. In 24 in-domain and out-of-domain tests, the present invention achieved state-of-the-art performance on 21 test sets, demonstrating the effectiveness and generalization of the present invention.
[0082] The above solution provided by the embodiments of the present invention can answer the input table image and question of the user, and finally the trained MTU model can be used for tasks such as table data analysis and understanding in daily work. In implementation, it can be installed in devices such as computers and mobile phones in the form of software to communicate with users.
[0083] Through the description of the above embodiments, those skilled in the art can clearly understand that the above embodiments can be implemented by software or by means of software plus a necessary general hardware platform. Based on such an understanding, the technical solutions of the above embodiments can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.), including several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in the various embodiments of the present invention.
[0084] Embodiment III
[0085] The present invention also provides an MTU data decoupling and synthesis system, which is mainly used to implement the method provided in the foregoing Embodiment 1, as Figure 4 shown. This system mainly includes:
[0086] A data collection unit for collecting tabular data;
[0087] A tabular image rendering unit for diversifying different tabular data through data augmentation techniques to obtain visually diverse tabular images, where each tabular data corresponds to a tabular image;
[0088] A tabular question-answer pair generation unit for generating questions corresponding to each tabular data in combination with set question categories and guiding a first large language model to output answers, obtaining questions and answers corresponding to each tabular data;
[0089] An MTU data generation unit for synthesizing the tabular image, questions, and answers corresponding to each tabular data to form MTU data, where MTU is multi-modal tabular understanding.
[0090] Considering that the MTU data decoupling and synthesis process has been introduced in detail in the foregoing embodiments, it will not be elaborated here.
[0091] Embodiment 4
[0092] The present invention also provides an MTU model training system, which is mainly used to implement the method provided in the foregoing Embodiment 2, as Figure 5 shown. This system mainly includes:
[0093] A model construction unit for constructing an MTU model, including: a hybrid multi-resolution visual encoder, a projection layer, and a second large language model;
[0094] A model training unit for generating MTU data using the foregoing method and training the MTU model. The training steps include: for each MTU data, performing resolution increase and decrease operations on the resolution of the tabular image to obtain images called high-resolution images and low-resolution images; the hybrid multi-resolution visual encoder extracts local features from the text regions in the high-resolution images and extracts the spatial context relationship between cells from the low-resolution images to obtain global features, and fuses the local features and the global features to obtain hybrid resolution visual features; the projection layer converts the hybrid resolution visual features into the embedding space of the second large language model and outputs the converted visual features; the second large language model outputs answers according to the converted visual features and the corresponding questions; constructing a loss function by combining the answers output by the second large language model and the answers included in the MTU data, and optimizing the MTU model using the loss function.
[0095] Considering that the MTU model training process has been introduced in detail in the foregoing embodiments, it will not be elaborated herein.
[0096] Those skilled in the art can clearly understand that for the convenience and brevity of description, only the above-mentioned division of each functional module is used as an example in the above embodiments. In practical applications, the above functions can be allocated to different functional modules as needed, that is, the internal structure of the system is divided into different functional modules to complete all or part of the functions described above.
[0097] Embodiment Five
[0098] The present invention also provides a processing device, as Figure 6 shown, which mainly includes: one or more processors; a memory for storing one or more programs; wherein, when the one or more programs are executed by the one or more processors, the one or more processors implement the methods provided in the foregoing embodiments (that is, the MTU data decoupling and synthesis method, the MTU model training method).
[0099] Further, the processing device further includes at least one input device and at least one output device; in the processing device, the processor, the memory, the input device, and the output device are connected by a bus.
[0100] In the embodiments of the present invention, the specific types of the memory, the input device, and the output device are not limited; for example:
[0101] The input device can be a touch screen, an image acquisition device, a physical button, or a mouse, etc.;
[0102] The output device can be a display terminal;
[0103] The memory can be a random access memory (RAM), or a non-volatile memory, such as a disk memory.
[0104] Embodiment Six
[0105] The present invention also provides a readable storage medium storing a computer program, which implements the methods provided in the foregoing embodiments (that is, the MTU data decoupling and synthesis method, the MTU model training method) when the computer program is executed by a processor.
[0106] In the embodiments of the present invention, the readable storage medium, as a computer-readable storage medium, may be disposed in the aforementioned processing device. For example, it may be a memory in the processing device. In addition, the readable storage medium may also be various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a magnetic disk, or an optical disc.
[0107] As described above, the foregoing are only preferred specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present invention should be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims. The information disclosed in the background art part of this article is only intended to deepen the understanding of the overall background technology of the present invention, and should not be regarded as an admission or any form of implication that this information constitutes the prior art already known to those skilled in the art.
Claims
1. A MTU model training method, characterized in that: include: Build the MTU model, including: hybrid multi-resolution visual encoder, projection layer and second largest language model; The MTU data is generated by decoupling the MTU data into a method, and the MTU model is trained. The MTU data includes a table image, a question and an answer corresponding to each table data. The training steps include: for each MTU data, the resolution of the table image is increased and decreased respectively, and the obtained images are called high-resolution images and low-resolution images; the hybrid multi-resolution visual encoder extracts local features from the text area in the high-resolution image, extracts the spatial context relationship between cells from the low-resolution image, obtains global features, and fuses the local features with the global features to obtain mixed-resolution visual features; the projection layer converts the mixed-resolution visual features into the embedding space of the second largest language model, and outputs the converted visual features; the second largest language model outputs the answer according to the converted visual features and the corresponding question in the MTU data; a loss function is constructed by combining the answer output by the second largest language model with the answer contained in the MTU data, and the MTU model is optimized by using the loss function.
2. The MTU model training method according to claim 1, characterized in that: include: The MTU data decoupling method comprises: Collect tabular data; By using data enhancement technology, different table data are processed in a diversified manner to obtain visually diverse table images, wherein each table data corresponds to a table image; Generate prompt words corresponding to each table data based on the set question category, and guide the first language model to output questions and answers; The table image, question and answer corresponding to each table data are integrated to form MTU data, where the MTU is multimodal table understanding.
3. A MTU model training method according to claim 2, characterized in that: Also includes: Preprocessing the collected table data includes: converting the table data into table codes containing table content and structure; filtering the table codes and retaining the table codes whose number of rows and columns meets the set requirements; Using the table codes obtained through preprocessing, the corresponding table images, questions and answers are generated.
4. The MTU model training method according to claim 2, characterized in that: The diversified processing of different table data by using data enhancement technology includes: diversified processing of table attributes and table layouts in different table data.
5. The MTU model training method according to claim 2, characterized in that: The generating of prompt words corresponding to each table data in combination with the set question category includes: The set problem category is the MTU problem type, and the problem category is subdivided into specific subcategories; For each table data, a prompt word generator randomly selects several subcategories to generate corresponding prompt words.
6. The MTU model training method according to claim 2 or 5, characterized in that: The prompt word includes a specified output format, guiding the first language model to output questions and answers in a corresponding format, and the answer output by the first language model includes two answers, one is an answer including a reasoning process, and the other is a direct result.
7. An MTU model training system, characterized in that: include: Model building unit, used to build the MTU model, including: hybrid multi-resolution visual encoder, projection layer and second largest language model; A model training unit is used to use the MTU data decoupling method to generate MTU data and train the MTU model. The MTU data includes a table image, a question and an answer corresponding to each table data; the training steps include: for each MTU data, the resolution of the table image is respectively increased and decreased, and the obtained images are called high-resolution images and low-resolution images; the hybrid multi-resolution visual encoder extracts local features from the text area in the high-resolution image, extracts the spatial context relationship between cells from the low-resolution image, obtains global features, and fuses the local features with the global features to obtain mixed-resolution visual features; the projection layer converts the mixed-resolution visual features to the embedding space of the second largest language model and outputs the converted visual features; the second largest language model outputs the answer according to the converted visual features and the corresponding question in the MTU data; a loss function is constructed by combining the answer output by the second largest language model with the answer contained in the MTU data, and the MTU model is optimized using the loss function.
8. A processing device, characterized in that: include: one or more processors; A memory for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 6.
9. A readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
Model training data set generation method and device, computer equipment, readable storage medium and program product
CN119600619A