Table processing method and device, medium, equipment and product
The multimodal large language model uses end-to-end processing of table images, solving the problem of inaccurate table processing in the prior art, achieving high-accurate table structure detection and cell content recognition, and generating editable table files.
Patent Information
- Application Number
- CN202510560670.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-29
- Publication Date
- 2025-08-01
AI Technical Summary
The table files converted from the table processing method in the prior art are not accurate enough, and the multi-module processing method can easily lead to error accumulation, making it difficult to accurately identify borderless and complex nested table structures.
The first multimodal large language model is used to process the table images end-to-end, and feature information is extracted through a visual encoder, and table structure and cell content information are generated by combining the input projector and the large language model to generate corresponding table codes.
Improve the accuracy of table processing, avoid error accumulation in multi-module processing, and accurately identify and generate standard form table files.
Smart Images

Figure CN120411999A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the technical field of data processing, and specifically, to a table processing method, apparatus, medium, device, and product. Background Art
[0002] A table is a means of organizing and arranging data, which can summarize a large amount of data and facilitate the search and comparison of data. In some cases, the table exists in the form of an image. In order to be able to edit the table, it is necessary to process the table image to convert it into an editable table file, so as to meet the application requirements such as data analysis and information extraction. However, the table file converted by the table processing method in the related technology may not be accurate enough. Summary of the Invention
[0003] This Summary of the Invention section is provided to introduce concepts in a brief form, which will be described in detail in the following Detailed Description section. This Summary of the Invention section is not intended to identify key features or essential features of the claimed technical solution, nor is it intended to be used to limit the scope of the claimed technical solution.
[0004] In a first aspect, the present disclosure provides a table processing method, the method including: Obtaining a table image of a first table; Processing the table image based on a first multimodal large language model to obtain a first code, where the first code is used to generate a second table corresponding to the first table, the first multimodal large language model is used to determine first table information of the first table according to the table image, and generate the first code according to the first table information, and the first table information includes table structure information and / or cell content information of the first table; Generating the second table according to the first code.
[0005] In a second aspect, the present disclosure provides a table processing apparatus, the apparatus including: An obtaining module, configured to obtain a table image of a first table; A processing module, configured to process the table image based on a first multimodal large language model to obtain a first code, where the first code is used to generate a second table corresponding to the first table, the first multimodal large language model is used to determine first table information of the first table according to the table image, and generate the first code according to the first table information, and the first table information includes table structure information and / or cell content information of the first table; A first generating module, configured to generate the second table according to the first code.
[0006] In a third aspect, the present disclosure provides a computer-readable medium having a computer program stored thereon, and when the computer program is executed by a processing device, the steps of the table processing method provided in the first aspect of the present disclosure are implemented.
[0007] In a fourth aspect, the present disclosure provides an electronic device, including: a storage device having a computer program stored thereon; a processing device configured to execute the computer program in the storage device to implement the steps of the table processing method provided in the first aspect of the present disclosure.
[0008] In a fifth aspect, the present disclosure provides a computer program product including a computer program, and when the computer program is executed by a processor, the steps of the table processing method provided in the first aspect of the present disclosure are implemented.
[0009] Through the above technical solution, the table image of the first table is processed based on the first multi-modal large language model to obtain a first code, and the first code is used to generate a second table corresponding to the first table, and then the second table is generated according to the first code. In the present disclosure, the first multi-modal large language model can complete multiple tasks, including table structure detection, cell content recognition, code generation, etc. In this way, multiple tasks such as table structure detection, cell content recognition, and code generation can all be completed by the first multi-modal large language model, rather than using multiple different models to perform different tasks respectively as in the related art. In the present disclosure, through the first multi-modal large language model, the first code can be output end-to-end according to the table image, avoiding the phenomenon of error accumulation that is likely to occur when using multiple different models. When generating the first code, the first multi-modal large language model can combine the information of the table image and the first table information at the same time and output the first code end-to-end. The processing process can be completed by this one model of the first multi-modal large language model, avoiding the situation of error accumulation between multiple modules and improving the accuracy of table processing.
[0010] Other features and advantages of the present disclosure will be described in detail in the subsequent specific implementation part. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] In combination with the drawings and with reference to the following specific implementation manners, the above and other features, advantages, and aspects of the embodiments of the present disclosure will become more obvious. Throughout the drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic, and the original components and elements are not necessarily drawn to scale. In the drawings: Figure 1 is a flowchart of a table processing method shown according to an exemplary embodiment.
[0012] Figure 2 is a schematic diagram of a table image shown according to an exemplary embodiment.
[0013] Figure 3 The figure is a schematic diagram showing a process of processing a table image using a first multimodal large language model according to an exemplary embodiment.
[0014] Figure 4 The figure is a schematic diagram showing a process of processing a table image using a first multimodal large language model according to another exemplary embodiment.
[0015] Figure 5 It is a schematic diagram exemplarily showing table structure information.
[0016] Figure 6 The figure is a block diagram of a table processing device according to an exemplary embodiment.
[0017] Figure 7 A schematic structural diagram of an electronic device suitable for implementing the embodiments of the present disclosure is shown. DETAILED DESCRIPTION
[0018] The following describes embodiments of the present disclosure in more detail with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments described herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are for illustrative purposes only and are not intended to limit the scope of protection of the present disclosure.
[0019] It should be understood that the various steps described in the method embodiments of the present disclosure may be performed in different orders and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this respect.
[0020] As used herein, the term "including" and its variations are open-ended, i.e., "including but not limited to." The term "based on" means "based, at least in part, on." The term "one embodiment" means "at least one embodiment," the term "another embodiment" means "at least one additional embodiment," and the term "some embodiments" means "at least some embodiments." Other terms are defined in the following description.
[0021] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.
[0022] It should be noted that the modifications of "one" and "multiple" mentioned in this disclosure are illustrative rather than restrictive. Those skilled in the art should understand that, unless otherwise clearly specified in the context, it should be understood as "one or more".
[0023] The names of the messages or information exchanged between multiple devices in the embodiments of this disclosure are only for illustrative purposes and are not used to limit the scope of these messages or information.
[0024] It can be understood that, before using the technical solutions disclosed in the embodiments of this disclosure, the types, usage scopes, usage scenarios, etc. of the personal information involved in this disclosure should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with relevant laws and regulations.
[0025] For example, when responding to receiving an active request from the user, a prompt message is sent to the user to clearly prompt the user that the operation requested by the user will require obtaining and using the user's personal information. Thus, the user can autonomously choose whether to provide personal information to software or hardware such as an electronic device, an application program, a server, or a storage medium that performs the operations of the technical solutions of this disclosure according to the prompt message.
[0026] As an optional but non-limiting implementation manner, the manner of sending a prompt message to the user in response to receiving an active request from the user can be, for example, in the form of a pop-up window. The prompt message can be presented in text in the pop-up window. In addition, the pop-up window can also carry selection controls for the user to choose "agree" or "disagree" to provide personal information to the electronic device.
[0027] It can be understood that the above process of notifying and obtaining the user's authorization is only illustrative and does not limit the implementation manners of this disclosure. Other manners that meet relevant laws and regulations can also be applied to the implementation manners of this disclosure.
[0028] At the same time, it can be understood that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) should comply with the requirements of the corresponding laws, regulations and related provisions.
[0029] Table recognition is a process of parsing two-dimensional structured table data. In related technologies, during the process of recognizing a table image, usually a table structure detection model is adopted, such as a Convolutional Neural Networks (CNN) model or a Graph Convolutional Networks (GCN) model, to detect table structure information such as the table area, row and column frameworks, and cell boundaries in the table image. Additionally, an OCR (Optical Character Recognition) model is used to recognize the text content in the table. Then, a post-processing model is employed to post-process the table structure information and text content information to obtain a converted table file.
[0030] The methods in related technologies involve multiple different models in the processing process, that is, the table structure is detected by the table structure detection model, the table text content is recognized by the text recognition model, and then the results are integrated by the post-processing model. The processing method of multiple models and multiple modules is likely to lead to incorrect table recognition. For example, it is not easy to detect the structures of borderless tables and complex nested tables. If the table structure detection model makes mistakes in detecting the table structure, the basis for the post-processing model to perform post-processing is not accurate enough, making it easy for errors to accumulate among modules, resulting in inaccurate table recognition.
[0031] The present disclosure provides a table processing method, device, medium, equipment, and product, which processes the table image through a multimodal large language model and end-to-end outputs the code for generating the corresponding table, improving the accuracy of table processing.
[0032] Figure 1 is a flowchart of a table processing method shown according to an exemplary embodiment. This method can be applied to an electronic device, such as a terminal or a server. As Figure 1 shown, the table processing method may include steps 11 to 13.
[0033] In step 11, obtain the table image of the first table.
[0034] The table image of the first table may be an image containing the first table. Figure 2 is a schematic diagram of a table image shown according to an exemplary embodiment. As Figure 2As shown, where AAA, BB, CC, DD, EE, and FF are column headers, which can also be called field names. The column header AAA is divided into two column headers, EE and FF. b1 to b4 are field information under the column header BB, c1 to c4 are field information under the column header CC, d1 to d4 are field information under the column header DD, e1 to e4 are field information under the column header EE, and f1 to f4 are field information under the column header FF.
[0035] The first table in the table image can be in a standard table form. For example, the border of the table has no deformation and is in a horizontal or vertical form. In one embodiment, if the first table in the original image is deformed, and the original image is, for example, an image of the first table taken by a user, image processing can be performed on the original image to obtain the table image, so that the first table in the table image is a table in a standard form, which can improve the accuracy of recognition based on the table image.
[0036] In step 12, the table image is processed based on the first multimodal large language model to obtain the first code.
[0037] Among them, the first code is used to generate a second table corresponding to the first table. The first multimodal large language model is used to determine the first table information of the first table according to the table image and generate the first code according to the first table information. The first table information includes the table structure information and / or cell content information of the first table.
[0038] Multimodal large language models (MLLM, Multimodal Large Language Models) are based on large language models (LLM, Large Language Models) and can process and understand information in different modalities (such as images, audio). In the present disclosure, the first multimodal large language model can be a pre-trained multimodal large language model that can complete multiple tasks, and multiple tasks include, for example, table structure detection, cell content recognition, code generation, etc.
[0039] Among them, the first multimodal large language model can determine the first table information according to the table image. The first table information can include the table structure information of the first table, or the first table information can include the cell content information of the first table, or the first table can also include both the table structure information and cell content information of the first table.
[0040] Exemplarily, the table structure information includes, for example, information related to the table structure such as the rows, columns, and merged cells of the first table. The cell content information can include text information such as words and numbers in each cell of the first table, and can also include images added to the cells. A cell is the smallest unit formed by the intersection of a row and a column. Figure 2Taking the shown table image as an example, Figure 2 In the shown table image, there is no border in the first table, but each data input position can be regarded as a cell. For example, the position where b1 is located corresponds to a cell, and the cell content information of this cell is b1.
[0041] After the first multimodal large language model determines the first table information, it can generate a first code according to the first table information. The first code is used to generate a second table corresponding to the first table. The second table corresponding to the first table can refer to a table with the same table structure and the same cell content as the first table. The second table can be an editable form of the table. The first code is, for example, HTML (Hyper Text Markup Language) code, and the second table can be an editable table presented through a web page. In addition, the first code can also be other types of codes, such as markdown code. The type of the first code is not restricted.
[0042] In step 13, generate the second table according to the first code.
[0043] After obtaining the first code output by the first multimodal large language model, for example, the first code is HTML code, the HTML code can be rendered to generate a second table in the form of a web page. Table 1 below is a schematic diagram of the generated second table.
[0044] Table 1
[0045] Among them, the cells where BB, CC, DD, and AAA are located are in the form of merged cells.
[0046] Through the above technical solution, the table image of the first table is processed based on the first multi-modal large language model to obtain a first code, and the first code is used to generate a second table corresponding to the first table, and then the second table is generated according to the first code. In the present disclosure, the first multi-modal large language model can complete a variety of tasks, including table structure detection, cell content recognition, code generation and other tasks. In this way, a variety of tasks such as table structure detection, cell content recognition, and code generation can all be completed by the first multi-modal large language model, rather than using multiple different models to perform different tasks separately as in the related art. In the present disclosure, through the first multi-modal large language model, the first code can be output end-to-end according to the table image, avoiding the phenomenon of error accumulation that is likely to occur when using multiple different models. When generating the first code, the first multi-modal large language model can combine the information of the table image and the first table information at the same time to output the first code end-to-end, and the processing process can be completed by this one model of the first multi-modal large language model, avoiding the situation of error accumulation between multiple modules and improving the accuracy of table processing.
[0047] Figure 3 is a schematic diagram of a process of a first multi-modal large language model processing a table image shown according to an exemplary embodiment. As Figure 3 shown, in the present disclosure, the first multi-modal large language model includes a vision encoder; the first multi-modal large language model is used to generate the first code in the following manner: Extract feature information from the table image through the vision encoder to obtain visual feature information; Determine the first table information according to the visual feature information, and generate the first code according to the first table information.
[0048] Among them, the vision encoder is a type of modality encoder in the multi-modal large language model. The vision encoder is, for example, CLIP VIT (Contrastive Language-ImagePre-training Vision Transformer).
[0049] After inputting the table image into the first multi-modal large language model, the visual encoder can first extract features from the table image to obtain visual feature information. Among them, the visual feature information can be an image feature vector, which can be expressed as a visual token. The visual feature information can include several visual tokens, and each visual token can be a 1*1024 feature vector. 1024 is the dimension of each visual token feature vector. The value 1024 is only an example, and the dimensions of the feature vectors output by different visual encoders can be different. For example, if the visual feature information includes m visual tokens, then the visual feature information can be an m*1024 feature vector. After obtaining the visual feature information output by the visual encoder, the first multi-modal large language model can determine the first table information based on the visual feature information and generate the first code according to the first table information.
[0050] In this way, through the visual encoder, the feature extraction of the table image is realized. Subsequently, data processing can be carried out based on the extracted visual feature information, enabling the first multi-modal large language model to perceive and process the table image, thereby achieving the goal of generating the first code end-to-end according to the table image.
[0051] As Figure 3 shown, the first multi-modal large language model also includes an input projector; determining the first table information based on the visual feature information and generating the first code according to the first table information can include: Through the input projector, convert the visual feature information into text feature information under a specified dimension; Determine the first table information based on the text feature information and generate the first code according to the first table information.
[0052] Among them, the input projector (Input Projector) is responsible for projecting the visual feature information into the text feature space, that is, aligning the visual feature information to the input space of the first large language model. The input projector can be, for example, a linear projector, or can also be implemented by MLP (Multilayer Perceptron), Cross-Attention, Q-Former (Querying Transformer, a transformer model for vision-language modeling).
[0053] As Figure 3As shown, the first multi-modal large language model may further include a first large language model. Since the first large language model can recognize text feature information, in order to align the visual feature information to the input space of the first large language model, through an input projector, the visual feature information is converted into text feature information that can be understood by the first large language model in a specified dimension, and this text feature information can be represented as text tokens. Among them, the specified dimension can be the dimension of the hidden state of the first large language model, and the dimension of the hidden state is also the hidden_size. The hidden_size parameter usually refers to the number of hidden units in each layer of the large language model, that is, the dimension of the embedding vector or the dimension size of information transmission in the large language model. The dimension of the hidden state of the first large language model is a model parameter of the first large language model, and this dimension of the hidden state can be preset. For example, for the large language models qwen2.5-7b and vicuna-7b, the dimension of the hidden state is 4096. Exemplarily, the input projector can convert 1024-dimensional visual tokens into 4096-dimensional text tokens.
[0054] In this way, after the visual feature information is converted into text feature information in a specified dimension through the input projector, the first large language model can understand and process the text feature information in the specified dimension, so as to generate the first code through the first large language model.
[0055] In the present disclosure, determining the first table information according to the text feature information and generating the first code according to the first table information may include: Generating the first table information and the first code through the first large language model, where the first large language model is used to generate the first table information according to the text feature information and a preset first prompt text, and generate the first code according to the text feature information, the first table information and a preset second prompt text. The first prompt text is used to guide the first large language model to generate the first table information, and the second prompt text is used to guide the first large language model to generate the first code according to the first table information.
[0056] The first prompt text may include a prompt text for guiding the first large language model to generate the table structure information of the first table (hereinafter referred to as the first prompt text 1), and / or a prompt text for guiding the first large language model to generate the cell content information of the first table (hereinafter referred to as the first prompt text 2). The prompt text is used as the prompt of the first large language model. Exemplarily, the first prompt text 1 is, for example, "Parse the structure information of the table in the table image", and the first prompt text 2 is, for example, "Identify the cell content of the table in the table image". The second prompt text is, for example, "Generate HTML code according to the image features of the table image, as well as the structure information and cell content of the table".
[0057] Among them, the first prompt text 1, the first prompt text 2, and the second prompt text can be pre-converted into corresponding text tokens in a specified dimension.
[0058] In one example, taking the first table information including both table structure information and cell content information as an example, the text tokens converted according to the visual feature information, the text tokens corresponding to the first prompt text 1, the text tokens corresponding to the first prompt text 2, and the text tokens corresponding to the second prompt text can be concatenated, and the concatenated features are input into the first large language model. For example, if the number of text tokens converted according to the visual feature information is m, and the total number of text tokens corresponding to each of the three prompt texts is n, taking the dimension of the text token as 4096 as an example, the feature vector of (m + n) * 4096 is input into the first large language model. The first large language model can first generate the table structure information and cell content information of the first table, and then generate the first code according to the text feature information, the table structure information and cell content information of the first table.
[0059] In another example, the text tokens converted according to the visual feature information, the text tokens corresponding to the first prompt text 1, and the text tokens corresponding to the first prompt text 2 can also be concatenated first, and the concatenated features are input into the first large language model. After the first large language model generates the table structure information and cell content information of the first table, the text tokens corresponding to the second prompt text are input into the first large language model, so that the first large language model generates the first code according to the text feature information, the table structure information and cell content information of the first table.
[0060] In this way, converting the visual feature information into text feature information in a specified dimension can align the visual feature information of the table image to the input space of the first large language model, so that the understanding ability and reasoning ability of the first large language model can be utilized to guide the first large language model to complete the tasks of table structure detection, cell content recognition and code generation, and obtain the first code generated end-to-end by the first large language model.
[0061] Figure 4 It is a schematic diagram of a process of processing a table image by a first multimodal large language model shown according to another exemplary embodiment. In this embodiment, the table image includes multiple first table images of a first table, and the image resolutions of the multiple first table images are different; among them, the visual encoder extracts features from the table image to obtain visual feature information, which may include: For each first table image, feature extraction is performed on the first table image through the visual encoder corresponding to the first table image to obtain the visual feature information of the first table image.
[0062] As Figure 4 shown, taking two first table images as examples, they are a low-resolution first table image (such as a resolution of 224×224) and a high-resolution first table image (such as a resolution of 2560×1080). Figure 4 The embodiments shown are only examples, and the number of first table images is not limited. Exemplarily, the resolution of the table images of the first table can be adjusted, such as upsampling and downsampling, to obtain multiple first table images with different resolutions.
[0063] Since the image resolutions of multiple first table images are different, different visual encoders can be used to perform feature extraction on the first table images. Due to the quadratic growth of the computational complexity of the self-attention mechanism in the VIT-based encoder, it can usually only process low-resolution images. As an example, as Figure 4 shown, feature extraction is performed on the low-resolution first table image through visual encoder 1 to obtain the visual feature information 1 of the low-resolution first table image. Visual encoder 1 can be, for example, a CLIP VIT encoder. Feature extraction is performed on the high-resolution first table image through visual encoder 2 to obtain the visual feature information 2 of the high-resolution first table image. Visual encoder 2 can be, for example, a ConvNext-L encoder based on convolution, a Swin Transformer encoder, etc. The dimensions of visual feature information 1 and visual feature information 2 can be the same or different.
[0064] After that, for each first table image, the visual feature information of the first table image can be converted into text feature information in a specified dimension through the input projector corresponding to the first table image. Exemplarily, as Figure 4 shown, the input projector corresponding to the low-resolution image, that is, input projector 1 is, for example, a linear projector. The input projector corresponding to the high-resolution image, that is, input projector 2 is, for example, a projector based on Q-Former.
[0065] Input projector 1 is used to convert visual feature information 1 into text feature information 1 in a specified dimension. Input projector 2 is used to convert visual feature information 2 into text feature information 2 in a specified dimension. After that, the first table information can be determined according to the text feature information corresponding to each first table image respectively, and the first code can be generated according to the first table information.
[0066] In one example, the text feature information 1, text feature information 2, and the text tokens of the prompt text can be concatenated. For example, the number of text tokens corresponding to the low-resolution image in text feature information 1 is i, the number of text tokens corresponding to the high-resolution image in text feature information 2 is j, and the prompt text has been introduced above. For example, the total number of text tokens corresponding to each of the three prompt texts is n. Taking the dimension of the text token as 4096 as an example, the feature vector of (i + j + n) * 4096 after feature concatenation is input into the first large language model. Among them, the method for the first large language model to perform table structure detection, cell content recognition, and code generation can refer to the above introduction.
[0067] In the present disclosure, Figure 4 The illustrated embodiments can be used as preferred embodiments. The first table images with different resolutions can provide the first large language model with information from different perspectives, so that when the first large language model performs table structure detection, cell content recognition, and code generation tasks, it has richer analysis and reasoning bases. For example, the visual feature information extracted from the high-resolution first table image contains richer and clearer local image features, and the visual feature information extracted from the low-resolution first table image can retain the overall layout information of the first table. In this way, the first large language model can perform more accurate analysis and reasoning based on the text feature information corresponding to the first table images with different resolutions, thereby improving the accuracy of table processing.
[0068] In the present disclosure, the cell content information includes the content in each cell of the first table. Taking Figure 2 the illustrated table image as an example, the cell content information may include: AAA, BB, CC, DD, EE, FF, b1 to b4, c1 to c4, d1 to d4, e1 to e4, f1 to f4. Among them, b1 to b4 include b1, b2, b3, and b4, and other information is similar.
[0069] The table structure information may include at least one of the following: the position information of the rectangular frame corresponding to each row in the first table in the table image; the position information of the rectangular frame corresponding to each column in the first table in the table image; the position information of the rectangular frame corresponding to the column header of the first table in the table image; the position information of the rectangular frame corresponding to the merged cells in the first table in the table image.
[0070] Figure 5 is an exemplary schematic diagram related to the table structure information. As Figure 5As shown, the table image can be cropped and a coordinate system can be established. For example, point o is the origin of coordinates, the x-axis is the horizontal direction, and the y-axis is the vertical direction. The position information of the rectangular frame in the table image can include the normalized coordinate information of the upper left corner of the rectangular frame and the normalized coordinate information of the lower right corner of the rectangular frame. This table image can be the table image before resolution processing.
[0071] As Figure 5 shown, the position information of the rectangular frame corresponding to each row in the first table in the table image can include: row[0.016,0.433,0.970,0.455] row[0.016,0.455,0.970,0.488] row[0.016,0.488,0.970,0.513] row[0.016,0.513,0.970,0.528] row[0.016,0.528,0.970,0.543] row[0.016,0.543,0.970,0.558] Among them, row represents a row. row[0.016,0.433,0.970,0.455] is the position information of the rectangular frame corresponding to the row where AAA is located. This rectangular frame is a rectangular frame with point 21 as the upper left corner and point 24 as the lower right corner. 0.016 is the abscissa of point 21, 0.433 is the ordinate of point 21, 0.970 is the abscissa of point 24, and 0.455 is the ordinate of point 24. row[0.016,0.455,0.970,0.488] is the position information of the rectangular frame corresponding to the row where BB, CC, DD, EE, and FF are located. This rectangular frame is a rectangular frame with point 27 as the upper left corner and point 25 as the lower right corner. 0.016 is the abscissa of point 27, 0.455 is the ordinate of point 27, 0.970 is the abscissa of point 25, and 0.488 is the ordinate of point 25. The dotted line shown in the figure is for facilitating the marking of point 27 and is not used to represent the border line in the table. row[0.016,0.488,0.970,0.513] is the position information of the rectangular frame corresponding to the row where b1, c1, d1, e1, and f1 are located. This rectangular frame is a rectangular frame with point 28 as the upper left corner and point 26 as the lower right corner. 0.016 is the abscissa of point 28, 0.488 is the ordinate of point 28, 0.970 is the abscissa of point 26, and 0.513 is the ordinate of point 26. The position information of the corresponding rectangular frames of other rows is similar.
[0072] As Figure 5As shown, the position information of the rectangular boxes corresponding to each column in the first table in the table image may include: column[0.016,0.433.0.185,0.558] column[0.185,0.433,0.342,0.558] column[0.342,0.433,0.594,0.558] column[0.594,0.433,0.795,0.558] column[0,795,0.433,0.970,0.558] Among them, column represents a column. column[0.016,0.433.0.185,0.558] is the position information of the rectangular box corresponding to the column where BB, b1, b2, b3, and b4 are located. This rectangular box is a rectangular box with point 21 as the upper left corner and point 29 as the lower right corner. 0.185 represents the abscissa of point 29, and 0.558 represents the ordinate of point 29. column[0.594,0.433,0.795,0.558] represents the position information of the rectangular box corresponding to the column where EE, e1, e2, e3, and e4 are located. This rectangular box is a rectangular box with point 23 as the upper left corner and point 30 as the lower right corner. 0.594 is the abscissa of point 23, 0.433 is the ordinate of point 23, 0.795 is the abscissa of point 30, and 0.558 is the ordinate of point 30. The position information of the rectangular boxes corresponding to other columns is similar.
[0073] Such as Figure 5 As shown, the position information of the rectangular box corresponding to the column header of the first table in the table image may include: column header[0.016,0.433,0.970,0.488] Among them, column header represents the column header. AAA, BB, CC, DD, EE, and FF are the column headers of the first table. The rectangular box corresponding to the column header is a rectangular box with point 21 as the upper left corner and point 25 as the lower right corner.
[0074] Such as Figure 5 As shown, the position information of the rectangular box corresponding to the merged cells in the first table in the table image may include: spanning cell[0.016,0.433,0.185,0.488] spanning cell[0.185,0.433,0.342,0.488] spanning cell[0.342,0.433,0.594,0.488] spanning cell[0.594,0.433,0.970,0.455] Among them, "spanning" means merging cells. The position information of the rectangular box corresponding to the spanning cell [0.016, 0.433, 0.185, 0.488] where BB is located is the rectangular box formed with point 21 as the upper left corner and point 22 as the lower right corner. 0.185 is the abscissa of point 22, and 0.488 is the ordinate of point 22. The position information of the rectangular box corresponding to the spanning cell [0.594, 0.433, 0.970, 0.455] where AAA is located is the rectangular box formed with point 23 as the upper left corner and point 24 as the lower right corner. 0.970 is the abscissa of point 24, and 0.455 is the ordinate of point 24. The position information of the rectangular box corresponding to the spanning cell where CC is located and the rectangular box corresponding to the spanning cell where DD is located is similar.
[0075] In the above examples, the abscissa and ordinate values of each point are relative coordinates, that is, the normalized coordinate information. The value of the abscissa is the actual abscissa of the point in the table image divided by the width of the table image, and the width of the table image is the length of the table image in the x-axis direction. The value of the ordinate is the actual ordinate of the point in the table image divided by the height of the table image, and the height of the table image is the length of the table image in the y-axis direction. Taking point 21 as an example, the abscissa 0.016 of point 21 is obtained by dividing the actual abscissa of point 21 in the table image by the width of the table image, and the ordinate 0.433 of point 21 is obtained by dividing the actual ordinate of point 21 in the table image by the height of the table image.
[0076] In this way, through the position information of rows, columns, and merged cells in the first table, the structure of the first table can be accurately identified. Moreover, when the first large language model recognizes the cell content information, it can also recognize the position of the cell content information in the image, thus corresponding the cell content to the table structure.
[0077] In the present disclosure, the first multi-modal large language model can be trained in the following manner: Obtain training samples, where the training samples include sample images of sample tables, second table information of the annotated sample tables, and second codes for generating sample tables. The second table information includes table structure information and / or cell content information of the annotated sample tables; Process the sample image through a multimodal large language model to obtain the third table information of the predicted sample table and the third code for generating the sample table. The third table information includes the table structure information and / or cell content information of the predicted sample table. Update the parameters of the multimodal large language model according to the difference information between the third table information and the second table information, and the difference information between the third code and the second code. When the training is completed, obtain the first multimodal large language model.
[0078] Among them, the training samples can be obtained from the PubTables-1M dataset, the FinTabNet dataset, and the PubTabNet dataset. The number of training samples is multiple. The table structure information of the labeled sample table may include at least one of the following: the position information of the rectangular box corresponding to each row in the sample table in the sample image; the position information of the rectangular box corresponding to each column in the sample table in the sample image; the position information of the rectangular box corresponding to the column header of the sample table in the sample image; the position information of the rectangular box corresponding to the merged cells in the sample table in the sample image. The cell content information of the labeled sample table may include information such as text, numbers, and letters in each cell of the sample table.
[0079] The multimodal large language model for training may include a basic visual encoder, an input projector, and a large language model. The process of processing the sample image through the multimodal large language model is similar to the process of processing the table image by the first multimodal large language model introduced above. Among them, the visual encoder may first extract the feature information of the sample image to obtain the visual feature information of the sample image, and then through the input projector, convert the visual feature information of the sample image into text feature information in a specified dimension. According to the prompt text and the text feature information corresponding to the sample image, the third table information of the predicted sample table and the third code for generating the sample table are obtained through the large language model. The third table information may include the table structure information and / or cell content information of the predicted sample table. The type included in the table structure information of the predicted sample table is the same as the type of the table structure information of the above-mentioned labeled sample table.
[0080] According to the difference information between the predicted third table information and the labeled second table information, and the difference information between the predicted third code and the labeled second code, the loss value can be calculated. The method of calculating the loss value can refer to the related technology. Then, the parameters of the multimodal large language model can be updated according to the loss value. The update method can be full parameter fine-tuning. Among them, updating the parameters of the multimodal large language model may refer to updating the parameters of the visual encoder, the input projector, and the large language model simultaneously.
[0081] The training stop condition can be, for example, that multiple training samples have been traversed, or the loss value has converged, indicating that the training is completed. In the case of training completion, the first multimodal large language model is obtained. In this way, the first multimodal large language model obtained through training has the ability to complete table structure detection, cell content recognition, and code generation, and the table structure information output by the first multimodal large language model obtained through training, that is, the position information of the rectangles corresponding to rows, columns, and merged cells, conforms to the required format.
[0082] Through the above technical solution, when generating the first code, the first multimodal large language model can combine the information of the table image, table structure information, and cell content information at the same time, and end-to-end output the first code, simplifying the table processing process, and the processing process can be completed by this single model of the first multimodal large language model, avoiding the accumulation of errors between multiple modules and improving the accuracy of table processing.
[0083] Based on the same inventive concept, the present disclosure also provides a table processing device, Figure 6 which is a block diagram of a table processing device shown according to an exemplary embodiment, as Figure 6 shown. The table processing device 60 may include: An acquisition module 61, configured to acquire a table image of a first table; A processing module 62, configured to process the table image based on the first multimodal large language model to obtain a first code, where the first code is used to generate a second table corresponding to the first table, and the first multimodal large language model is configured to determine first table information of the first table according to the table image and generate the first code according to the first table information, and the first table information includes table structure information and / or cell content information of the first table; A first generation module 63, configured to generate the second table according to the first code.
[0084] Optionally, the first multimodal large language model includes a visual encoder; the first multimodal large language model is configured to generate the first code through the following modules: An extraction module, configured to extract feature information from the table image through the visual encoder to obtain visual feature information; A second generation module, configured to determine the first table information according to the visual feature information and generate the first code according to the first table information.
[0085] Optionally, the table image includes multiple first table images of the first table, and the image resolutions of the multiple first table images are different; The extraction module includes: An extraction sub-module, configured to, for each of the first table images, extract features of the first table image through the visual encoder corresponding to the first table image, to obtain the visual feature information of the first table image.
[0086] Optionally, the first multimodal large language model further includes an input projector; The second generation module includes: A conversion sub-module, configured to convert the visual feature information into text feature information in a specified dimension through the input projector; A generation sub-module, configured to determine the first table information according to the text feature information, and generate the first code according to the first table information.
[0087] Optionally, the first multimodal large language model further includes a first large language model, and the specified dimension is the dimension of the hidden state of the first large language model; The generation sub-module is configured to: Generate the first table information and the first code through the first large language model, where the first large language model is configured to generate the first table information according to the text feature information and a preset first prompt text, and generate the first code according to the text feature information, the first table information and a preset second prompt text, the first prompt text is used to guide the first large language model to generate the first table information, and the second prompt text is used to guide the first large language model to generate the first code according to the first table information.
[0088] Optionally, the table structure information of the first table includes at least one of the following: the position information of the rectangular frame corresponding to each row in the first table in the table image; the position information of the rectangular frame corresponding to each column in the first table in the table image; the position information of the rectangular frame corresponding to the column title of the first table in the table image; the position information of the rectangular frame corresponding to the merged cells in the first table in the table image.
[0089] Optionally, the first multimodal large language model is trained through the following module: A sample acquisition module, configured to acquire training samples, where the training samples include sample images of sample tables, the second table information of the labeled sample tables, and the second codes used to generate the sample tables, and the second table information includes the table structure information and / or cell content information of the labeled sample tables; A sample processing module for processing the sample image through a multimodal large language model to obtain the third table information of the predicted sample table and the third code for generating the sample table, where the third table information includes the table structure information and / or cell content information of the predicted sample table; An update module for updating the parameters of the multimodal large language model according to the difference information between the third table information and the second table information, and the difference information between the third code and the second code; An obtaining module for obtaining the first multimodal large language model when the training is completed.
[0090] Next, refer to Figure 7 , which shows a schematic structural diagram of an electronic device 600 suitable for implementing the embodiments of the present disclosure. The terminal device in the embodiments of the present disclosure may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Tablet Computers), PMPs (Portable Multimedia Players), in-vehicle terminals (such as in-vehicle navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 7 The electronic device shown is only an example and should not impose any limitations on the functions and usage scopes of the embodiments of the present disclosure.
[0091] As Figure 7 shown, the electronic device 600 may include a processing device (such as a central processing unit, a graphics processing unit, etc.) 601, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 602 or the program loaded from the storage device 608 into the random access memory (RAM) 603. In the RAM 603, various programs and data required for the operation of the electronic device 600 are also stored. The processing device 601, the ROM 602, and the RAM 603 are connected to each other through a bus 604. The input / output (I / O) interface 605 is also connected to the bus 604.
[0092] Generally, the following devices may be connected to the I / O interface 605: an input device 606 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 607 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 608 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 609. The communication device 609 can allow the electronic device 600 to communicate with other devices wirelessly or wiredly to exchange data. Although Figure 7 shows the electronic device 600 with various devices, it should be understood that it is not required to implement or have all the shown devices. More or fewer devices may be alternatively implemented or had.
[0093] In particular, according to an embodiment of the present disclosure, the processes described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product that includes a computer program carried on a non-transitory computer-readable medium, and the computer program includes program code for performing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication device 609, or installed from the storage device 608, or installed from the ROM 602. When the computer program is executed by the processing device 601, the above functions defined in the method of the embodiment of the present disclosure are performed.
[0094] It should be noted that the above computer-readable medium in the present disclosure can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, the computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In the present disclosure, the computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, in which computer-readable program code is carried. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium, and the computer-readable signal medium can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted by any suitable medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.
[0095] In some embodiments, the client and the server can communicate using any currently known or future-developed network protocol such as HTTP (HyperText Transfer Protocol), and can be interconnected with digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include local area networks ("LAN"), wide area networks ("WAN"), the Internet (e.g., the Internet), and end-to-end networks (e.g., ad hoc end-to-end networks), as well as any currently known or future-developed networks.
[0096] The above computer-readable medium can be included in the above electronic device; or can exist separately without being assembled into the electronic device.
[0097] The above computer-readable medium carries one or more programs, which, when executed by the electronic device, cause the electronic device to: obtain a table image of a first table; Process the table image based on a first multimodal large language model to obtain a first code, where the first code is used to generate a second table corresponding to the first table, and the first multimodal large language model is used to determine first table information of the first table according to the table image and generate the first code according to the first table information, and the first table information includes table structure information and / or cell content information of the first table; Generate the second table according to the first code.
[0098] Computer program code for performing the operations of the present disclosure can be written in one or more programming languages or combinations thereof. The above programming languages include, but are not limited to, object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (e.g., by using an Internet service provider to connect through the Internet).
[0099] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a part of code, which contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions noted in the blocks may occur in a different order than noted in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0100] The modules described in the embodiments of the present disclosure can be implemented in software or in hardware. Among them, the name of the module does not constitute a limitation on the module itself in some cases. For example, the acquisition module can also be described as "the module for acquiring the form image".
[0101] The functions described above herein can be performed, at least in part, by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that can be used include: Field Programmable Gate Array (FPGA), Application Specific Integrated Circuit (ASIC), Application Specific Standard Product (ASSP), System on Chip (SOC), Complex Programmable Logic Device (CPLD), and so on.
[0102] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a Random Access Memory (RAM), a Read Only Memory (ROM), an Erasable Programmable Read Only Memory (EPROM or Flash Memory), an optical fiber, a portable Compact Disc Read Only Memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0103] According to one or more embodiments of the present disclosure, Example 1 provides a form processing method, the method comprising: Obtain the table image of the first table; Process the table image based on the first multimodal large language model to obtain a first code, wherein the first code is used to generate a second table corresponding to the first table, and the first multimodal large language model is used to determine first table information of the first table according to the table image and generate the first code according to the first table information, and the first table information includes table structure information and / or cell content information of the first table; Generate the second table according to the first code.
[0104] According to one or more embodiments of the present disclosure, Example 2 provides the method of Example 1, and the first multimodal large language model includes a visual encoder; the first multimodal large language model is used to generate the first code in the following manner: Extract features from the table image through the visual encoder to obtain visual feature information; Determine the first table information according to the visual feature information and generate the first code according to the first table information.
[0105] According to one or more embodiments of the present disclosure, Example 3 provides the method of Example 2, and the table image includes multiple first table images of the first table, and the image resolutions of the multiple first table images are different; The extracting features from the table image through the visual encoder to obtain visual feature information includes: For each of the first table images, extract features from the first table image through the visual encoder corresponding to the first table image to obtain the visual feature information of the first table image.
[0106] According to one or more embodiments of the present disclosure, Example 4 provides the method of Example 2, and the first multimodal large language model further includes an input projector; The determining the first table information according to the visual feature information and generating the first code according to the first table information includes: Convert the visual feature information into text feature information in a specified dimension through the input projector; Determine the first table information according to the text feature information and generate the first code according to the first table information.
[0107] According to one or more embodiments of the present disclosure, Example 5 provides the method of Example 4, and the first multimodal large language model further includes a first large language model, and the specified dimension is the dimension of the hidden state of the first large language model; Determining the first table information according to the text feature information and generating the first code according to the first table information includes: Generating the first table information and the first code through the first large language model, where the first large language model is used to generate the first table information according to the text feature information and a preset first prompt text, and generate the first code according to the text feature information, the first table information and a preset second prompt text. The first prompt text is used to guide the first large language model to generate the first table information, and the second prompt text is used to guide the first large language model to generate the first code according to the first table information.
[0108] According to one or more embodiments of the present disclosure, Example 6 provides the method of Example 1. The table structure information of the first table includes at least one of the following: the position information of the rectangular box corresponding to each row in the first table in the table image; the position information of the rectangular box corresponding to each column in the first table in the table image; the position information of the rectangular box corresponding to the column header of the first table in the table image; the position information of the rectangular box corresponding to the merged cells in the first table in the table image.
[0109] According to one or more embodiments of the present disclosure, Example 7 provides the method of Example 1. The first multimodal large language model is trained in the following manner: Obtain training samples, where the training samples include the sample images of the sample tables, the second table information of the annotated sample tables, and the second codes for generating the sample tables. The second table information includes the table structure information and / or cell content information of the annotated sample tables. Process the sample images through a multimodal large language model to obtain the predicted third table information of the sample tables and the third codes for generating the sample tables. The third table information includes the predicted table structure information and / or cell content information of the sample tables. Update the parameters of the multimodal large language model according to the difference information between the third table information and the second table information, and the difference information between the third code and the second code. When the training is completed, the first multimodal large language model is obtained.
[0110] According to one or more embodiments of the present disclosure, Example 8 provides a table processing device, and the device includes: An acquisition module, configured to acquire the table image of the first table; A processing module for processing the table image based on a first multimodal large language model to obtain a first code, where the first code is used to generate a second table corresponding to the first table, and the first multimodal large language model is used to determine first table information of the first table according to the table image and generate the first code according to the first table information, and the first table information includes table structure information and / or cell content information of the first table; A first generation module for generating the second table according to the first code.
[0111] According to one or more embodiments of the present disclosure, Example 9 provides a computer-readable medium having a computer program stored thereon, and when the computer program is executed by a processing device, the steps of the method described in any one of Examples 1 to 7 are implemented.
[0112] According to one or more embodiments of the present disclosure, Example 10 provides an electronic device, including: A storage device having a computer program stored thereon; A processing device for executing the computer program in the storage device to implement the steps of the method described in any one of Examples 1 to 7.
[0113] According to one or more embodiments of the present disclosure, Example 11 provides a computer program product including a computer program, and when the computer program is executed by a processor, the steps of the method described in any one of Examples 1 to 7 are implemented.
[0114] The above description is only a preferred embodiment of the present disclosure and an explanation of the applied technical principles. Those skilled in the art should understand that the scope of disclosure involved in the present disclosure is not limited to the technical solutions formed by the specific combination of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above disclosure concept. For example, technical solutions formed by mutually replacing the above features with (but not limited to) technical features having similar functions disclosed in the present disclosure.
[0115] In addition, although the operations are depicted in a particular order, this should not be construed as requiring that the operations be performed in the particular order shown or in sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, although a number of specific implementation details are included in the above discussion, these should not be construed as limiting the scope of the present disclosure. Certain features described in the context of separate embodiments may also be implemented combinatorially in a single embodiment. Conversely, the various features described in the context of a single embodiment may also be implemented separately or in any suitable sub-combination in multiple embodiments.
[0116] Although the subject matter has been described in language specific to structural features and / or methodological logical acts, it is to be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. On the contrary, the specific features and acts described above are merely example forms for implementing the claims. Regarding the apparatus in the above embodiments, the specific manner in which each module performs operations has been described in detail in the embodiments related to the method and will not be elaborated herein.
Claims
1. A table processing method, characterized in that, The method includes: Obtaining a table image of a first table; Processing the table image based on a first multimodal large language model to obtain a first code, wherein the first code is used to generate a second table corresponding to the first table, and the first multimodal large language model is used to determine first table information of the first table according to the table image and generate the first code according to the first table information, and the first table information includes table structure information and / or cell content information of the first table; Generating the second table according to the first code.
2. The method according to claim 1, wherein The first multimodal large language model includes a visual encoder; the first multimodal large language model is used to generate the first code in the following manner: Performing feature extraction on the table image through the visual encoder to obtain visual feature information; Determining the first table information according to the visual feature information and generating the first code according to the first table information.
3. The method according to claim 2, wherein The table image includes multiple first table images of the first table, and the image resolutions of the multiple first table images are different; The performing feature extraction on the table image through the visual encoder to obtain visual feature information includes: For each of the first table images, performing feature extraction on the first table image through the visual encoder corresponding to the first table image to obtain the visual feature information of the first table image.
4. The method according to claim 2, wherein The first multimodal large language model further includes an input projector; The determining the first table information according to the visual feature information and generating the first code according to the first table information includes: Converting the visual feature information into text feature information in a specified dimension through the input projector; Determining the first table information according to the text feature information and generating the first code according to the first table information.
5. The method according to claim 4, wherein The first multimodal large language model further includes a first large language model, and the specified dimension is the dimension of the hidden state of the first large language model; The determining the first table information according to the text feature information and generating the first code according to the first table information includes: Generating the first table information and the first code through the first large language model, wherein the first large language model is used to generate the first table information according to the text feature information and a preset first prompt text, and generate the first code according to the text feature information, the first table information and a preset second prompt text, the first prompt text is used to guide the first large language model to generate the first table information, and the second prompt text is used to guide the first large language model to generate the first code according to the first table information.
6. The method according to claim 1, wherein The table structure information of the first table includes at least one of the following: the position information of the rectangular box corresponding to each row in the first table in the table image; the position information of the rectangular box corresponding to each column in the first table in the table image; the position information of the rectangular box corresponding to the column header of the first table in the table image; the position information of the rectangular box corresponding to the merged cells in the first table in the table image.
7. The method according to claim 1, wherein The first multi-modal large language model is trained in the following manner: Obtain training samples, where the training samples include sample images of sample tables, the second table information of the annotated sample tables, and the second code for generating the sample tables. The second table information includes the table structure information and / or cell content information of the annotated sample tables; Process the sample images through a multi-modal large language model to obtain the third table information of the predicted sample tables and the third code for generating the sample tables. The third table information includes the table structure information and / or cell content information of the predicted sample tables; Update the parameters of the multi-modal large language model according to the difference information between the third table information and the second table information, and the difference information between the third code and the second code; When the training is completed, the first multi-modal large language model is obtained.
8. A table processing device, characterized in that, The device includes: An acquisition module for acquiring a table image of a first table; A processing module for processing the table image based on a first multi-modal large language model to obtain a first code, where the first code is used to generate a second table corresponding to the first table. The first multi-modal large language model is used to determine the first table information of the first table according to the table image and generate the first code according to the first table information. The first table information includes the table structure information and / or cell content information of the first table; A first generation module for generating the second table according to the first code.
9. A computer-readable medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processing device, it implements the steps of the method according to any one of claims 1 to 7.
10. An electronic device, characterized in that, Including: A storage device on which a computer program is stored; A processing device for executing the computer program in the storage device to implement the steps of the method according to any one of claims 1 to 7.
11. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 7.
Citation Information
Cited By
Sample data acquisition method, model training method, electronic equipment and medium
CN121527796A