Table identification method and system based on multi-modal large model, medium and equipment
By combining a multimodal large model with table images and coordinate features, a unified semantic vector is generated, which solves the problem of insufficient modeling of the topological relationship of merged cells, realizes accurate identification and structural reconstruction of cross-cell semantic association, and improves the accuracy of table recognition.
Patent Information
- Application Number
- CN202511071350.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-31
- Publication Date
- 2025-11-18
AI Technical Summary
Existing multimodal table recognition technologies suffer from insufficient modeling of the topological relationships of merged cells when dealing with complex scenarios, resulting in the loss of semantic associations across cells. Furthermore, the weight allocation of a single loss function cannot balance the accuracy requirements of text recognition and structural reconstruction.
A table recognition method based on a multimodal large model is adopted. By extracting the cell coordinate matrix and image features of the table, and combining HTML structural tags and a token segmentation reference library, a neural network model is trained to generate a unified semantic vector. The feature vector dimension is adjusted through a fully connected layer to generate a decoding sequence to obtain the HTML text content of the table.
Precise cell topological relationships were established, which improved the accuracy of table recognition and semantic association across cells. The weight allocation of the loss function was optimized, which improved the accuracy of recognizing complex tables and the precision of structure reconstruction.
Smart Images

Figure CN120976949A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of table recognition technology, and in particular to a table recognition method, system, medium and device based on a multimodal large model. Background Technology
[0002] In the fields of digital office and intelligent information processing, table recognition technology serves as a crucial link between unstructured documents and structured data, and its accuracy and robustness directly impact the quality of subsequent data analysis. Currently, traditional image processing-based table recognition methods (such as edge detection and connected component analysis) achieve good results when dealing with simple tables, but they have significant limitations when handling complex scenarios, especially complex text content.
[0003] In recent years, with the rise of multimodal large models, the ability to parse complex tables has been significantly improved by the joint modeling strategy that integrates text semantics, spatial coordinates and visual features. However, existing multimodal table recognition technologies still have the following bottlenecks: (1) insufficient modeling of the topological relationship of merged cells, resulting in the loss of semantic associations across cells; (2) the single loss function weight allocation cannot balance the accuracy requirements of text recognition and structure reconstruction. Summary of the Invention
[0004] To address the problem in existing technologies that fail to accurately identify the topological relationships of merged cells when extracting table information, resulting in the loss of semantic associations across cells, a table recognition method, system, medium, and device based on a multimodal large model are proposed.
[0005] This application provides a table recognition method, system, medium, and device based on a multimodal large model, which adopts the following technical solution:
[0006] On the one hand, this application provides a table recognition method based on a multimodal large model, including the following steps:
[0007] Extract the cell coordinate matrix from the table and generate a coordinate feature map;
[0008] A multimodal large model is established, and image encoding features are concatenated by combining table images and coordinate feature maps to obtain image feature vectors;
[0009] Based on the multimodal large model, the text feature vector is obtained by concatenating the cell coordinate matrix with the set prompt words and text encoding features.
[0010] By fusing text feature vectors and image feature vectors, a unified semantic vector is generated;
[0011] Based on a unified semantic vector, a decoding sequence is generated to obtain the HTML text content of the table.
[0012] To optimize the above technical solution, the specific measures also include:
[0013] Furthermore, a multimodal large model is established, specifically as follows:
[0014] Collect HTML structural tags and text tags;
[0015] Establish a token segmentation reference library based on HTML structural tags;
[0016] The table image, table coordinate feature map, HTML structural tags, text tags, and token segmentation reference library are input into the neural network model for training;
[0017] Calculate the loss function of the neural network model;
[0018]
[0019] Where: L represents the loss value, V represents the vocabulary size, and y i p represents the tag of the token segment at position i. i w represents the probability value corresponding to the prediction result of the i-th position. i This is the training weight factor corresponding to the Token at this position;
[0020] By combining the loss function of the neural network model, the trained neural network model is evaluated. If the loss function calculation result meets the set threshold range, the neural network model is saved as a multimodal large model. If the loss function calculation result does not meet the set threshold range, the neural network model is trained again until the loss function calculation result meets the set threshold range, and then the multimodal large model at this time is output.
[0021] Furthermore, by combining the table image and the coordinate feature map, image encoding features are concatenated to obtain the image feature vector, specifically:
[0022] Extract visual semantic information from the table image and generate a semantic representation vector. The visual semantic information includes text data, line data, and background data.
[0023] Extract relevant features from the coordinate feature map, including the spatial layout information of each cell in the table area in the two-dimensional matrix, the coordinate position of each cell in the image, and the row and column span of each cell;
[0024] The coordinate feature map is encoded, and a structural representation vector is extracted. The structural representation vector contains the table's layout relationship and hierarchical structure.
[0025] The semantic representation vector and the structural representation vector are concatenated dimensionally and fused to form the image feature vector.
[0026] Furthermore, by combining the cell coordinate matrix with the set prompt words, text encoding features are concatenated to obtain the text feature vector, specifically:
[0027] Encode the cell coordinate matrix into structured text information;
[0028] Structured text information is concatenated with the set prompt words to generate enhanced feature words;
[0029] Encode the enhanced feature words to obtain coordinate features. Figure 1 The text feature vectors are consistent.
[0030] Furthermore, the text feature vectors and image feature vectors are fused, specifically as follows:
[0031] Construct a unified multimodal representation space;
[0032] Within the multimodal representation space, text features and image features with the same semantic content are clustered.
[0033] The fully connected layer adjusts the dimensions of the text feature vector and the image feature vector to the same size, and then concatenates and combines the adjusted text feature vector and image feature vector with respect to structure, semantics and task objectives to generate a unified semantic vector.
[0034] Furthermore, based on the unified semantic vector, a decoding sequence is generated to obtain the HTML text content of the table, specifically:
[0035] Based on a unified semantic vector, several token words are generated step by step according to time steps;
[0036] After generating each token segment, the tag type is determined. The tag type includes HTML structure closing tags, HTML structure starting tags, and text content tags.
[0037] Based on the corresponding tag type, dynamically select the category of the next token segmentation and output the corresponding token segmentation;
[0038] The output tokens are segmented to form a decoding sequence, which yields the HTML text content of the table.
[0039] Furthermore, based on the corresponding tag type, select the next token segmentation category and output the corresponding token segmentation, specifically:
[0040] Perform an initial category determination on the current token segment. If the current token segment is a text content label, select a language model to identify the entire word space and output the next token segment.
[0041] If the current token segmentation is an HTML structure tag, then perform another category judgment. If the current token segmentation is an HTML structure start tag, then HTML structure tags are blocked, and the next token segmentation is selected from text tags. If the current token segmentation is an HTML structure closing tag, then non-HTML structure tags are blocked, and the next token segmentation is selected from text tags.
[0042] A token sequence is formed from multiple tokens.
[0043] Furthermore, this application provides a table recognition system based on a multimodal large model, comprising:
[0044] The data acquisition module is used to acquire table images;
[0045] The extraction module is used to extract the cell coordinate matrix from the table image and generate a coordinate feature map;
[0046] The modeling module is used to build large multimodal models.
[0047] The image stitching module is used to stitch together image encoding features based on a multimodal large model, combined with a table image and coordinate feature map, to obtain an image feature vector.
[0048] The text concatenation module is used to concatenate text encoding features based on the multimodal large model, combined with the cell coordinate matrix and the set prompt words, to obtain the text feature vector;
[0049] The fusion module is used to fuse text feature vectors and image feature vectors to generate a unified semantic vector;
[0050] The decoding module is used to generate a decoding sequence based on a unified semantic vector to obtain the HTML text content of the table.
[0051] On another front, this application provides a medium that stores a set of program instructions that can be read by a machine. When the set of program instructions stored in the medium is read and executed by a machine, the machine can implement the aforementioned table recognition method.
[0052] In another aspect, this application provides a device, characterized in that the device includes a processor and a memory connected together; the memory stores a set of program instructions; when the set of program instructions stored in the memory is read and executed by the processor, the device can implement the aforementioned table recognition method.
[0053] In summary, this application includes at least one of the following beneficial technical effects:
[0054] 1. The table recognition method of this application can establish accurate cell topological relationships, especially semantic relationships across cells, which improves the accuracy of recognition.
[0055] 2. In this application, multiple types of data were collected when collecting tabular data, and different labels were set for different categories of data to facilitate subsequent identification and judgment.
[0056] 3. The multimodal large model established in this application is trained multiple times using pre-set parameters and combined with a loss function, thereby improving the accuracy of the prediction results.
[0057] 4. During the token segmentation process, by making multiple judgments on the current token segmentation type, the type of the next token segmentation can be accurately predicted, so that different probabilities can be assigned according to the type of token segmentation. Attached Figure Description
[0058] Figure 1 This is a flowchart of the steps of the table recognition method based on a multimodal large model according to the present invention;
[0059] Figure 2 This is a flowchart of the cell coordinate matrix extraction process of the present invention;
[0060] Figure 3 This is a flowchart of the feature word encoding process;
[0061] Figure 4 This is a cross-modal fusion framework diagram of the present invention;
[0062] Figure 5 This is a flowchart of the dynamic weight allocation process;
[0063] Figure 6 This is a flowchart of the multimodal model training steps;
[0064] Figure 7 This is a data graph of the multimodal model training process;
[0065] Figure 8 This is a module relationship diagram of a table recognition system based on a multimodal large model;
[0066] Figure 9 This is a diagram showing the module relationships of the equipment. Detailed Implementation
[0067] The present application will be further described in detail below with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are only for explaining the present application and are not intended to limit the present application.
[0068] In the following description, numerous specific details are set forth for illustrative purposes in order to provide a thorough understanding of the inventive concept. As part of this specification, some of the accompanying drawings of this disclosure are block diagrams illustrating structures and devices to avoid complicating the disclosed principles. For clarity, not all features of the actual embodiment need to be described. References to “an embodiment” or “an embodiment” in this disclosure mean that a particular feature, structure, or characteristic described in connection with that embodiment is included in at least one embodiment, and multiple references to “an embodiment” or “an embodiment” should not be construed as necessarily referring to the same embodiment.
[0069] Unless explicitly defined, the terms “a,” “an,” and “the” are not intended to refer to a singular entity, but rather to include a general category whose specific examples can be used for illustration. Therefore, the use of the terms “a” or “an” can mean any number of at least one, including “a,” “one or more,” “at least one,” and “one or more.” The term “or” means any of the options and any combination of the options, including all options unless explicitly indicated that the options are mutually exclusive. The phrase “at least one of” when combined with a list of items refers to a single item in the list or any combination of items in the list. The phrase does not require all items listed unless explicitly defined as such.
[0070] First implementation method:
[0071] The first embodiment of the present invention provides a table recognition method based on a multimodal large model, comprising the following steps:
[0072] Step 1: Extract the cell coordinate matrix of the table and generate a coordinate feature map.
[0073] Specifically, when extracting the cell coordinate matrix of a table, a lightweight table detection model (such as PP-StructureV2) can be used to perform structured analysis on the original table image, extract the cell coordinate matrix information, and generate the corresponding coordinate feature map accordingly.
[0074] like Figure 2 As shown, during the extraction process, the areas of the table that need to be extracted must first be identified, such as annotations, table headers, and table text with different colored boxes. Then, the coordinate positions of the corresponding areas are extracted, and finally, a coordinate feature map is formed.
[0075] Step 2: Establish a multimodal large model, and combine the table image and coordinate feature map to perform image encoding feature splicing to obtain the image feature vector.
[0076] Specifically, this step first requires establishing a multimodal large model. In the process of establishing the multimodal large model, a large amount of historical data is needed to train the neural network model so that the final multimodal large model can meet the usage requirements and ensure the accuracy of table recognition.
[0077] S21: The process of establishing a multimodal large model is as follows:
[0078] S211: Collect HTML structural tags and text tags;
[0079] Among them, HTML structure tags are the HTML structured descriptions corresponding to the tables in the table image, including...
[0080] 、 、 Tags such as colspan and rowspan.
[0081] S212: Establish a token segmentation reference library based on HTML structural tags;
[0082] Specifically: the token segmentation reference library can be automatically parsed by the model based on historical data, or it can be added manually. In this implementation, considering the special characteristics of cells spanning multiple rows and columns, to ensure the semantic integrity of structural tokens during the model input stage, it is necessary to avoid their segmentation during the token recognition process. Therefore, the collected HTML structural tags (such as...) 、 、 、 <td rowspan="i">The HTML structure tags are manually added to the token segmentation reference library in complete form. In this way, the above-mentioned pre-set HTML structure tags are identified as independent token segmentation that cannot be split in the preprocessing stage, thereby ensuring the consistency of the structure and semantics in the subsequent encoding, fusion and decoding process. The hierarchical level and boundary of the structure tags can be accurately identified and restored in the training and inference stage, thereby improving the accuracy of multi-modal large model identification.
[0083] S213: inputting the table image, the coordinate feature map of the table, the HTML structure tag, the text label and the token segmentation reference library into a neural network model for training;
[0084] S214: calculating a loss function of the neural network model;
[0085]
[0086] wherein L represents a loss value, V represents a size of a word table, w i represents a training weight factor corresponding to the position token, y i represents a label of the token segmentation at the i-th position, p i represents a probability value corresponding to the prediction result at the i-th position;
[0087] In actual use, the weight factor is initialized and set according to the category to which the token segmentation belongs. For token segmentation belonging to the HTML structure tag label (for example 、 、 <td colspan="i">、 <td rowspan="i">For structural tags, their initial weight is set to 2.0; for ordinary text tags, their initial weight is set to 1.0. The weight factor ranges from [0, +∞), and the weight value is set based on task requirements and training objectives. The above weight value division is one method in this embodiment. In actual division, it can be divided as needed. The key point is that the weights of the token segments corresponding to text tags and the token segments corresponding to structural tags should not be inconsistent.
[0088] S215: Combine the loss function of the neural network model to evaluate the trained neural network model. If the loss function calculation result meets the set threshold range, the neural network model is saved as a multimodal large model. If the loss function calculation result does not meet the set threshold range, the neural network model is trained again until the loss function calculation result meets the set threshold range, and then the multimodal large model at this time is output.
[0089] Specifically: During model training, the current model is evaluated at fixed intervals, using the loss function value calculated in S142 to determine the trend of model performance changes. If multiple consecutive evaluations show that the loss function value tends to stabilize or no longer decreases significantly, it indicates that the model's learning effect has basically reached its optimal state, and the model training is considered to have converged. The output at this point is the final multimodal large model. Figure 3 As shown, after a large number of training iterations, the overall model tends to converge, and the loss value of the loss function is within the set threshold range of 0.6. This is the final multimodal large model.
[0090] S22: Combine the table image and coordinate feature map to perform image encoding feature concatenation to obtain the image feature vector.
[0091] Specifically:
[0092] S221: Extract visual semantic information from the table image and generate a semantic representation vector. The visual semantic information includes text data, line data, and background data.
[0093] S222: Extract relevant features from the coordinate feature map. These relevant features include the spatial layout information of each cell in the table area in the two-dimensional matrix, the coordinate position of each cell in the image, and the row and column span of each cell.
[0094] S223: Encode the coordinate feature map and extract the structure representation vector, which contains the table's layout and hierarchical structure;
[0095] S224: Concatenate the semantic representation vector and the structural representation vector dimensionally to form an image feature vector.
[0096] Step 3: Based on the multimodal large model, combine the cell coordinate matrix with the set prompt words to concatenate the text encoding features to obtain the text feature vector.
[0097] Specifically:
[0098] S31: Encode the cell coordinate matrix into structured text information;
[0099] Specifically, the structured text information of the cell matrix can be in the form of a JSON array. That is, each cell contains two sets of coordinates: the x-coordinate of the top-left corner, the y-coordinate of the top-left corner, the x-coordinate of the bottom-right corner, and the y-coordinate of the bottom-right corner. The two sets of coordinates are combined to form the corresponding cell coordinate matrix. For example: [[0,0,100,30],[100,0,200,30],...].
[0100] S32: Concatenate structured text information with the set prompt words to generate enhanced feature words;
[0101] S33: Encode the enhanced feature words to obtain coordinate features. Figure 1 The text feature vectors are consistent.
[0102] For example, the prompt could be: "Based on the original table image and its coordinate feature map given above, and combined with the following cell coordinate matrix, please fully parse the table's row and column structure, merged cells, and text content to generate an HTML5-compliant output."
[0103]
[0104]
[0105]
[0106]
[0107]
[0108]
[0109]
[0110]
[0111]
[0112]
[0113]
[0114]
[0115]
[0116]
[0117]
[0118]
[0119]
[0120]
[0121]
[0122]
[0123] Code. Requirements: 1. Keep 、 The nesting relationship is correct. Specifically, (table row) must directly contain (Table cell), and each Must be fully nested in Intracellularly, a parent-child hierarchical relationship is formed; 2, using row span, column span for merged cells; 3, output code format is beautiful, indentation is standard. After the prompt word, the cell coordinate matrix extracted in S1 (for example, in the form of a JSON array) is appended as structured input content. The enhanced feature words formed at this time contain the accurate coordinate matrix corresponding to each cell. In the subsequent recognition process, the cell will be recognized as a whole and will not be separated, ensuring the accuracy of the recognition. Step 4, fuse the text feature vector and the image feature vector to generate a unified semantic vector. Specifically: S41: construct a unified multi-modal representation space; S42: in the multi-modal representation space, cluster the text feature vector and the image feature vector with the same semantic content; S43: adjust the dimensions of the text feature vector and the image feature vector to the same size through a fully connected layer, and combine the adjusted text feature vector and image feature vector in structure, semantics, and task target to generate a unified semantic vector. Step 5, generate a decoding sequence according to the unified semantic vector to obtain the HTML text content of the table. Specifically, first, the semantic vector is analyzed through a multi-head attention mechanism to extract the distributed representation of the table structure and the text content; then, the decoder generates the output sequence step by step in a self-recurrent manner, and at each time step, the probability distribution of the next Token is calculated based on the current hidden state and the generated content; finally, the hidden state is mapped to the vocabulary space through a linear transformation layer, and the Token with the highest probability is selected as the output, and the loop is iterated until the complete Token sequence is generated. S51: According to the unified semantic vector, generate several token words according to the time step. S52: After generating each token word, judge the label type, which includes HTML structure closing label, HTML structure starting label and text content label. S53: According to the corresponding label type, dynamically select the next token word and output the corresponding token word. Specifically: S531: make a preliminary classification of the current token word, if the current token word is a text content label, select the language model to identify the full vocabulary space, and output the next token word. S532: If the current token word is an HTML structure label, make a second classification, if the current token word is an HTML structure starting label, shield the HTML structure label, and select the next token word from the text label; if the current token word is an HTML structure closing label, shield the non-HTML structure label, and select the next token word from the text label. S533: Form a token sequence according to multiple tokens.Specifically, in each generation step, the system first performs type judgment on the current generated Token, and dynamically switches the selectable set of the next Token according to whether it is an HTML structural tag or text content. Specifically, the following three control strategies are included: 1. If the current Token is an HTML structural start tag (for example. 、 、
[0124] When the system determines that the model has entered a table content segment, the next token should be natural language text content. At this time, the structural constraint module is activated, and the output probability of all HTML structural tokens is set to zero through a masking strategy, retaining only text tokens as candidates to ensure that the structural tag is immediately followed by valid text content. 2. If the current token is an HTML structural closing tag (e.g., <html>, <body ... When the system determines that it is currently at a structure block switching node, the next token should be a new structure tag. At this time, it reverse-blocks the output probability of all non-structure text tokens, retaining only the candidates of valid structure tokens to ensure the correct progression of the tag hierarchy logic.
[0125] 3. If the current Token is ordinary text content (i.e., it does not belong to the set of structural tags), the structural constraint module is not activated. The system performs prediction and sampling in the entire vocabulary space according to the method of the general language model, without interfering with the distribution of Tokens.
[0126] S54: Segment the multiple output tokens to form a decoding sequence and obtain the HTML text content of the table.
[0127] After completing multimodal feature fusion and structured token-level decoding, the final output is complete HTML text content. The HTML code is generated token by token by the decoding module, with a standardized structure hierarchy and complete tag closures, which can accurately restore the row and column relationships, merged cells, and nested structure information in the original table image.
[0128] The generated HTML code will be in standard HTML5 format.
[0129] Tags are the core, including 、 、 The structural units are combined with attributes such as row span and column span to accurately express the layout logic of the cell crossing rows or columns. Meanwhile, the text content remains consistent with the original Figure 1 , and good visual consistency and semantic correspondence are achieved.
[0130] The above HTML code can be directly used for DOM rendering engine in the front-end browser environment to visually display the table, and can also be used to provide a structured data output interface to the downstream system to realize functions such as data extraction, format conversion, structure comparison, intelligent form filling, etc.
[0131] The step division of the above various methods is only for clear description, and in implementation, one step can be combined or some steps can be split and decomposed into multiple steps, as long as the same logical relationship is included, all within the protection scope of the patent, and adding irrelevant modifications or introducing irrelevant designs in the algorithm or process, but not changing the core design of the algorithm and process, are within the protection scope of the patent.
[0132] Second embodiment:
[0133] The second embodiment of the present application provides a table recognition system based on a multi-modal large model, as shown in the following formula (1): Figure 7 , comprising:
[0134] The acquisition module 201 is configured to acquire a table image.
[0135] The extraction module 202 is configured to extract a cell coordinate matrix in the table image to generate a coordinate feature map.
[0136] The modeling module 203 is configured to establish a multi-modal large model.
[0137] The image splicing module 204 is configured to perform image encoding feature splicing on the table image and the coordinate feature map according to the multi-modal large model to obtain an image feature vector.
[0138] The text splicing module 205 is configured to perform text encoding feature splicing on the cell coordinate matrix and the set prompt word according to the multi-modal large model to obtain a text feature vector.
[0139] The fusion module 206 is configured to fuse the text feature vector and the image feature vector to generate a unified semantic vector.
[0140] The decoding module 207 is configured to generate a decoding sequence according to the unified semantic vector to obtain an HTML text content of the table.
[0141] Third embodiment:
[0142] A third embodiment of the present invention provides a computer-readable storage medium storing a set of program instructions that can be read by a machine. When the set of program instructions stored in the medium is read and executed by a machine, the machine can execute the table recognition method described in the foregoing embodiments of this application.
[0143] Fourth implementation method:
[0144] like Figure 6 As shown, a third embodiment of the present invention provides an apparatus including: at least one processor 401; and a memory 402 communicatively connected to the at least one processor 401; wherein the memory 402 stores instructions executable by the at least one processor 401, the instructions being executed by the at least one processor 401 to enable the at least one processor 401 to perform the table recognition method described in the foregoing embodiments of this application.
[0145] The memory 402 and processor 401 are connected via a bus, which may include any number of interconnecting buses and bridges. The bus connects various circuits of one or more processors 401 and memory 402 together. The bus can also connect various other circuits, such as peripheral devices, voltage regulators, and power management circuits, which are well known in the art and therefore will not be described further herein. A bus interface provides an interface between the bus and the transceiver. The transceiver can be a single element or multiple elements, such as multiple receivers and transmitters, providing a unit for communicating with various other devices over a transmission medium. Data processed by processor 401 is transmitted over a wireless medium via an antenna, which further receives data and transmits it to processor 401.
[0146] Processor 401 is responsible for managing the bus and general processing, and can also provide various functions, including timing, peripheral interfaces, voltage regulation, power management, and other control functions. Memory 402 can be used to store data used by processor 401 during operation.
[0147] The above descriptions are merely embodiments of the present invention. Commonly known structures and characteristics are not described in detail here. Those skilled in the art are aware of all common technical knowledge in the field prior to the application date or priority date, are aware of all existing technologies in that field, and have the ability to apply conventional experimental methods prior to that date. Those skilled in the art can, under the guidance of this application, improve and implement this solution in combination with their own capabilities. Some typical known structures or methods should not be obstacles for those skilled in the art to implement this application. It should be noted that those skilled in the art can make several modifications and improvements without departing from the structure of the present invention. These should also be considered within the scope of protection of the present invention, and will not affect the effectiveness of the implementation of the present invention or the practicality of the patent. The scope of protection claimed in this application should be determined by the content of its claims, and the specific embodiments described in the specification can be used to interpret the content of the claims.
Claims
1. A table recognition method based on a multimodal large model, characterized in that, Includes the following steps: Extract the cell coordinate matrix from the table and generate a coordinate feature map; A multimodal large model is established, and image encoding features are concatenated by combining table images and coordinate feature maps to obtain image feature vectors; Based on the multimodal large model, the text feature vector is obtained by concatenating the cell coordinate matrix with the set prompt words and text encoding features. By fusing text feature vectors and image feature vectors, a unified semantic vector is generated; Based on a unified semantic vector, a decoding sequence is generated to obtain the HTML text content of the table.
2. The table recognition method according to claim 1, characterized in that, Establish a large multimodal model, specifically as follows: Collect HTML structural tags and text tags; Establish a token segmentation reference library based on HTML structural tags; The table image, the table's coordinate feature map, HTML structural tags, text tags, and the token segmentation reference library are input into the neural network model for training. Calculate the loss function of the neural network model; Where: L represents the loss value, V represents the vocabulary size, and y i p represents the tag of the token segment at position i. i w represents the probability value corresponding to the prediction result of the i-th position. i This is the training weight factor corresponding to the Token at this position; By combining the loss function of the neural network model, the trained neural network model is evaluated. If the loss function calculation result meets the set threshold range, the neural network model is saved as a multimodal large model. If the loss function calculation result does not meet the set threshold range, the neural network model is trained again until the loss function calculation result meets the set threshold range, and then the multimodal large model at this time is output.
3. The table recognition method according to claim 2, characterized in that, By combining the table image and the coordinate feature map, image encoding features are concatenated to obtain the image feature vector, specifically: Extract visual semantic information from the table image and generate a semantic representation vector. The visual semantic information includes text data, line data, and background data. Extract relevant features from the coordinate feature map, including the spatial layout information of each cell in the table area in the two-dimensional matrix, the coordinate position of each cell in the image, and the row and column span of each cell; The coordinate feature map is encoded, and a structural representation vector is extracted. The structural representation vector contains the table's layout relationship and hierarchical structure. The semantic representation vector and the structural representation vector are concatenated dimensionally and fused to form the image feature vector.
4. The table recognition method according to claim 3, characterized in that, By combining the cell coordinate matrix with the set prompt words, text encoding features are concatenated to obtain the text feature vector, specifically: Encode the cell coordinate matrix into structured text information; Structured text information is concatenated with the set prompt words to generate enhanced feature words; The enhanced feature words are encoded to obtain text feature vectors that are consistent with the coordinate feature map.
5. The table recognition method according to claim 4, characterized in that, The text feature vector and image feature vector are fused to generate a unified semantic vector, specifically as follows: Construct a unified multimodal representation space; Within the multimodal representation space, text feature vectors and image feature vectors with the same semantic content are clustered. The fully connected layer adjusts the dimensions of the text feature vector and the image feature vector to the same size, and then concatenates and combines the adjusted text feature vector and image feature vector with respect to structure, semantics and task objectives to generate a unified semantic vector.
6. The table recognition method according to claim 5, characterized in that, Based on a unified semantic vector, a decoding sequence is generated to obtain the HTML text content of the table, specifically: Based on a unified semantic vector, several token words are generated step by step according to time steps; After generating each token segment, the tag type is determined. The tag type includes HTML structure closing tags, HTML structure starting tags, and text content tags. Based on the corresponding tag type, dynamically select the category of the next token segmentation and output the corresponding token segmentation; The output tokens are segmented to form a decoding sequence, which yields the HTML text content of the table.
7. The table recognition method according to claim 6, characterized in that, Based on the corresponding tag type, dynamically select the next token segmentation category and output the corresponding token segmentation, specifically: Perform an initial category determination on the current token segment. If the current token segment is a text content label, select a language model to identify the entire word space and output the next token segment. If the current token segmentation is an HTML structure tag, then perform another category judgment. If the current token segmentation is an HTML structure start tag, then HTML structure tags are blocked, and the next token segmentation is selected from text tags. If the current token segmentation is an HTML structure closing tag, then non-HTML structure tags are blocked, and the next token segmentation is selected from text tags. A token sequence is formed from multiple tokens.
8. A table recognition system based on a multimodal large model, characterized in that, include: The data acquisition module is used to acquire table images; The extraction module is used to extract the cell coordinate matrix from the table image and generate a coordinate feature map; The modeling module is used to build large multimodal models. The image stitching module is used to stitch together image encoding features based on a multimodal large model, combined with a table image and coordinate feature map, to obtain an image feature vector. The text concatenation module is used to concatenate text encoding features based on the multimodal large model, combined with the cell coordinate matrix and the set prompt words, to obtain the text feature vector; The fusion module is used to fuse text feature vectors and image feature vectors to generate a unified semantic vector; The decoding module is used to generate a decoding sequence based on a unified semantic vector to obtain the HTML text content of the table.
9. A medium, characterized in that: The medium stores a set of program instructions that can be read by a machine. When the set of program instructions stored in the medium is read and executed by the machine, the machine can implement the table recognition method according to any one of claims 1 to 7.
10. A device, characterized in that, The device includes a connected processor and a memory; the memory stores a set of program instructions; when the program instructions stored in the memory are read and executed by the processor, the device can implement the table recognition method according to any one of claims 1 to 7.
Citation Information
Cited By
Multi-source complex table-oriented trusted question and answer agent construction method and system
CN121766460A
Multi-source complex table-oriented trusted question and answer agent construction method and system
CN121766460B
Structured processing method and device for medical instrument data, medium and equipment
CN122045231A