Table processing method and device, storage medium and electronic equipment
By generating and verifying table processing results in multimodal large models, the problem of lack of accuracy in table data processing of multimodal large models is solved, and higher accuracy and interpretability are achieved.
Patent Information
- Application Number
- CN202510797482.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-16
- Publication Date
- 2025-07-11
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The prior art lacks effective results verification methods when processing table data using multimodal large models, resulting in a lack of accuracy in output results, especially when facing complex structures or irregular formats.
By obtaining the pending table image and problem text, inputting a pre-trained multimodal large model, generating the first processing result, and using the code generation template to generate executable code, executing the code with the code executor, obtaining the second processing result, and finally verifying it based on the first and second processing results to determine the target processing result.
It improves the accuracy of the multimodal large model's table processing results, and makes its inference process more interpretable, which improves users' trust in the output results.
Smart Images

Figure CN120297429A_ABST
Abstract
Description
Technical Field
[0001] This specification relates to the field of artificial intelligence, and particularly to a method, device, storage medium, and electronic device for table processing. Background Art
[0002] Currently, in the era of big data, users usually face the situation of needing to process a large amount of tabular data. To improve the processing efficiency of tabular data, manual rules or simple structured parsing algorithms such as regular expressions, template matching, and simple layout analysis are usually adopted. However, these methods are prone to failure when dealing with tables with complex structures or irregular formats. Especially in the process of cross-platform and cross-document format processing, the compatibility of different devices and file formats also brings parsing difficulties.
[0003] In the prior art, intelligent algorithms such as machine learning and deep learning have been gradually applied to table processing. Especially the development of multi-modal large model algorithms has made it possible to extract image features from pictures and map them to the language space. Through the large language model for understanding and analysis in the language dimension, the accuracy and generalization performance of table understanding can be improved.
[0004] However, currently in the process of using multi-modal large models to process tabular data, there is a lack of effective methods to verify the output results of the models, which makes the output results of these models often lack sufficient accuracy. Based on this, this application provides a method, device, storage medium, and electronic device for table processing. Summary of the Invention
[0005] This specification provides a method, device, storage medium, and electronic device for table processing to partially solve the above problems existing in the prior art.
[0006] This specification adopts the following technical solutions: A method for table processing includes: Obtaining a table image to be processed and a problem text, where the problem text represents a processing task for the table in the table image; Inputting the table image and the problem text into a pre-trained multi-modal large model to make the multi-modal large model output a first processing result; According to a code generation template, making the multi-modal large model generate executable code corresponding to the processing task according to the text content in the table image, and using a code executor to execute the code to obtain a second processing result; Obtaining a target processing result of the table image to be processed according to the first processing result and the second processing result.
[0007] Optionally, obtaining the problem text specifically includes: Obtain the input text of the user; According to the input text, select a target prompt word from the preset prompt words; Concatenate the target prompt word and the input text to obtain a question text.
[0008] Optionally, according to the code generation template, enable the multi-modal large model to generate the executable code corresponding to the processing task according to the text content in the table image, specifically including: Through the multi-modal large model, determine the text data of the table image to be processed; According to the question text and the text data, generate a code generation instruction according to the code generation template; Input the code generation instruction into the multi-modal large model to obtain the executable code corresponding to the processing task.
[0009] Optionally, according to the first processing result and the second processing result, obtain the target processing result of the table image to be processed, specifically including: Calculate the first similarity between the first processing result and the second processing result, and determine whether the first similarity exceeds a threshold; If so, use the first processing result as the target processing result; If not, repeat the step of determining the second processing result to determine multiple second processing results, calculate the second similarities between the multiple second processing results, and determine the target processing result with the most occurrences from the multiple second processing results, where two second processing results with a second similarity exceeding the threshold represent two identical second processing results.
[0010] Optionally, according to the first processing result and the second processing result, obtain the target processing result of the table image to be processed, specifically including: Repeat the step of determining the second processing result and / or the step of determining the first processing result to determine each processing result to be verified, where each processing result to be verified includes at least one second processing result and at least one first processing result; Calculate the third similarity between each processing result to be verified, and determine the target processing result with the most occurrences from each processing result to be verified according to the third similarity, where two processing results to be verified with a third similarity exceeding the threshold represent two identical processing results to be verified.
[0011] This specification provides a table processing device, including: An acquisition module, configured to acquire a table image to be processed and a question text, where the question text represents a processing task for the table in the table image; A first processing module, configured to input the table image and the question text into a pre-trained multimodal large model, so that the multimodal large model outputs a first processing result; A second processing module, configured to generate a template according to the code, so that the multimodal large model generates executable code corresponding to the processing task according to the text content in the table image, and execute the code by using a code executor to obtain a second processing result; A determination module, configured to obtain a target processing result of the table image to be processed according to the first processing result and the second processing result.
[0012] Optionally, the determination module is specifically configured to calculate a first similarity between the first processing result and the second processing result, and determine whether the first similarity exceeds a threshold; if so, use the first processing result as the target processing result; if not, repeat the step of determining the second processing result to determine multiple second processing results, calculate each second similarity between the multiple second processing results, and determine a target processing result whose occurrence times reach a preset value from the multiple second processing results, where two second processing results with a second similarity exceeding the threshold indicate two identical second processing results.
[0013] Optionally, the determination module is specifically configured to repeat the step of determining the second processing result and / or the step of determining the first processing result to determine each processing result to be verified, where each processing result to be verified includes at least one second processing result and at least one first processing result; calculate a third similarity between each processing result to be verified, and determine a target processing result with the most occurrences from each processing result to be verified, where two processing results to be verified with a third similarity exceeding the threshold indicate two identical processing results to be verified.
[0014] This specification provides a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the above table processing method is implemented.
[0015] This specification provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the program, the above table processing method is implemented.
[0016] At least one of the technical solutions adopted in this specification can achieve the following beneficial effects: In the table processing method provided in this specification, a table image to be processed and a problem text are obtained. The problem text represents a processing task for the table in the table image. The table image and the problem text are input into a pre-trained multi-modal large model, so that the multi-modal large model outputs a first processing result. According to a code generation template, the multi-modal large model generates executable code corresponding to the processing task based on the text content in the table image, and uses a code executor to execute the code to obtain a second processing result. According to the first processing result and the second processing result, the target processing result of the table image to be processed is obtained.
[0017] In the above method, table processing is implemented based on a multi-modal large model, and at the same time, a code executor is combined to verify the processing result, improving the accuracy of the multi-modal large model for table processing results. At the same time, based on this self-verification method, the inference process of the multi-modal large model is more interpretable. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] The drawings described herein are used to provide a further understanding of this specification and form a part of this specification. The illustrative embodiments of this specification and their descriptions are used to explain this specification and do not constitute an improper limitation of this specification. In the drawings: Figure 1 is a schematic flowchart of a table processing method provided in this specification; Figure 2 is a schematic diagram of generating a first processing result provided in this specification; Figure 3 is a schematic diagram of generating a second processing result provided in this specification; Figure 4 is a schematic diagram of a table processing device provided in this specification; Figure 5 corresponding to that provided in this specification Figure 1 is a schematic structural diagram of an electronic device. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0019] To make the purpose, technical solutions, and advantages of this specification clearer, the technical solutions of this specification will be clearly and completely described below in conjunction with the specific embodiments of this specification and the corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of this specification, rather than all the embodiments. Based on the embodiments in this specification, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of this application.
[0020] It should be understood that although the terms first, second, third, etc. may be used in this application to describe various information, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from each other. For example, without departing from the scope of this application, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the word "if" as used herein may be interpreted as "when" or "while" or "in response to a determination".
[0021] With the development of computer technology, tables have been widely used in many fields due to their ability to contain structured information and support quick viewing and analysis. However, with the continuous increase in the amount of information, the structure and content of tables have become increasingly complex, including multiple rows and columns, merged cells, nested relationships, and elements such as text, pictures, and symbols with different formats. These complex features make the automatic parsing and understanding of tables extremely difficult, and thus pose higher requirements for the effective processing of table data.
[0022] Most current table processing methods rely on simple structured parsing algorithms such as regular expressions, template matching, and simple layout analysis, or other manual rules. However, these methods have low accuracy when dealing with tables with complex structures or irregular formats. Especially in the process of cross-platform and cross-document format processing, the compatibility of different devices and file formats also brings processing difficulties. In addition, tasks such as information extraction, semantic association, and row-column relationship parsing in table understanding require extremely high precision and stability of the table understanding system, and simple structured parsing is difficult to meet the actual needs.
[0023] With the development of large models, multi-modal models for table understanding and analysis are expected to further improve the accuracy and generalization performance of table processing. However, currently, in the process of using multi-modal large models to process table data, there is a lack of effective methods to verify the output results of the models, which makes the output results of these models often lack sufficient accuracy. Based on this, this application provides a table processing method, device, storage medium, and electronic device.
[0024] The following will, with reference to the accompanying drawings, elaborate on the technical solutions provided by each embodiment of this specification.
[0025] Figure 1 The following is a schematic flowchart of a table processing method provided by an embodiment of this specification, including the following steps: S100: Obtain a table image to be processed and a question text, where the question text represents a processing task for the table in the table image.
[0026] In one or more embodiments of this specification, there is no limitation on which device specifically executes this table processing method. For example, it can be a mobile terminal, a server, etc. However, since subsequent steps involve model pre-training, prompt word splicing, etc., these steps are generally executed by the server. Therefore, in the following of this specification, the server executing this geometric structure evaluation method is taken as an example for description. Among them, the server can be a single device or composed of multiple devices. For example, a distributed server, a server for cloud services, etc. This specification does not limit this.
[0027] Since the original intention of the multi-modal large model is to be able to understand and generate data across multiple modalities, such as text, images, etc., and generally does not directly process the original file, in order to complete the processing of the table through the pre-trained large model, the server can first obtain the table image to be processed and the question text representing the processing task of the table. Among them, the processing tasks include but are not limited to table structure recognition, table classification, table fine understanding, table editing, general table description, table recognition, table numerical question answering, table analysis code generation, table structure layout question answering, table editing, cross-table question answering, etc. Since there are many types of table processing tasks, they are not listed one by one here.
[0028] However, it should be noted that in one or more embodiments of this specification, there is no limitation on how the server specifically obtains the table to be processed and the question text. It can directly obtain the table image to be processed input by the user from the user's input, or it can obtain the table uploaded by the user and then take a screenshot of the table to obtain the table image to be processed. Of course, if the table input by the user is text data, the text data can also be directly obtained without taking a screenshot of the table, instead of obtaining the table image to be processed.
[0029] Furthermore, in one or more embodiments of this specification, there is also no limitation on the specific manner in which the server obtains the question text representing the processing task of the table in the table image. It can use the text data input by the user as the question text. For example, if the text data input by the user is text data with a clear processing task such as "Please analyze what the maximum value of the price in the table is", the user's input can be directly used as the question text.
[0030] In order for the multi-modal large model to better execute the processing task, the server can also splice the user's input text with a preset system prompt word to generate the question text. For example, the system prompt word is "You are an expert in table recognition. Please answer according to the table picture provided by the user and the following question: ( )", and the user's input text is spliced into the "( )" of the system prompt word to generate the question text.
[0031] Since there are many types of processing tasks for tables, the server can also generate a corresponding prompt word for each processing task. Then, after obtaining the user's input text, according to the semantics of the text data, it matches the target prompt word among the preset prompt words. Then, it concatenates the target prompt word and the input text to obtain the question text.
[0032] S102: Input the table image and the question text into a pre-trained multi-modal large model, so that the multi-modal large model outputs a first processing result.
[0033] After obtaining the table image and the question text to be processed, the server can input them into a pre-trained multi-modal large model, so that the multi-modal large model performs processing tasks according to the question text and the table image and outputs a first processing result.
[0034] As Figure 2 shown, Figure 2 FIG. is a schematic diagram of generating a first processing result provided in this specification. After the server inputs the question text and the table image into the multi-modal large model, the multi-modal large model generates a reply to the question text based on the table image, that is, the first processing result.
[0035] S104: According to the code generation template, make the multi-modal large model generate the executable code corresponding to the processing task according to the text content in the table image, and use the code executor to execute the executable code to obtain a second processing result.
[0036] Since there may be problems of "large model hallucinations" in the multi-modal large model, which results in a low trust level of users in the first processing result output by the multi-modal large model. In order to improve the trust level of users in the first processing result output by the multi-modal large model and the accuracy of the first processing result output by the multi-modal large model, the server can introduce a verification mechanism.
[0037] Currently, the technology of generating executable code through a multi-modal large model is becoming increasingly mature. The server can obtain a preset code generation template, and then generate the executable code corresponding to the processing task according to the text content in the table image through the multi-modal large model, and use the code executor to execute the code to obtain a second processing result.
[0038] Specifically, the server can first convert the table image to be processed into text data, and then, according to the question text and the converted text data, generate a code generation instruction according to the code generation template. Then, input the code generation instruction into the multi-modal large model to obtain the executable code corresponding to the processing task.
[0039] As Figure 3 shown, Figure 3A schematic diagram for generating a second processing result provided in this specification. The server inputs "You are an expert in table recognition. Please recognize the table image and output it in text form" and the table image into the multimodal large model. Then, the text data output by the large model is spliced with the code generation template to generate a code generation instruction. The code generation instruction is input into the multimodal large model to enable the multimodal large model to generate executable code. Then, the executable code is executed using the code executor to obtain the second processing result.
[0040] It should be noted that when generating the code generation instruction, the code generation template can be "You are an expert in table recognition. Please generate the corresponding Python code according to the table data file <text data> and the question <question text>". Here, <text data> and <question text> represent placeholders, which are replaced by specific text data and question text during execution. Of course, the above is only one embodiment provided in this specification, and the code generation instruction can also be generated by other methods, which are not listed one by one in this specification.
[0041] Furthermore, due to the length limitation of the input of the multimodal large model, the text data can be saved to a csv file for calling Python code. The above code is executed through the Python interpreter, and the second processing result is determined according to the return result of the program execution.
[0042] S106: Obtain the target processing result of the to-be-processed table image according to the first processing result and the second processing result.
[0043] After determining the first processing result and the second processing result respectively according to the above two methods, the server can then determine the target processing result according to the first processing result and the second processing result, so as to verify the output result of the multimodal large model.
[0044] Specifically, after determining the first processing result and the second processing result, the server can first determine the first similarity between the first processing result and the second processing result, and judge whether the first similarity exceeds the threshold. If so, it means that the output result of the multimodal large model is accurate, and the server can use the first processing result as the target processing result. If not, it means that the output result of the multimodal large model is inaccurate and needs further verification. Then the server can repeat the step of determining the second processing result, determine multiple second processing results, and then calculate the second similarities between the multiple second processing results. According to the second similarities, the target processing result with the most occurrences is determined from the multiple second processing results. Among them, two second processing results with a second similarity exceeding the threshold indicate two identical second processing results.
[0045] It should be noted that in one or more embodiments of this specification, there is no limitation on the specific method adopted by the server to determine the first similarity and the second similarity. What can be calculated are semantic similarity, cosine similarity, vector similarity, etc. It is also possible to weight each similarity according to a certain weight to determine the total similarity. This specification does not limit this and can be set according to actual needs. Moreover, the threshold can be set according to actual needs, and this specification does not limit it.
[0046] Based on Figure 1 In the table processing method shown, by obtaining the table image to be processed and the question text, where the question text represents the processing task for the table in the table image, inputting the table image and the question text into a pre-trained multimodal large model, enabling the multimodal large model to output a first processing result, according to the code generation template, enabling the multimodal large model to generate executable code corresponding to the processing task based on the text content in the table image, and using the code executor to execute the code to obtain a second processing result, and obtaining the target processing result of the table image to be processed according to the first processing result and the second processing result.
[0047] In the above method, table processing is implemented based on a multimodal large model, and at the same time, a code executor is combined to verify the processing result, improving the accuracy of the table processing result of the multimodal large model and making the reasoning process of the multimodal large model more interpretable.
[0048] In addition, in step S100, in one or more embodiments of this specification, there is no limitation on the number of table images to be processed obtained, nor on the number of table images input into the multimodal large model, and they can be obtained or input according to the actual needs of the user.
[0049] Furthermore, in step S106, when verifying the output result of the multimodal large model, in order to further improve the accuracy of the target processing result, the server can also directly repeat the steps of determining the second processing result and / or the steps of determining the first processing result to determine each processing result to be verified, where each processing result to be verified includes at least one second processing result and at least one first processing result. That is, each processing result to be verified can include one first processing result and multiple second processing results, or each processing result to be verified includes multiple first processing results and one second processing result, or each processing result to be verified includes multiple first processing results and multiple second processing results.
[0050] Then, the server calculates the third similarity between each processing result to be verified, and determines the target processing result with the most occurrences from each processing result to be verified according to the third similarity, where two processing results to be verified with a third similarity exceeding the threshold represent two identical processing results to be verified.
[0051] It should be noted that in one or more embodiments of this specification, the calculation method of the third similarity is the same as that of the first similarity and the second similarity, and can be set according to actual needs, and this specification does not limit this.
[0052] Furthermore, in order to improve the verification efficiency, the server can also first determine the first similarity between the first processing result and the second processing result, and determine whether the first similarity exceeds a threshold. If so, it means that the output result of the multi-modal large model is accurate, and the server can use the first processing result as the target processing result. If not, it means that the output result of the multi-modal large model is inaccurate and further verification is required, and then directly repeat the steps of determining the second processing result and / or determining the first processing result to determine each processing result to be verified, where each processing result to be verified includes at least one second processing result and at least one first processing result. That is, each processing result to be verified can include one first processing result and multiple second processing results, or each processing result to be verified includes multiple first processing results and one second processing result, or each processing result to be verified includes multiple first processing results and multiple second processing results. Then, the server calculates the third similarity between each processing result to be verified, and determines the target processing result with the most occurrences from each processing result to be verified according to the third similarity, where two processing results to be verified with a third similarity exceeding the threshold represent two identical processing results to be verified.
[0053] In addition, when training the multi-modal large model, in order to improve the generality of the large model, the server can also obtain training samples for each type of processing task and use them to train the multi-modal large model. In order to further ensure that the multi-modal large model has the general ability of text and image understanding, in addition to tabular data, conventional text and image question-and-answer data, plain text data, code and mathematical data, etc. are also added.
[0054] Furthermore, in view of the problem that the logical structure in complex tables is difficult to be represented by a single structured data, the server can also split such tables into multiple parts for processing.
[0055] Furthermore, when training the processing ability of the multi-modal large model on cross-tabular data, when training the multi-modal large model, the data synthesis method can also be used to obtain a training data set, for example, splicing at least two tables.
[0056] In addition, in order to further improve the user's trust in the output result of the multi-modal large model, the server can also describe the above verification process of the first processing result and the second processing result in natural language form and display it to the user to further improve the interpretability of the inference process of the multi-modal large model.
[0057] Based on the same idea of a table processing method provided by one or more embodiments of this specification, this specification also provides a corresponding table processing device, as Figure 4 shown.
[0058] Figure 4 It is a schematic diagram of a table processing device provided by this specification, specifically including: An acquisition module 400, configured to acquire a table image to be processed and a question text, where the question text represents a processing task for the table in the table image; A first processing module 401, configured to input the table image and the question text into a pre-trained multimodal large model, so that the multimodal large model outputs a first processing result; A second processing module 402, configured to generate an executable code corresponding to the processing task according to a code generation template by the multimodal large model based on the text content in the table image, and execute the executable code by a code executor to obtain a second processing result; A determination module 403, configured to obtain a target processing result of the table image to be processed according to the first processing result and the second processing result.
[0059] Optionally, the acquisition module 400 is specifically configured to acquire an input text of a user; select a target prompt word from preset prompt words according to the input text; splice the target prompt word and the input text to obtain a question text.
[0060] Optionally, the second processing module 402 is specifically configured to determine text data of the table image to be processed through the multimodal large model; generate a code generation instruction according to the question text and the text data according to a code generation template; input the code generation instruction into the multimodal large model to obtain an executable code corresponding to the processing task.
[0061] Optionally, the determination module 403 is specifically configured to calculate a first similarity between the first processing result and the second processing result, and determine whether the first similarity exceeds a threshold; if so, use the first processing result as the target processing result; if not, repeat the step of determining the second processing result to determine multiple second processing results, calculate second similarities between the multiple second processing results, and determine a target processing result whose occurrence times reach a preset value from the multiple second processing results, where two second processing results with the second similarity exceeding the threshold represent two identical second processing results.
[0062] Optionally, the determining module 403 is specifically configured to repeatedly execute the step of determining the second processing result and / or the step of determining the first processing result to determine each processing result to be verified, where each processing result to be verified includes at least one second processing result and at least one first processing result; calculate a third similarity between each of the processing results to be verified, and determine, according to the third similarity, a target processing result with the most occurrences from each of the processing results to be verified, where two processing results to be verified with a third similarity exceeding a threshold represent two identical processing results to be verified.
[0063] This specification also provides a computer-readable storage medium storing a computer program that can be used to execute the above Figure 1 provided table processing method.
[0064] This specification also provides Figure 5 a corresponding Figure 1 structural schematic diagram of an electronic device as shown in Figure 5 the figure. As shown, at the hardware level, the electronic device includes a processor, an internal bus, a network interface, a memory, and a non-volatile memory. Of course, other hardware required for other services may also be included. The processor reads the corresponding computer program from the non-volatile memory into the memory and then runs it to implement the above Figure 1 described table processing method.
[0065] Of course, in addition to the software implementation, this specification does not exclude other implementation manners, such as a logic device or a combination of software and hardware. That is to say, the execution subject of the following processing flow is not limited to each logic unit, and may also be hardware or a logic device.
[0066] In the 1990s, it was obvious to distinguish whether an improvement to a technology was an improvement in hardware (e.g., improvement to the circuit structure of diodes, transistors, switches, etc.) or an improvement in software (improvement to the method flow). However, with the development of technology, many improvements to method flows today can be regarded as direct improvements to the hardware circuit structure. Almost all designers obtain the corresponding hardware circuit structure by programming the improved method flow into the hardware circuit. Therefore, it cannot be said that an improvement to a method flow cannot be implemented with a hardware entity module. For example, a Programmable Logic Device (PLD) (e.g., a Field Programmable Gate Array (FPGA)) is such an integrated circuit whose logical function is determined by the user programming the device. Designers can program by themselves to "integrate" a digital system on a PLD without having to ask a chip manufacturer to design and fabricate a dedicated integrated circuit chip. Moreover, nowadays, instead of manually fabricating integrated circuit chips, this programming is mostly implemented using "logic compiler" software, which is similar to the software compiler used in program development and writing. The original code before compilation also has to be written in a specific programming language, which is called a Hardware Description Language (HDL), and there is not only one kind of HDL, but many kinds, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, RHDL (Ruby Hardware Description Language), etc. The most commonly used ones currently are VHDL (Very-High-Speed Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art should also be aware that by simply performing a little logical programming on the method flow with the above-mentioned several hardware description languages and programming it into the integrated circuit, it is easy to obtain the hardware circuit that implements the logical method flow.
[0067] The controller can be implemented in any suitable manner. For example, the controller can take the form of, for example, a microprocessor or a processor and a computer-readable medium storing computer-readable program code (such as software or firmware) executable by the (micro)processor, logic gates, switches, an application specific integrated circuit (ASIC), a programmable logic controller, and an embedded microcontroller. Examples of the controller include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicone Labs C8051F320. The memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art also know that, in addition to implementing the controller in the form of pure computer-readable program code, it is entirely possible to logically program the method steps to enable the controller to be implemented in the form of logic gates, switches, application specific integrated circuits, programmable logic controllers, and embedded microcontrollers to achieve the same function. Therefore, such a controller can be considered a hardware component, and the devices included therein for implementing various functions can also be regarded as the structures within the hardware component. Or even, the devices for implementing various functions can be regarded as either software modules for implementing the method or structures within the hardware component.
[0068] The systems, devices, modules, or units illustrated in the above embodiments can be specifically implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, the computer can be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or a combination of any of these devices.
[0069] For the convenience of description, when describing the above devices, they are described separately as various units according to their functions. Of course, when implementing this specification, the functions of each unit can be implemented in the same or multiple software and / or hardware.
[0070] Those skilled in the art should understand that the embodiments of this specification can be provided as a method, a system, or a computer program product. Therefore, this specification can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, this specification can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0071] This specification is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the specification. It should be understood that each flow and / or block in the flowchart and / or block diagram, and combinations of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processors of general-purpose computers, special-purpose computers, embedded processors, or other programmable data processing devices to generate a machine, such that the instructions executed by the processors of the computer or other programmable data processing devices produce means for implementing the functions specified in one or more of the flows Figure 1 one or more of the flows and / or blocks Figure 1 or means for implementing the functions specified in one or more of the blocks.
[0072] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufacture including instruction means for implementing the functions specified in one or more of the flows Figure 1 one or more of the flows and / or blocks Figure 1 or means for implementing the functions specified in one or more of the blocks.
[0073] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more of the flows Figure 1 one or more of the flows and / or blocks Figure 1 or means for implementing the functions specified in one or more of the blocks.
[0074] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory.
[0075] The memory may include non-permanent memory in the form of computer-readable media, random access memory (RAM), and / or non-volatile memory such as read-only memory (ROM) or flash RAM. The memory is an example of computer-readable media.
[0076] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.
[0077] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.
[0078] It should be understood by those skilled in the art that the embodiments of this specification may be provided as methods, systems or computer program products. Therefore, this specification may take the form of a complete hardware embodiment, a complete software embodiment or an embodiment combining software and hardware. Moreover, this specification may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0079] This specification may be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. This specification may also be practiced in distributed computing environments where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules may be located in local and remote computer storage media, including storage devices.
[0080] Each embodiment in this specification is described in a progressive manner. For the identical or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the system embodiments, since they are basically similar to the method embodiments, the description is relatively simple. For the relevant parts, reference can be made to the corresponding descriptions in the method embodiments.
[0081] The above are only the embodiments of this specification and are not intended to limit this specification. For those skilled in the art, various modifications and changes can be made to this specification. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of this specification shall be included within the scope of the claims of this specification.
Claims
1. A table processing method, characterized in that, Including: Obtain a table image to be processed and a question text, where the question text represents a processing task for the table in the table image; Input the table image and the question text into a pre-trained multimodal large model, so that the multimodal large model outputs a first processing result; According to a code generation template, enable the multimodal large model to generate executable code corresponding to the processing task according to the text content in the table image, and use a code executor to execute the executable code to obtain a second processing result; According to the first processing result and the second processing result, obtain the target processing result of the table image to be processed.
2. The method according to claim 1, wherein Obtaining the question text specifically includes: Obtain the input text of the user; According to the input text, select a target prompt word from the preset prompt words; Concatenate the target prompt word and the input text to obtain a question text.
3. The method according to claim 1, wherein According to the code generation template, enabling the multimodal large model to generate executable code corresponding to the processing task according to the text content in the table image specifically includes: Determine the text data of the table image to be processed through the multimodal large model; According to the question text and the text data, generate a code generation instruction according to the code generation template; Input the code generation instruction into the multimodal large model to obtain the executable code corresponding to the processing task.
4. The method according to claim 1, wherein According to the first processing result and the second processing result, obtaining the target processing result of the table image to be processed specifically includes: Calculate a first similarity between the first processing result and the second processing result, and determine whether the first similarity exceeds a threshold; If so, use the first processing result as the target processing result; If not, repeat the step of determining the second processing result to determine multiple second processing results, calculate the second similarities between the multiple second processing results, and determine the target processing result with the most occurrences from the multiple second processing results, where two second processing results with a second similarity exceeding the threshold represent two identical second processing results.
5. The method according to claim 1, characterized in that, According to the first processing result and the second processing result, obtaining the target processing result of the table image to be processed specifically includes: Repeat the step of determining the second processing result and / or the step of determining the first processing result to determine each processing result to be verified, where each processing result to be verified includes at least one second processing result and at least one first processing result; Calculate a third similarity between each processing result to be verified, and determine the target processing result with the most occurrences from each processing result to be verified according to the third similarity, where two processing results to be verified with a third similarity exceeding the threshold represent two identical processing results to be verified.
6. A table processing device, characterized in that, Including: An acquisition module for acquiring a table image to be processed and a question text, where the question text represents a processing task for the table in the table image; A first processing module for inputting the table image and the question text into a pre-trained multimodal large model, so that the multimodal large model outputs a first processing result; A second processing module, configured to generate an executable code according to a code generation template, cause the multi-modal large model to generate the executable code corresponding to the processing task according to the text content in the table image, and execute the executable code by using a code executor to obtain a second processing result; A determination module, configured to obtain a target processing result of the table image to be processed according to the first processing result and the second processing result.
7. The device according to claim 6, characterized in that, Specifically, the determination module is configured to calculate a first similarity between the first processing result and the second processing result, and determine whether the first similarity exceeds a threshold; if so, use the first processing result as the target processing result; If not, repeat the step of determining the second processing result to determine a plurality of second processing results, calculate each second similarity between the plurality of second processing results, and determine, according to each second similarity, a target processing result whose occurrence times reach a preset value from the plurality of second processing results, where two second processing results with the second similarity exceeding the threshold indicate two identical second processing results.
8. The device according to claim 6, characterized in that, Specifically, the determination module is configured to repeat the step of determining the second processing result and / or the step of determining the first processing result to determine each processing result to be verified, where each processing result to be verified includes at least one second processing result and at least one first processing result; calculate a third similarity between each processing result to be verified, and determine, according to the third similarity, a target processing result with the most occurrences from each processing result to be verified, where two processing results to be verified with the third similarity exceeding the threshold indicate two identical processing results to be verified.
9. A computer-readable storage medium, characterized in that, The storage medium stores a computer program, and when the computer program is executed by a processor, the method described in any one of claims 1 to 5 above is implemented.
10. An electronic device, characterized in that, It includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the method described in any one of claims 1 to 5 above is implemented.
Citation Information
Patent Citations
Table data processing large language model training method and device, medium and equipment
CN118132969A
Language large model training method, system and device and computer readable storage medium
CN118210895A
Image generation method and device, electronic equipment and storage medium
CN118298062A
File layout analysis and picture information extraction method for big language model RAG questions and answers
CN118364785A
Large-model natural language document query system and method based on full-text search
CN118838984A