Method for analyzing table based on multi-modal large language model

Through LoRA fine-tuning of multimodal large language model, the semantic understanding and high cost problems of traditional table analysis methods are solved, and efficient and fine table analysis and parameter extraction are achieved, which is suitable for the field of test and measurement.

CN120279571APending Publication Date: 2025-07-08CHENGDU TIANHENG INSTR EQUIP CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510367484.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-26
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

Traditional table analytics are difficult to understand the relationship between the semantic meanings and measurement parameters behind the table, and the full parameter fine-tuning of multimodal large language models is expensive and lacks lightweight customized training methods.

Method used

The multimodal large language model is trained by LoRA lightweight fine-tuning method, and the input and mean square error loss functions are optimized through image-text, multimodal feature vectors are generated, and measurement parameters and equipment characteristics are extracted.

Benefits of technology

It realizes the fine customization and training efficiency of table analysis without significantly increasing parameters, provides deep semantic table structure and parameter interpretation, and improves the practical value of the analysis results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120279571A_ABST
    Figure CN120279571A_ABST
Patent Text Reader

Abstract

The invention discloses a method for analyzing a table based on a multi-modal large language model, which comprises the following steps of: S1, collecting a data set, and training the multi-modal large language model; s2, analyzing and redrawing the chart of the data set in the step S1; and S3, the trained multi-modal large language model extracts measurement parameters and information from the edited table. The method has the beneficial effects that the LoRA training is performed on the data set in the field of test and measurement, and the multi-modal large language model is utilized to analyze and redraw the given table and chart, so that the visual display of the original data and the highlighting of key parameters are realized, a user can conveniently and quickly obtain key information, and the user experience is improved. According to the method, a table structure with deep semantics and parameter interpretation are obtained, on the basis of LoRA fine tuning, on the premise that parameters are not greatly increased, fine customization of the multi-modal large language model on table analysis in the field of test and measurement is achieved, and the training and reasoning efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of test and measurement, and particularly to a method for parsing tables based on a multimodal large language model. Background Art

[0002] Data in the field of test and measurement usually contains rich information such as experimental parameters, units, orders of magnitude, measurement ranges, quantization errors, and device specifications. This information often appears in tables, charts, and data lists with special formats. Traditional table parsing mainly relies on rule-based or structured extraction methods, which simply map the rows and columns of the table and are difficult to understand the semantic meaning behind the table, the relationships between measurement parameters, and the differences between parameters of specific instruments (such as oscilloscopes and spectrum analyzers).

[0003] With the development of multimodal large language model technology, it has become possible to integrate images, text, and structured data into a unified semantic vector space. Such models can cross the format and visual structure of the table and directly extract useful semantic information from its content. Especially in the field of test and measurement, through multimodal models, the actual meaning of the table can be deeply understood from the given table and relevant indication information, such as the frequency band information, vertical sensitivity, time base, measurement accuracy, quantization unit, and order of magnitude of a certain type of oscilloscope, thus providing users with high-level knowledge extraction.

[0004] However, the parameter scale of multimodal large language models is huge, and full-parameter fine-tuning is costly and time-consuming. In practical applications, it is often desirable to perform lightweight fine-tuning on existing base models to adapt to tasks in specific fields (such as the test and measurement field). At this time, the LoRA (Low-Rank Adaptation) fine-tuning method came into being. LoRA greatly reduces the number of fine-tuning parameters by adding low-rank incremental parameters to the original weights, thereby improving training efficiency and reducing resource consumption. Currently, there is still a lack of a method that combines multimodal large language models with LoRA lightweight fine-tuning technology in table parsing in the test and measurement field. Traditional solutions mostly rely on rule-based parsing and manual post-processing, unable to truly achieve in-depth understanding of table semantics and lacking effective lightweight customization training means to adapt to specific measurement field datasets. Summary of the Invention

[0005] The purpose of the present invention is to overcome the shortcomings of the prior art and provide a method for parsing tables based on a multimodal large language model.

[0006] The purpose of the present invention is achieved through the following technical solutions: A method for parsing tables based on a multimodal large language model, comprising the following steps:

[0007] S1: Collect a dataset and train the multimodal large language model;

[0008] S2: Parse and redraw the charts in the dataset in step S1;

[0009] S3: The trained multimodal large language model extracts measurement parameters and information from the edited table.

[0010] Preferably, in step S1, the following steps are further included:

[0011] S11: Select the architecture and base model of the multimodal large language model;

[0012] S12: Convert the table data in the dataset into images and corresponding text descriptions to form image-text pairs;

[0013] S13: Input the image-text pairs into the multimodal model and train the multimodal large language model by the LoRA fine-tuning method.

[0014] Preferably, in step S1, the dataset includes the parameter table of the test measurement device, the instruction manual, the standard unit, and the order of magnitude specification.

[0015] Preferably, in step S13, by introducing a low-rank increment matrix ΔW into the weight matrix of the pre-trained model, the fine-tuned weight matrix is W + ΔW.

[0016] Preferably, the increment matrix ΔW is decomposed into the product of low-rank matrices A and B, where d and k are the dimensions of the original weight matrix W, and r is a preset low-rank parameter.

[0017] Preferably, in step S13, the training process specifically includes the following steps:

[0018] S13.1: Input the image-text pairs into the multimodal large language model and generate multimodal feature vectors by processing the input through the base model;

[0019] S13.2: Adopt the mean square error loss function to adapt to the continuous value features and parameter relationships of the table data,

[0020] LME = 1Ni = 1N|y LoRA (i)-y true (i)|2;

[0021] where, y LoRA (i) is the output of the model after LoRA fine-tuning, y true (i) is the true target output, and N is the number of samples;

[0022] S13.3: By minimizing the mean square error loss function, use the gradient descent method to optimize the parameters of the increment matrices A and B,

[0023]

[0024] Among them, η is the learning rate.

[0025] The present invention has the following advantages:

[0026] 1. The present invention conducts LoRA training on the dataset in the field of test and measurement, and uses a multi-modal large language model to parse and redraw the given tables and charts, realizing the intuitive display of the original data and highlighting the key parameters, facilitating users to quickly obtain key information, obtaining a table structure and parameter interpretation with deep semantics. Based on LoRA fine-tuning, the present invention realizes the fine customization of the multi-modal large language model for table parsing in the field of test and measurement without significantly increasing the parameters, improving the training and inference efficiency.

[0027] 2. The present invention extracts the measurement parameters and device characteristics in the table through multi-modal feature understanding and semantic modeling, rather than simply reading the row and column structure of the table, so that the parsing result has more practical value.

[0028] 3. The present invention better adapts to the parsing of quantization indicators, units, and order of magnitude characteristics in the table by using the mean square error loss function in training and the image-text semantic fusion of the multi-modal model. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] Figure 1 is a schematic diagram of the method flow for parsing a table based on a multi-modal large language model;

[0030] Figure 2 is a schematic diagram of parsing and redrawing the original table in the field of test and measurement;

[0031] Figure 3 is a schematic diagram of LoRA training the multi-modal large language model. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0032] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Usually, the components of the embodiments of the present invention described and illustrated herein can be arranged and designed in various different configurations.

[0033] Accordingly, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely represents selected embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.

[0034] It should be noted that, without conflict, the embodiments in the present invention and the features in the embodiments may be combined with each other.

[0035] It should be noted that similar reference numerals and letters indicate similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings.

[0036] In the description of the present invention, it should be noted that the orientation or positional relationship indicated by the terms "center", "upper", "lower", "left", "right", "vertical", "horizontal", "inner", "outer", etc. is based on the orientation or positional relationship shown in the drawings, or the orientation or positional relationship in which the invention product is usually placed during use, or the orientation or positional relationship commonly understood by those skilled in the art. It is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as a limitation to the present invention. In addition, the terms "first", "second", etc. are only used for distinguishing descriptions and cannot be understood as indicating or implying relative importance.

[0037] In the description of the present invention, it should also be noted that unless otherwise clearly specified and defined, the terms "set", "install", "connect", and "couple" should be understood in a broad sense. For example, it may be a fixed connection, a detachable connection, or an integral connection; it may be a mechanical connection or an electrical connection; it may be directly connected or indirectly connected through an intermediate medium, and it may be the communication inside two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific situations.

[0038] In this embodiment, as Figure 1 and Figure 2 shown, a method for parsing a table based on a multimodal large language model includes the following steps:

[0039] S1: Collect a data set and train the multimodal large language model; preferably, the data set includes parameter tables of test measurement devices, instructions, standard units, and order-of-magnitude specifications.

[0040] S2: Parse and redraw the charts in the data set in step S1;

[0041] S3: The trained multi-modal large language model extracts measurement parameters and information from the edited table. By performing LoRA training on the dataset in the test measurement field and using the multi-modal large language model to parse and redraw the given table and chart, the intuitive display of the original data and the highlighting of key parameters are realized, facilitating users to quickly obtain key information, obtaining a table structure and parameter interpretation with deep semantics. Based on LoRA fine-tuning, the present invention realizes the fine customization of the multi-modal large language model for table parsing in the test measurement field without significantly increasing parameters, improving the training and inference efficiency.

[0042] Further, in step S1, the following steps are further included:

[0043] S11: Select the architecture and basic model of the multi-modal large language model;

[0044] S12: Convert the table data in the dataset into images and corresponding text descriptions to form image-text pairs;

[0045] S13: Input the image-text pairs into the multi-modal model and train the multi-modal large language model by the LoRA fine-tuning method. Specifically, select a pre-trained multi-modal large language model as the basic model, such as CLIP or a similar model, which can process the input of images and text. Through multi-modal feature understanding and semantic modeling, the measurement parameters and device characteristics in the table are extracted, rather than simply reading the row and column structure of the table, so that the parsing result has more practical value.

[0046] In this embodiment, in step S13, by introducing a low-rank incremental matrix ΔW into the weight matrix of the pre-trained model, the fine-tuned weight matrix is W + ΔW. Specifically, the incremental matrix ΔW is decomposed into the product of low-rank matrices A and B, where, d and k are the dimensions of the original weight matrix W, and r is a preset low-rank parameter.

[0047] Further, as Figure 3 shown, in step S13, the training process specifically includes the following steps:

[0048] S13.1: Input the image-text pairs into the multi-modal large language model and generate multi-modal feature vectors by processing the input through the basic model;

[0049] S13.2: Adopt the mean square error loss function to adapt to the continuous value features and parameter relationships of the table data,

[0050] LME = 1Ni = 1N|y LoRA (i)-y true (i)|2;

[0051] where, yLoRA (i) is the model output after LoRA fine-tuning, y true (i) is the true target output, and N is the number of samples;

[0052] S13.3: By minimizing the mean squared error loss function, use the gradient descent method to optimize the parameters of the incremental matrices A and B,

[0053]

[0054]

[0055] where η is the learning rate. By using the mean squared error loss function in training and leveraging the image-text semantic fusion of the multimodal model, the parsing of the quantization metrics, units, and order of magnitude characteristics in the table is better adapted.

[0056] The following is an example: Select the oscilloscope parameter table as the test object,

[0057] (1). Data collection: Collect table data containing oscilloscope parameters of different models, a total of 150 oscilloscope parameter table images, which cover key information such as bandwidth, sampling rate, number of channels, sensitivity, and time base unit.

[0058] Data preprocessing: Process the collected table images, including adjusting the image resolution, removing noise, unifying the table format, etc.; and pair each table image with the corresponding text description to form an image-text pair for use in the training of the multimodal model.

[0059] LoRA fine-tuning training:

[0060] 1). Selection of the base model: Select the pre-trained CLIP model as the base multimodal large language model;

[0061] 2). LoRA configuration: Set the low-rank parameter of LoRA to r = 4 to achieve low-rank approximation of the parameters;

[0062] 3). Training settings: The learning rate η is 0.001, the batch size is 32, the number of training epochs is 15, and the loss function is the mean squared error;

[0063] 4). Training process:

[0064] Input the processed image-text pairs into the CLIP model;

[0065] Use the mean squared error loss function to optimize the LoRA incremental parameters A and B, so that the vector features output by the model better match the continuous value features and parameter relationships of the table data; during the training process, only update the LoRA incremental parameters and keep the original model weights unchanged to achieve lightweight fine-tuning.

[0066] (2), Chart Redrawing and Editing:

[0067] Initial Chart Selection: Select a typical oscilloscope parameter table containing information such as bandwidth, sampling rate, number of channels, sensitivity, and time base unit.

[0068] Redrawing Process: According to the preset instructions, structurally redraw the table, including: adjusting the bandwidth unit from MHz to GHz to meet the measurement requirements of higher frequencies; adding annotations to highlight the highest sampling rate and key parameters; optimizing the axis units to ensure the accuracy and readability of data representation.

[0069] Generate a structurally and visually optimized oscilloscope parameter chart for efficient parsing by the multimodal model.

[0070] (3), Semantic Parsing and Information Extraction:

[0071] Model Parsing: Input the edited table image into a multimodal large language model fine-tuned with LoRA for semantic parsing.

[0072] Information Extraction: The model automatically extracts the key measurement parameters and information of the oscilloscope, including but not limited to: frequency band range: DC to 1 GHz; sampling rate: 2 GS / s; number of channels: 4; sensitivity: 100 mV / div; rise time: 350 ps.

[0073] (4) Model Output:

[0074] Parsing Results: Frequency band range: DC to 1 GHz; sampling rate: 2 GS / s; number of channels: 4; sensitivity: 100 mV / div; rise time: 350 ps

[0075] Performance Metrics: Parsing accuracy: 95%; average MSE loss: 0.02; parameter update amount: LoRA incremental parameters account for 10% of the original model parameters; training time: approximately 2 hours (saving approximately 70% of the training time compared to full-parameter fine-tuning).

[0076] (5) Experimental Conclusions:

[0077] 1), High Accuracy: The parsing accuracy reaches 95%, significantly better than the 80% accuracy of traditional rule-based or structured extraction methods, indicating that the multimodal large language model has significant advantages in semantic understanding and information extraction.

[0078] 2), Lightweight Fine-tuning: Using LoRA for fine-tuning only requires updating the incremental parameters, and the parameter update amount is only 10% of the original model, significantly reducing the training and storage costs. At the same time, the training time is shortened by approximately 70%, improving the training efficiency.

[0079] 3) Deep semantic analysis: Through the semantic understanding ability of the multimodal model, not only the key numerical information in the table is accurately extracted, but also the correlation between parameters is understood, providing a more practical knowledge refinement result.

[0080] 4) Strong practicability: The edited chart is redrawn visually, facilitating users to quickly obtain key information, assisting test engineers in making rapid and accurate decisions and selections in actual work, and improving work efficiency and data accuracy.

[0081] Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments or perform equivalent replacements for some of the technical features. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A method for parsing tables based on a multi-modal large language model, characterized in that: It includes the following steps: S1: Collect a dataset and train a multimodal large language model; S2: Parse and redraw the charts in the dataset in step S1; S3: The trained multimodal large language model extracts measurement parameters and information from the edited table.

2. The method for parsing a table based on a multimodal large language model according to claim 1, wherein: In step S1, the following steps are further included: S11: Select the architecture and basic model of the multimodal large language model; S12: Convert the table data in the dataset into images and corresponding text descriptions to form image-text pairs; S13: Input the image-text pairs into the multimodal model and train the multimodal large language model by the LoRA fine-tuning method.

3. The method for parsing a table based on a multimodal large language model according to claim 2, wherein: In step S1, the dataset includes a parameter table of a test measurement device, an instruction manual, standard units, and magnitude specifications.

4. The method for parsing a table based on a multimodal large language model according to claim 3, wherein: In step S13, by introducing a low-rank incremental matrix ΔW into the weight matrix of the pre-trained model, the fine-tuned weight matrix is W + ΔW.

5. The method for parsing a table based on a multimodal large language model according to claim 4, wherein: The incremental matrix ΔW is decomposed into the product of low-rank matrices A and B, where d and k are the dimensions of the original weight matrix W, and r is a preset low-rank parameter.

6. The method for parsing a table based on a multimodal large language model according to claim 5, wherein: In step S13, the training process specifically includes the following steps: S13.1: Input the image-text pairs into the multimodal large language model and generate multimodal feature vectors by processing the input through the basic model; S13.2: Adopt the mean square error loss function to adapt to the continuous value features and parameter relationships of the table data, LME = 1Ni = 1N|y LoRA (i)-y true (i)|2; Among them, y LoRA (i) is the output of the model after LoRA fine-tuning, y true (i) is the true target output, and N is the number of samples; S13.3: By minimizing the mean square error loss function, use the gradient descent method to optimize the parameters of the incremental matrices A and B, where η is the learning rate.