Processing method and system for adaptively cutting long table image to realize content analysis
By adaptively cutting images and visual-language model (VLM) to extract text information from long table pictures, the problem of poor recognition effect in the prior art when processing ultra-long and complex tables is solved, and efficient content analysis of ultra-long tables is achieved.
Patent Information
- Application Number
- CN202411885631.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-19
- Publication Date
- 2025-05-23
AI Technical Summary
The prior art has problems with narrow applicable fields and poor identification results when dealing with ultra-long and complex tables, especially in scenarios where the table boundaries are unclear or the structure is irregular, making it difficult to effectively analyze.
The text information in long table pictures is extracted by adaptively cutting images and visual-language model (VLM), and the content analysis of the ultra-long tables is realized. This method does not require detecting cell vertices or predicting cell relationships, but directly uses VLM to organize cell content, supporting the processing of ultra-long tables.
It significantly simplifies the operation process, improves processing efficiency and accuracy, can directly support the processing of ultra-long tables, and avoids complex considerations of the content and relationships between large numbers of cells.
Smart Images

Figure CN120032383A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of information technology, and in particular to a processing method and system for adaptively cutting long table images to implement content analysis. Background Art
[0002] With the rapid advancement of informatization and digitalization, a large amount of important data is stored in the form of electronic documents. Tables, as the core form of data organization and display, are widely used in various fields, such as financial statements, scientific research documents, government reports and commercial contracts. Therefore, the research and application of table image parsing technology has received widespread attention, but there are still deficiencies in processing extra-long and complex tables.
[0003] Existing technologies rely on frame detection or cell relationship prediction, which makes it difficult to cope with scenarios where table boundaries are unclear or structures are irregular; methods based on graph neural networks require a large amount of labeled data and perform poorly with noise and irregular tables; text positioning and cell segmentation methods have strict requirements on the input image size, and the recognition effect is easily affected by scaling and cropping; training parsing models consumes a lot of time and resources, and it is difficult to quickly handle complex tasks.
[0004] Therefore, there is an urgent need for a method and a corresponding system that can process table images and implement content analysis. Summary of the invention
[0005] The present disclosure provides a processing method and system for adaptively cutting long table images to realize content analysis. By utilizing adaptive cutting images and a vision-language model (VLM) to extract text information from long table images, the present disclosure at least solves the technical problems of the existing table analysis technology, such as narrow application fields and poor recognition effects.
[0006] According to a first aspect of the present disclosure, a method for adaptively cutting a long table image to implement content parsing is provided, comprising the following steps:
[0007] Obtain a text file, convert the text file into an image, input the image into a layout detection model for analysis, and obtain a json file according to the analysis result;
[0008] Process the json file, extract the coordinates of the table category and crop them to obtain a table image;
[0009] The table image is input into the adaptive image function, and a judgment is made as to whether the table image needs to be cut and whether there is a cutting quantity, to obtain a judgment result, and a VLM analysis is performed according to the judgment result to obtain an output result.
[0010] According to the above aspects and any possible implementation, an implementation is further provided, wherein the json file includes a picture area and a page number corresponding to the picture;
[0011] The image region includes the category of the region determined by the layout detection model and the coordinates of the selected region.
[0012] According to the above aspects and any possible implementation, an implementation is further provided, wherein the categories of the areas determined by the layout detection model include a json file of the positioning box coordinates corresponding to the titles, texts, tables and parts to be ignored;
[0013] The selected area coordinates include the coordinates of the upper left corner and the lower right corner of the rectangular area, specifically: the upper left corner horizontal coordinate, the upper left corner vertical coordinate, the lower right corner horizontal coordinate and the lower right corner vertical coordinate.
[0014] According to the above aspects and any possible implementation, an implementation is further provided, wherein the table image is input into an adaptive image function, whether the table image needs to be cut and whether there is a cutting quantity, a judgment result is obtained, and a VLM analysis is performed according to the judgment result to obtain an output result. The process is:
[0015] Determine whether the table image needs to be cut. If not, input the table image and the preset first prompt word into the VLM model, completely extract the content of the table image through the VLM model and output the extraction result;
[0016] If so, determine whether there is a cutting number. If there is a cutting number, process the table image according to the cutting number to obtain the output result; if there is no cutting number, segment the table image based on the adaptive image segmentation function to obtain the segmented image, and use the VLM model and the pre-set second prompt word to completely extract the content of the segmented image and output the extraction result.
[0017] According to the above aspects and any possible implementation manner, an implementation manner is further provided, wherein the process of processing the table image according to the number of cuts to obtain the output result is:
[0018] The parameter cutting quantity value is judged. If the cutting quantity value is greater than 5, the output model reports an error "the model cannot process more than 5 pictures";
[0019] If the cutting quantity value is less than or equal to 5, a corresponding number of the table images are cut based on the cutting function to obtain a cutting result.
[0020] According to the above aspects and any possible implementation manner, an implementation manner is further provided, wherein the process of segmenting the table image based on the adaptive image segmentation function to obtain the segmented image is:
[0021] Input the table image into the adaptive image segmentation function, extract the height and width values corresponding to the table image, and calculate the aspect ratio value;
[0022] The ratio is input into a piecewise function to calculate the number of cuts required, and the height of the equal-width intercepted image is calculated according to the number of cuts;
[0023] The image is cropped according to the height of the equal-width intercepted image to obtain a segmented image.
[0024] According to the above aspects and any possible implementation, an implementation is further provided, wherein the process of inputting the ratio into the piecewise function to calculate the required number of cuts is:
[0025] When the ratio is greater than 0 and less than or equal to 1.0, the number of cuts is set to 1;
[0026] When the ratio is greater than 1.0 and less than or equal to 1.5, the number of cuts is set to 2;
[0027] When the ratio is greater than 1.5 and less than or equal to 2.0, the number of cuts is set to 3;
[0028] When the ratio is greater than 2.0 and less than or equal to 2.5, the number of cuts is set to 4;
[0029] When the ratio is greater than 2.5, the number of cuts is set to 5.
[0030] According to the above aspects and any possible implementation manner, an implementation manner is further provided, wherein the method for calculating the height of the equal-width intercepted pictures according to the number of cuts is:
[0031]
[0032] Among them, H is the height of the equal-width cut picture, h is the height of the table picture, δ is the picture overlap, and Q is the number of cuts. When the number of cuts is 1, it is regarded as an undivided image.
[0033] According to the above aspect and any possible implementation, there is further provided an implementation, wherein the first prompt word is “Please extract text content from the image”;
[0034] The second prompt word is "You are an image recognition expert, good at accurately extracting text information from table images, and can output semantic structures consistent with the table content. These images are split by sliding on a long image. Please sort out all the text content of the corresponding long image based on the text information of each image."
[0035] According to a second aspect of the present disclosure, there is provided a processing system for adaptively cutting a long table image to implement content parsing, comprising: a text file processing module, a cropping module and a content parsing module;
[0036] The text file processing module is used to obtain a text file, convert the text file into a picture, and input the picture into a layout detection model for analysis, and obtain a json file according to the analysis result;
[0037] The cropping module is used to process the json file, extract the coordinates of the table category and crop them to obtain a table image;
[0038] The content analysis module is used to determine whether the table image needs to be cut and whether there is a certain number of cuts, obtain a determination result, and perform VLM analysis based on the determination result to obtain an output result
[0039] Compared with the prior art, the present invention has the following technical effects:
[0040] The present invention uses adaptive image segmentation and a vision-language model (VLM) and sets appropriate prompt words, so that the model can accurately extract text information from long table images. The present invention does not need to predict the relationship between cells, does not need to detect cell vertices, does not need to prepare training data to train the model, and does not need to locate text in advance. This is due to the powerful ability of the vision-language model (VLM) in reading table content and arranging cell relationships. In particular, the present invention can directly support the processing of ultra-long tables. Compared with the existing method that needs to additionally consider the complexity of the content and relationship between a large number of cells, the present invention significantly simplifies the operation process and improves processing efficiency and accuracy.
[0041] It should be understood that the contents described in the summary of the invention are not intended to limit the key or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] The above and other features, advantages and aspects of the embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. The accompanying drawings are used to better understand the present solution and do not constitute a limitation of the present disclosure. In the accompanying drawings, the same or similar reference numerals represent the same or similar elements, among which:
[0043] Figure 1 A schematic flow chart of a processing method for adaptively cutting a long table image to implement content parsing according to an embodiment of the present disclosure is shown;
[0044] Figure 2 A schematic diagram of the structure of a processing system for adaptively cutting a long table image to implement content parsing according to an embodiment of the present disclosure is shown;
[0045] Figure 3 A block diagram of an exemplary electronic device according to an embodiment of the present disclosure is shown. DETAILED DESCRIPTION
[0046] In order to make the purpose, technical solution and advantages of the embodiments of the present disclosure clearer, the technical solution in the embodiments of the present disclosure will be clearly and completely described below in conjunction with the drawings in the embodiments of the present disclosure. Obviously, the described embodiments are part of the embodiments of the present disclosure, not all of the embodiments. Based on the embodiments in the present disclosure, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present disclosure.
[0047] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the present invention is further described in detail below with reference to the accompanying drawings and specific embodiments.
[0048] From the perspective of application prospects at the macro level, with the deepening of informatization and digitalization, a large amount of important data is stored in the form of electronic documents, and tables, as the core form of data organization and display, are widely used in various fields, such as financial statements, scientific research documents, government reports and commercial contracts. However, many tables are stored in the form of scans or pictures, which are difficult to use directly for data analysis or decision support. How to efficiently parse these unstructured table images into structured data has always been a core research issue in data processing and document automation. This demand has promoted the development of table image parsing technology.
[0049] From the perspective of specific technical implementation, table image parsing technology mainly involves text recognition and table structure extraction, among which the extraction of table structure, that is, the relationship between cells, is a major difficulty. Currently, there are several known methods:
[0050] 1. Method based on frame detection
[0051] By detecting the borders and cell vertex coordinates of the table, text is extracted and the table structure is constructed. However, this method relies too much on the accuracy of the borders and is difficult to handle complex tables with blurred, missing or irregular borders.
[0052] 2. Methods based on Graph Neural Network (GNN)
[0053] The graph model is constructed using cell content and vertex relationships, and the row and column structure is identified through classification. However, GNN model training relies on a large amount of labeled data, and the parsing effect is limited when faced with noisy or irregular tables.
[0054] 3. Text positioning and cell segmentation methods
[0055] This type of method locates the text first and then divides the cells to extract the layout information. However, it has strict requirements on the image input size and needs to crop or scale complex table images, which may lead to a decrease in recognition accuracy, especially in the case of very long tables.
[0056] 4. Training methods based on deep learning
[0057] Building a table parsing model for a specific task usually takes a lot of time and computing resources, has high requirements for hardware performance, and is difficult to efficiently handle complex tasks.
[0058] The method of the present invention does not need to detect cell vertices or predict the relationship between cells. The visual-language model (VLM) has the ability to read tables and organize the relationships between cell contents. It supports ultra-long tables (for other methods, ultra-long tables require consideration of more content and relationships between cells). It does not need to prepare training data to train the model and has high robustness.
[0059] Reference Figure 1 As shown, this embodiment provides a processing method for adaptively cutting a long table image to implement content parsing, comprising the following steps:
[0060] S101. Obtain a text file, convert the text file into an image, and input the image into a layout detection model for analysis, and obtain a json file according to the analysis result.
[0061] In this embodiment, the local detection model outputs a json file containing the positioning box coordinates corresponding to the categories of title, text, table, and parts to be ignored. The format of the json file is as follows:
[0062]
[0063]
[0064] Among them, "blocks" are different areas in the PDF image, "page_idx" is the page number corresponding to the PDF image, and "type" is the category of the area determined by the layout analysis model. "type" corresponds to four categories: "title", "text", "table", and "discarded". The category corresponding to "title" is the title, the category corresponding to "text" is the text, the category corresponding to "table" is the table, and the category corresponding to "discarded" is the part that needs to be ignored, such as unimportant information such as headers and footers; "bbox" refers to the coordinates of the selected area, which consists of four numbers, namely the coordinates of the upper left corner and the lower right corner of the rectangular area, specifically [horizontal coordinate of the upper left corner, vertical coordinate of the upper left corner, horizontal coordinate of the lower right corner, vertical coordinate of the lower right corner]. According to the above information, the position of different areas can be quickly located.
[0065] By matching the layout analysis model, the part of the json file with "type" as "table" is output, and the corresponding area coordinates "bbox" and page number "page_idx" are extracted and stored as a new json file. Extracting the page number is convenient for quickly locating the page number of the image in the file, and cropping the image according to the extracted area coordinates to obtain a .png, .jpg or .jpeg format image.
[0066] S102, processing the json file, extracting the coordinates of the table category and cropping them to obtain a table image.
[0067] S103, inputting the table image into the adaptive image function, judging whether the table image needs to be cut and whether there is a certain number of cuts, obtaining a judgment result, and performing VLM analysis according to the judgment result to obtain an output result.
[0068] In this embodiment, a cropped table image is input into the adaptive image function, and parameters are input at the same time, including whether cutting is required and the number of cuts; the parameters corresponding to whether cutting is required are "yes" and "no", and the default value of the parameter is "yes", that is, when the value of this parameter is not specifically set, the function needs to cut the image by default; the number of cuts parameter should be input with a non-zero positive integer, and if the number of cuts is unknown, it can be omitted.
[0069] If the parameter corresponding to whether to cut is "No", then the picture and the set prompt word-1 are directly input into the visual-language model (VLM), so that the model can completely extract the content of the table picture and obtain the output. In this embodiment, the set prompt word-1 is: "Please extract the text content in the picture".
[0070] If the parameter corresponding to whether to cut is "yes", check whether there is a cut quantity. Specifically, if the cut quantity value is less than or equal to 5, cut the corresponding number of pictures, otherwise the model reports an error output "the model cannot process more than 5 pictures".
[0071] If the parameter number of cuts is not entered, the image is input into the adaptive image segmentation function proposed in this patent to determine how to segment the image. The specific process is:
[0072] Input the table image into the adaptive image segmentation function, extract the height and width values corresponding to the table image, and calculate the aspect ratio;
[0073] The ratio is input into the piecewise function to calculate the number of cuts required, and the height of the equally wide intercepted image is calculated based on the number of cuts;
[0074] The image is cropped according to the height of the image with equal width to obtain the segmented image.
[0075] Specifically, in this embodiment, the piecewise function is as follows:
[0076]
[0077] Here, "ratio" is the ratio of height to width calculated in the previous step.
[0078] When the ratio is greater than 0 and less than or equal to 1.0, the number of segmentations is set to 1, that is, the image is not segmented, and the image and prompt word -1 are directly input into the visual-language model (VLM) to extract the text content in the table image; when the ratio is greater than 1.0 and less than or equal to 1.5, the number of segmentations is set to 2; when the ratio is greater than 1.5 and less than or equal to 2.0, the number of segmentations is set to 3, when the ratio is greater than 2.0 and less than or equal to 2.5, the number of segmentations is set to 4, and when the ratio is greater than 2.5, the number of segmentations is set to 5.
[0079] After obtaining the number of segments, the height of each equally-width intercepted image is calculated according to the following formula:
[0080]
[0081] Among them, H is the height of the equal-width cut picture, h is the height of the table picture, δ is the picture overlap, and Q is the number of cuts. When the number of cuts is 1, it is regarded as an undivided image.
[0082] Image overlap is a parameter of the adaptive image cutting function. The default value is 50%. This parameter is used to set the percentage of overlap between the next cropped image and the previous cropped image during the cropping process. There is a certain degree of overlap between the cropped images, which is conducive to the model's content splicing of multiple cropped images.
[0083] After the cropped images are output, the cropped images and the set prompt word-2 are input into the visual-language model (VLM) to completely extract the table image content, and the complete and structured table image content is output. The prompt word-2 set in this embodiment is: "You are an image recognition expert, good at accurately extracting text information from table images, and can output semantic structures consistent with the table content. These images are images that are segmented by sliding on a long image. Please sort out all the text content of the corresponding long image based on the text information of each image."
[0084] The setting of the prompt word -2 tells the model what identity to answer the question in. Secondly, it introduces the relationship between multiple input images to the model, and finally requires the model to extract the complete content of the table image.
[0085] like Figure 2 As shown, this embodiment also provides a processing system for adaptively cutting long table images to achieve content analysis, including: a text file processing module 1, a cutting module 2 and a content analysis module 3;
[0086] The text file processing module 1 is used to obtain a text file, convert the text file into an image, and input the image into a layout detection model for analysis to obtain an analysis result, and obtain a json file according to the analysis result;
[0087] The cropping module 2 is used to process the json file, extract the coordinates of the table category and crop them to obtain the table image;
[0088] The content analysis module 3 is used to judge whether the table image needs to be cut and whether there is a certain number of cuts, obtain a judgment result, and perform VLM analysis according to the judgment result to obtain an output result.
[0089] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working process of the described module can refer to the corresponding process in the aforementioned method embodiment, and will not be repeated here.
[0090] Figure 3A schematic block diagram of an electronic device 300 that can be used to implement an embodiment of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or required herein.
[0091] The electronic device 300 includes a computing unit 301, which can perform various appropriate actions and processes according to a computer program stored in a ROM 302 or a computer program loaded from a storage unit 308 into a RAM 303. In the RAM 303, various programs and data required for the operation of the electronic device 300 can also be stored. The computing unit 301, the ROM 302, and the RAM 303 are connected to each other via a bus 304. An I / O interface 305 is also connected to the bus 304.
[0092] A number of components in the electronic device 300 are connected to the I / O interface 305, including: an input unit 306, such as a keyboard, a mouse, etc.; an output unit 307, such as various types of displays, speakers, etc.; a storage unit 308, such as a disk, an optical disk, etc.; and a communication unit 309, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 309 allows the electronic device 300 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.
[0093] The computing unit 301 may be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 301 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, digital signal processors (DSPs), and any appropriate processors, controllers, microcontrollers, etc. The computing unit 301 performs the various methods and processes described above, such as the processing method for implementing content parsing by adaptively cutting long table images. For example, in some embodiments, the processing method for implementing content parsing by adaptively cutting long table images may be implemented as a computer software program, which is tangibly included in a machine-readable medium, such as a storage unit 308. In some embodiments, part or all of the computer program may be loaded and / or installed on the electronic device 300 via ROM 302 and / or communication unit 309. When the computer program is loaded into RAM 303 and executed by the computing unit 301, one or more steps of the processing method for implementing content parsing by adaptively cutting long table images described above may be executed. Alternatively, in other embodiments, the computing unit 301 may be configured to execute the processing method for adaptively cutting a long table image to implement content parsing in any other appropriate manner (for example, by means of firmware).
[0094] Various implementations of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), load programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0095] The program code for implementing the method of the present disclosure may be written in any combination of one or more programming languages. These program codes may be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, so that the program code, when executed by the processor or controller, enables the functions / operations specified in the flow chart and / or block diagram to be implemented. The program code may be executed entirely on the machine, partially on the machine, partially on the machine and partially on a remote machine as a stand-alone software package, or entirely on a remote machine or server.
[0096] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or equipment. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0097] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0098] The systems and techniques described herein may be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system may be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.
[0099] A computer system may include a client and a server. The client and the server are generally remote from each other and usually interact through a communication network. The relationship of client and server is generated by computer programs running on respective computers and having a client-server relationship with each other. The server may be a cloud server, a server of a distributed system, or a server combined with a blockchain.
[0100] It should be understood that the various forms of processes shown above can be used to reorder, add or delete steps. For example, the steps recorded in this disclosure can be executed in parallel, sequentially or in different orders, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and this document does not limit this.
[0101] The above specific implementations do not constitute a limitation on the protection scope of the present disclosure. It should be understood by those skilled in the art that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modification, equivalent substitution and improvement made within the spirit and principle of the present disclosure shall be included in the protection scope of the present disclosure.
Claims
1. A method for adaptively cutting long table images to achieve content parsing, characterized in that: The following steps are involved: Obtain a text file, convert the text file into an image, input the image into a layout detection model for analysis to obtain an analysis result, and obtain a json file according to the analysis result; Process the json file, extract the coordinates of the table category and crop them to obtain a table image; The table image is input into the adaptive image function, and a judgment is made as to whether the table image needs to be cut and whether there is a cutting quantity, to obtain a judgment result, and the judgment result is subjected to VLM analysis to obtain an output result.
2. The method for adaptively cutting long table images to achieve content analysis according to claim 1, characterized in that: The json file includes the image area and the page number corresponding to the image; The image region includes the category of the region determined by the layout detection model and the coordinates of the selected region.
3. The method for adaptively cutting long table images to achieve content analysis according to claim 2, characterized in that: The categories of the areas determined by the layout detection model include a json file of the positioning box coordinates corresponding to the titles, texts, tables and parts to be ignored; The selected area coordinates include the coordinates of the upper left corner and the lower right corner of the rectangular area, specifically: the upper left corner horizontal coordinate, the upper left corner vertical coordinate, the lower right corner horizontal coordinate and the lower right corner vertical coordinate.
4. The method for implementing content analysis by adaptively cutting long table images according to claim 1 is characterized in that: The process of inputting the table image into the adaptive image function, judging whether the table image needs to be cut and whether there is a cutting quantity, obtaining a judgment result, and performing VLM analysis on the judgment result to obtain an output result is as follows: Determine whether the table image needs to be cut. If not, input the table image and the preset first prompt word into the VLM model, completely extract the content of the table image through the VLM model and output the extraction result; If yes, determine whether there is a cutting quantity, and if there is a cutting quantity, process the table image according to the cutting quantity to obtain an output result; If there is no cut quantity, the table image is segmented based on the adaptive image segmentation function to obtain a segmented image, and the segmented image content is completely extracted through the VLM model and the preset second prompt word and the extraction result is output.
5. The method for implementing content analysis by adaptively cutting long table images according to claim 4 is characterized in that: The process of processing the table image according to the cutting quantity to obtain the output result is: Determine the parameter cutting quantity value. If the cutting quantity value is greater than 5, the output model reports an error "the model cannot process more than 5 pictures"; If the cutting quantity value is less than or equal to 5, a corresponding number of the table images are cut based on the cutting function to obtain a cutting result.
6. The method for implementing content analysis by adaptively cutting long table images according to claim 5 is characterized in that: The process of segmenting the table image based on the adaptive image segmentation function to obtain the segmented image is as follows: Input the table image into the adaptive image segmentation function, extract the height and width values corresponding to the table image, and calculate the aspect ratio value; The ratio is input into a piecewise function to calculate the required number of cuts, and the height of the equal-width intercepted image is calculated according to the number of cuts; The image is cropped according to the height of the equal-width intercepted image to obtain a segmented image.
7. The method for adaptively cutting a long table image to realize content analysis according to claim 6 is characterized in that: The process of inputting the ratio into the piecewise function to calculate the required number of cuts is: When the ratio is greater than 0 and less than or equal to 1.0, the number of cuts is set to 1; When the ratio is greater than 1.0 and less than or equal to 1.5, the number of cuts is set to 2; When the ratio is greater than 1.5 and less than or equal to 2.0, the number of cuts is set to 3; When the ratio is greater than 2.0 and less than or equal to 2.5, the number of cuts is set to 4; When the ratio is greater than 2.5, the number of cuts is set to 5.
8. The method for implementing content analysis by adaptively cutting long table images according to claim 6 is characterized in that: The method for calculating the height of the equal-width intercepted image according to the number of cuts is: Among them, H is the height of the equal-width cut picture, h is the height of the table picture, δ is the picture overlap, and Q is the number of cuts. When the number of cuts is 1, it is regarded as an undivided image.
9. The method for implementing content analysis by adaptively cutting long table images according to claim 4 is characterized in that: The first prompt word is "Please extract the text content in the picture"; The second prompt word is "You are an image recognition expert, good at accurately extracting text information from table images, and can output semantic structures consistent with the table content. These images are split by sliding on a long image. Please sort out all the text content of the corresponding long image based on the text information of each image." 10. A processing system for adaptively cutting long table images to achieve content analysis, characterized in that: include: A text file processing module (1), a cropping module (2) and a content parsing module (3); The text file processing module (1) is used to obtain a text file, convert the text file into a picture, and input the picture into a layout detection model for analysis, and obtain a json file according to the analysis result; The cropping module (2) is used to process the json file, extract the coordinates of the table category and crop them to obtain a table image; The content analysis module (3) is used to judge whether the table image needs to be cut and whether there is a certain number of cuts, obtain a judgment result, and perform VLM analysis based on the judgment result to obtain an output result.