Table generation method and device
By identifying and stitching content cells and header cells in table images, the problem of low accuracy in image table recognition is solved, and higher accuracy in table generation is achieved.
Patent Information
- Application Number
- CN202510442168.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-09
- Publication Date
- 2025-07-18
AI Technical Summary
In the prior art, the recognition accuracy of picture tables is low, especially when processing complex tables, it is difficult to meet actual needs.
By obtaining the image of the table, identifying the header area and content area, generating content cells based on text content information, and splicing them with the header cells based on cell position information to generate a table.
Improves the semantic smoothness and position matching accuracy of content cells, and enhances the accuracy of table generation.
Smart Images

Figure CN120337888A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the technical field of table reconstruction, and specifically relates to a table generation method and apparatus thereof. Background Art
[0002] Picture tables are usually saved in the form of pictures or scanned documents, lacking structured data and cannot be directly used for subsequent processing and analysis.
[0003] In related technologies, there is a method of using optical character recognition (OCR) technology to recognize tables in pictures or scanned documents. However, the accuracy in recognizing complex tables still cannot meet the actual requirements. Picture tables usually contain multiple lines of text, merged cells, and content with different layout methods, which makes the OCR technology have a problem of low recognition accuracy when recognizing picture tables. Summary of the Invention
[0004] The purpose of the embodiments of this application is to provide a table generation method and apparatus thereof, which solves the problem of low recognition accuracy in recognizing picture tables in related technologies.
[0005] In a first aspect, the embodiments of this application provide a table generation method, including: obtaining a first image including a table, where the first image includes a header area and a content area; generating at least one content cell based on text content information, where the text content information is content information determined by performing image recognition on the content area, and the semantic information in the same content cell is mutually related; splicing at least one content cell with a header cell based on the cell position information of the content cell to generate a first table, where the header cell corresponds to the header area, and the cell position information is position information determined according to the content position information of the text content information.
[0006] In a second aspect, the embodiments of this application provide a table generation apparatus, including: an obtaining module, configured to obtain a first image including a table, where the first image includes a header area and a content area; a generating module, configured to generate at least one content cell based on text content information, where the text content information is content information determined by performing image recognition on the content area, and the semantic information in the same content cell is mutually related; a splicing module, configured to splice at least one content cell with a header cell based on the cell position information of the content cell to generate a first table, where the header cell corresponds to the header area, and the cell position information is position information determined according to the content position information of the text content information.
[0007] In a third aspect, an embodiment of the present application provides an electronic device, including a processor and a memory. The memory stores a program or instructions that can be run on the processor. When the program or instructions are executed by the processor, the steps of the method in the first aspect are implemented.
[0008] In a fourth aspect, an embodiment of the present application provides a readable storage medium, on which a program or instructions are stored. When the program or instructions are executed by a processor, the steps of the method in the first aspect are implemented.
[0009] In a fifth aspect, an embodiment of the present application provides a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor, and the processor is used to run a program or instructions to implement the steps of the method in the first aspect.
[0010] In a sixth aspect, an embodiment of the present application provides a computer program product, which is stored in a storage medium. The program product is executed by at least one processor to implement the method in the first aspect.
[0011] In the embodiments of the present application, the electronic device can identify the content area in the first image including a table, generate content cells in the content area based on the semantic information of the recognized text content information. Each content cell includes content with semantic relevance, and further determines the cell position information of each content cell based on the content position information of the text content information. Therefore, it can be spliced with the header cell based on the cell position information, ensuring smooth semantics inside the content cell, accurate division of the content cell, and improving the matching accuracy of the display position of the content cell and the position of the header cell, thereby improving the accuracy of generating a table based on the table image. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] Figure 1 FIG. 1 shows one of the schematic flowcharts of the table generation method provided in some embodiments of the present application;
[0013] Figure 2 FIG. 2 shows one of the schematic diagrams of the display interface provided in some embodiments of the present application;
[0014] Figure 3 FIG. 3 shows another schematic diagram of the display interface provided in some embodiments of the present application;
[0015] Figure 4 FIG. 4 shows yet another schematic diagram of the display interface provided in some embodiments of the present application;
[0016] Figure 5a FIG. 5 shows still another schematic diagram of the display interface provided in some embodiments of the present application;
[0017] Figure 5bShows the fifth schematic diagram of the display interface provided in some embodiments of the present application;
[0018] Figure 5c Shows the sixth schematic diagram of the display interface provided in some embodiments of the present application;
[0019] Figure 5d Shows the seventh schematic diagram of the display interface provided in some embodiments of the present application;
[0020] Figure 6a Shows the eighth schematic diagram of the display interface provided in some embodiments of the present application;
[0021] Figure 6b Shows the ninth schematic diagram of the display interface provided in some embodiments of the present application;
[0022] Figure 7 Shows the splicing schematic diagram of the content cell and the table header cell in some embodiments of the present application;
[0023] Figure 8 Shows the second schematic diagram of the flow of the table generation method provided in some embodiments of the present application;
[0024] Figure 9 Shows the schematic block diagram of the table generation device provided in some embodiments of the present application;
[0025] Figure 10 Shows the structural block diagram of the electronic device in some embodiments of the present application;
[0026] Figure 11 Shows the schematic diagram of the hardware structure of the electronic device provided in some embodiments of the present application. Detailed implementation manners
[0027] Next, the technical solutions in the embodiments of the present application will be clearly described with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are some, rather than all, of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application belong to the scope of protection of the present application.
[0028] The terms "first", "second", etc. in the description and claims of this application are used to distinguish similar objects, rather than to describe a specific order or sequence. It should be understood that such terms can be interchanged under appropriate circumstances, so that the embodiments of this application can be implemented in an order other than those illustrated or described herein, and the objects distinguished by "first", "second", etc. are generally of the same category, and do not limit the number of objects. For example, the first object can be one or multiple. In addition, "and / or" in the description and claims means at least one of the connected objects, and the character " / ", generally represents an "or" relationship between the associated objects before and after.
[0029] The following will be combined with the attached Figures 1 to 11 drawings to describe in detail the table generation method and device provided by the embodiments of this application through specific embodiments and their application scenarios.
[0030] In some embodiments of this application, a table generation method is provided. Figure 1 FIG. shows one of the flow diagrams of the table generation method provided in some embodiments of this application. As Figure 1 shown, the table generation method includes:
[0031] Step 102, obtain a first image including a table, where the first image includes a header area and a content area;
[0032] In the embodiments of this application, the main body of the first image is a table, the table includes a table header and a table body, the image area where the table header is located is the header area, and the area where the table body is located is the content area. The first image can be an image captured by an electronic device, and the first image can also be an image selected from the album of the electronic device or an image downloaded from the network.
[0033] Figure 2 FIG. shows one of the schematic diagrams of the display interface provided in some embodiments of this application. As Figure 2 shown, after the user clicks the capture button 201 in the capture interface to capture the first image 202, and then clicks the completion control 203, the first image 202 is recognized.
[0034] Figure 3 FIG. shows another schematic diagram of the display interface provided in some embodiments of this application. As Figure 3 shown, the user clicks the album view button 301 to trigger the display of multiple image thumbnails 302. The user selects the first image to be recognized from the multiple image thumbnails 302, triggering the recognition of the first image. Among them, Picture 1 is the first image.
[0035] It should be noted that two types of parameters are required to reconstruct the table in the first image. One of them is the header content for coordinate positioning within the header area, and the table content for displaying relevant information. Among them, the header content is displayed within the header area, and the table content is displayed within the content area.
[0036] Exemplarily, taking a class schedule as an example, the header area displays header content including "Weekday", "Period", etc. The header content is used to position the displayed table content, and the content area displays information such as "Course Name", "Teacher", "Location", etc.
[0037] Step 104, generate at least one content cell based on the text content information, where the text content information is the content information determined by image recognition of the content area. The semantic information within the same content cell is mutually related.
[0038] In the embodiment of the present application, the text content information is the content information obtained by recognizing the content area in the first image through image recognition technology. Exemplarily, the content area is recognized through OCR technology to obtain the text content information within the content area. Specifically, for example: the text content information can be the text content within the content area.
[0039] In the embodiment of the present application, after recognizing the text content information, through parsing and splicing of the text content information, content cells can be obtained. Each content cell is a text content with coherent semantics, and all the content within each content cell is semantically mutually related.
[0040] Specifically, the text content information is parsed through a semantic aggregation model, so as to aggregate the text content information with related semantics to form a content cell. Since the first image includes at least one cell with complete semantics, at least one content cell can be aggregated through the semantic aggregation model. The content cell is a cell in the table, that is, the content cell is editable text content.
[0041] Step 106, splice at least one content cell with a header cell based on the cell position information of the content cell to generate a first table, where the header cell corresponds to the header area, and the cell position information is the position information determined according to the content position information of the text content information.
[0042] In the embodiments of the present application, since the content cell is an editable text content based on the text content information, the cell position information of the content cell can be determined according to the content position information corresponding to the text content information. Specifically, the text content information included in each content cell is determined, the content position information corresponding to each text content information is determined, and the cell position information is determined according to the content position information of the text content information located at the edge position of the content cell.
[0043] It should be noted that the file format of the first table can be set according to actual needs.
[0044] In the embodiments of the present application, the header cell is the cell in the first table corresponding to the header area in the first image, and the content in the header cell is editable. The cell position information of the content cell is the relative position information between the content cell and the header cell, that is, according to the cell position information, the content cell and the header cell can be spliced.
[0045] Exemplarily, the number of content cells is i, and the i cell position information is expressed by the following expression: Cell_1:(x_c1_left,x_c1_right,y_c1_left,y_c1_right,text_1), Cell_2:(x_c2_left,x_c2_right,y_c2_left,y_c2_right,text_2)…Cell_i:(x_ci_left,x_ci_right,y_ci_left,y_ci_right,text_i). Wherein, Cell_1 to Cell_i are the 1st content cell to the i-th cell respectively, x_c1_left to x_ci_left are the x-axis coordinates on the left side of the 1st content cell to the x-axis coordinates on the left side of the i-th cell, x_c1_right to x_ci_right are the x-axis coordinates on the right side of the 1st content cell to the x-axis coordinates on the right side of the i-th cell, y_c1_left to y_ci_left are the y-axis coordinates on the left side of the 1st content cell to the y-axis coordinates on the left side of the i-th cell, and y_c1_right to y_ci_right are the y-axis coordinates on the right side of the 1st content cell to the y-axis coordinates on the right side of the i-th cell. Specifically, after determining the cell position information of the content cell, the relative position between each content cell and the header cell can be determined based on the cell position information, so as to splice the content cell and the header cell.
[0046] In the embodiments of the present application, the header cell can be a cell automatically recognized and generated by the electronic device based on the header area in the first image, or a cell generated by the user's manual operation.
[0047] Figure 4 FIG. 3 shows a schematic diagram of a display interface provided in some embodiments of the present application. As Figure 4 shown, after the user clicks the "Recognize and Reconstruct Table" button 401, the electronic device is triggered to execute the first image recognition and reconstruction function. When the user clicks the function button 402, a floating window 403 is triggered to be displayed. There are 3 option buttons 404 displayed in the floating window 403, and the user can select the ways of "fully automatic", "semi-automatic", and "manual" to recognize and reconstruct the table through the 3 option buttons 404.
[0048] Exemplarily, the fully automatic way is that the electronic device recognizes the entire first image, including recognizing the content area and the header area, constructs content cells based on the recognized content area, and header cells corresponding to the header area, and automatically splices the content cells and the header cells to obtain the first table.
[0049] Exemplarily, the semi-automatic way is that the user manually selects at least part of the header area and at least part of the content area that the electronic device needs to recognize. The electronic device recognizes the at least part of the header area and at least part of the content selected by the user, and performs reconstruction based on the recognition result.
[0050] Exemplarily, the manual way is that the user manually creates header cells and manually selects at least part of the content cells that need to be recognized and reconstructed. The electronic device automatically recognizes and generates some content cells, and then splices the some content cells with the manually created header cells.
[0051] Figure 5a FIG. 4 shows a schematic diagram of a display interface provided in some embodiments of the present application. Figure 5b FIG. 5 shows a schematic diagram of a display interface provided in some embodiments of the present application. Figure 5c FIG. 6 shows a schematic diagram of a display interface provided in some embodiments of the present application. Figure 5d FIG. 7 shows a schematic diagram of a display interface provided in some embodiments of the present application. As Figures 5a to 5d shown, the header cells 501 and blank cells 502 are displayed in the display interface. After the user clicks the blank cell 502, the electronic device can obtain the position information of the blank cell 502 clicked by the user. Then the user selects the content area in the first image 503 by means of box selection, recognizes and extracts the content cells 504 in the content area, and triggers to splice the content cells 504 to the position of the blank cell 502.
[0052] In the embodiments of the present application, the electronic device can identify the content area in the first image, generate content cells within the content area based on the semantic information of the recognized text content information, and each content cell contains content with semantic relevance. Also, based on the content position information of the text content information, the cell position information of each content cell is determined, so that it can be spliced with the header cell based on the cell position information, ensuring smooth semantics inside the content cell, accurate division of the content cell, and improving the matching accuracy of the display position of the content cell and the position of the header cell, thereby improving the accuracy of generating the first table based on the first image.
[0053] In some embodiments of the present application, after obtaining the first image including the table, the table generation method further includes: identifying the table text information and text position information in the first image, where the table text information and the text position information are in one-to-one correspondence; determining the text content information and header content information in the table text information, and determining the content position information and header position information in the text position information; generating a header cell based on the header content information and the header position information.
[0054] In the embodiments of the present application, the electronic device can automatically identify the table text information in the first image. The table text information includes all the text information in the first image, and the text position information corresponding to the table text information. The text position information is the position information corresponding to each table text information, that is, it can identify the paired table text information and text position information.
[0055] Specifically, each table text information in the image table and the text position information corresponding to each table text information are identified through OCR technology, where the text position information is the pixel position of the table text information in the first image.
[0056] In the embodiments of the present application, after obtaining the table text information and the table position information, it is necessary to classify the text content information and the header content information in the table text information, that is, specifically identify whether each table text information belongs to the text content information or the header content information, and use the text position information corresponding to the text content information as the content position information, and use the text position information corresponding to the header content information as the header position information. After obtaining the header content information and the header position information, a header cell can be constructed.
[0057] Specifically, the text content information in the tabular text information is identified by the semantic aggregation model, and the header content information in the tabular text information is identified by the header expression template library. At least some of the text content information is correlated. Therefore, the text content information can be identified by the semantic aggregation model, and the header content information is usually fixed vocabulary. Therefore, the header content information can be identified by the header expression template library.
[0058] In the embodiment of the present application, when the user needs to perform table reconstruction on the table content in the entire first image, the user can select the fully automatic recognition and reconstruction table function. At this time, the electronic device identifies all the tabular text information in the first image and the corresponding text position information, and further screens the text content information and the header content information in the tabular text information to facilitate subsequent automatic construction of content cells and header cells, and the operation of splicing the content cells and the header cells based on the position relationship.
[0059] In some embodiments of the present application, after obtaining the first image including the table, the table generation method further includes: in response to a first selection input for at least part of the content area, identifying the text content information and the content position information in the content area, where the text content information is at least part of the information selected by the first selection input in the content area; in response to a second selection input for at least part of the header area, identifying the header content information and the header position information in the header area, where the header content information is at least part of the information selected by the second selection input in the header area, and the header position information is the position of the header content information in the first image; generating a header cell based on the header content information and the header position information.
[0060] In the embodiment of the present application, the user can select and input to control the electronic device to select and identify at least part of the content area and at least part of the header area. Specifically, the user identifies and extracts the text content information in at least part of the content area through the first selection input, and determines the content position information corresponding to the extracted text content information. The user then identifies and extracts the header content information in at least part of the header area through the second selection input, and determines the header position information corresponding to the header content information.
[0061] Figure 6a FIG. 8 shows a schematic diagram of a display interface provided in some embodiments of the present application. Figure 6b FIG. 9 shows a schematic diagram of a display interface provided in some embodiments of the present application. As Figure 6a shown, after the user performs a first selection input on the first image, a first selection box 601 is displayed in the content area of the first image. The electronic device identifies the text content information within the range of the first selection box 601 and obtains the content position information corresponding to the text content information within the first selection box 601. AsFigure 6b As shown, after the user performs a second selection input on the first image, a second selection box 602 is displayed within the header area of the first image. The electronic device recognizes the header content information within the range of the second selection box 602 and obtains the header position information corresponding to the header content information within the second selection box 602.
[0062] Specifically, each text content information selected by the first selection input, the content position information corresponding to each text content information, and each header content information selected by the second selection input, and the header position information corresponding to each header content information are recognized through OCR technology.
[0063] In the embodiments of the present application, after obtaining the table text information and the table position information, it is necessary to classify the text content information and the header content information in the table text information, that is, specifically identify whether each table text information belongs to the text content information or the header content information, and use the text position information corresponding to the text content information as the content position information, and use the text position information corresponding to the header content information as the header position information. After obtaining the header content information and the header position information, a header cell can be constructed.
[0064] In the embodiments of the present application, when the user needs to perform table reconstruction on at least part of the table content in the first image, the user can select the semi-automatic recognition and reconstruction table function. The user selects the text content information and the content position information in at least part of the content area to be recognized through the first selection input, and selects the at least part of the header content information and the header position information to be recognized through the second selection input, which is convenient for automatically constructing content cells and header cells according to the selected text content information and header content information, and performing the operation of splicing the content cells and the header cells based on the position relationship.
[0065] In some embodiments of the present application, the header content information includes multiple header content sub-information. Generating a header cell based on the header content information and the header position information includes:
[0066] Obtaining an expression template library corresponding to the header content information; sequentially extracting the header content sub-information in the header content information along a first preset direction, and searching for the header content sub-information in the expression template library; in the case where the header content sub-information is found in the expression template library, determining multiple header position sub-information corresponding to the multiple header content sub-information according to the first preset direction and the header position information; generating a header cell according to the multiple header content sub-information and the multiple header position sub-information.
[0067] In the embodiments of the present application, the header content information includes multiple header content sub-information. The header content sub-information can be text content such as phrases, word groups, words, characters, etc. Each header content sub-information corresponds to a header position sub-information, and each header content sub-information is located in a cell, that is, the header position sub-information is the pixel position of the corresponding header content sub-information in the first image.
[0068] Specifically, since the text content in the header content information is usually fixed text content, such as "Period", "Week" in the class schedule, "Personnel", "Date" in the attendance sheet, etc. Therefore, by traversing and querying the expression template library along the first preset direction and determining whether there is the header content sub-information in the header content information in the expression template library, the extension direction of the header content information can be determined. If the header content sub-information is traversed along the first preset direction and the corresponding content is found in the expression template library, it is determined that the header content sub-information extends along the first preset direction. After determining the first preset direction of the extension of the header content information and the header position information, the header position sub-information corresponding to each header content sub-information can be based on the header position information and the first preset direction, thereby constructing the header cell to ensure that the content and layout of the header cell match the header area in the first image. It should be noted that the first preset direction includes the horizontal direction and the vertical direction.
[0069] Such as Figure 6aAs shown, arrow A and arrow B show the first preset direction. Exemplarily, the header content sub-information includes information of the "week" type and information of the "class number" type. Specifically, for example, it is determined whether the information of the "week" type is on the horizontal axis or the vertical axis. First, an expression template library for the information of the "week" type is established. Secondly, the order from top to bottom or from left to right is used as the first preset direction, and each header content sub-information is sequentially matched with the content in the expression template library. If a match is found in the expression template library for the first time, the matched week header expression is marked as w0, and its coordinates are (w0_x1, w0_x2, w0_y1, w0_y2). Next, starting from w0, extend to the left and below respectively to obtain the header content sub-information and header position sub-information to be parsed. Subsequently, parse whether these header content sub-informations match the expression template library of the week header. Finally, according to the number of matches in the horizontal and vertical directions, select the one with the largest number of matches to determine whether the week is on the horizontal axis or the vertical axis, so as to determine the extension direction of the header content information, and then the header position sub-information of each header content sub-information can be determined according to the header position information. Specifically, for example, the header position sub-information of "clear" includes: Week_1: (x_w1_left, x_w1_right, y_w1_left, y_w1_right), Week_2: (x_w2_left, x_w2_right, y_w2_left, y_w2_right) …, Week_7: (x_w7_left, x_w7_right, y_w7_left, y_w7_right).The sub-information of the header position for "Section" includes: Section_1:(x_s1_left+x_hs_left,x_s1_right+x_hs_left,y_s1_left+y_hs_left,y_s1_right+y_hs_left), Section_2:(x_s2_left+x_hs_left,x_s2_right+x_hs_left,y_s2_left+y_hs_left,y_s2_right+y_hs_left)…Section_i:(x_si_left+x_hs_left,x_si_right+x_hs_left,y_si_left+y_hs_left,y_si_right+y_hs_left). Among them, Week_1 to Week_7 are the 1st to 7th sub-information of the header position respectively, x_w1_left to x_w7_left are the left x-axis coordinates of the 1st to 7th sub-information of the header position respectively, x_w1_right to x_w7_right are the left x-axis coordinates of the 1st to 7th sub-information of the header position respectively, y_w1_left to y_w7_left are the left y-axis coordinates of the 1st to 7th sub-information of the header position respectively, and y_w1_right to y_w7_right are the left y-axis coordinates of the 1st to 7th sub-information of the header position respectively. Section_1 to Section_i are the 1st to i-th sub-information of the header position respectively, x_s1_left to x_si_left are the left x-axis coordinates of the 1st to i-th sub-information of the header position respectively, x_s1_right to x_si_right are the left x-axis coordinates of the 1st to i-th sub-information of the header position respectively, y_s1_left to y_si_left are the left y-axis coordinates of the 1st to i-th sub-information of the header position respectively, and y_s1_right to y_si_right are the left y-axis coordinates of the 1st to i-th sub-information of the header position respectively.
[0070] In the embodiment of the present application, by identifying and analyzing the header content information, the extension direction of each sub-information of the header content information can be determined, so as to determine the corresponding sub-information of the header position, and based on the sub-information of the header position, the header cell can be constructed to ensure that the layout and content of the header cell match the header area in the first image.
[0071] In some embodiments of the present application, at least one content cell is spliced with a header cell based on the cell position information of the content cell to generate a first table, including:
[0072] Based on the cell position information and multiple header position sub-information, at least one content cell is spliced with a header cell to obtain a first table.
[0073] In the embodiments of the present application, after identifying the cell position information and multiple header position sub-information, through the cell position information and multiple header position sub-information, the positional relationship between the content cell and the header cell can be determined, and based on the positional relationship between the content cell and the header cell, the content cell and the header cell are spliced to obtain a first table.
[0074] Figure 7 The splicing schematic diagram of the content cell and the header cell in some embodiments of the present application is shown, such as Figure 7As shown, when the number of content cells is multiple, each content cell is traversed in sequence, and the horizontal axis cell header position sub-information and the vertical axis cell header position sub-information in the header cell are combined for splicing. Specifically, for example: cell week positioning matching. For cell Cell_i, first, taking Cell_i as the reference point, extend upward or leftward. If the first preset direction of week_mode is the up-down direction, extend vertically upward; if the first preset direction of week_mode is the left-right extension, extend horizontally leftward, where week_mode is the "week" cell in the header cell. Secondly, according to the header content sub-information Week_i of "week", use the coordinates to locate which week the extension line intersects or is compatible with in the header content sub-information. Locate the week expression of this cell Cell_i, denoted as Cell_i_week. Then perform cell session positioning matching. For cell Cell_i, first, taking Cell_i as the reference point, extend upward or leftward. If the first preset direction of section_mode is the up-down direction, extend vertically upward; if section_mode extends vertically upward, if the first preset direction of section_mode is the left-right extension, extend horizontally leftward. Secondly, according to the header content sub-information Section_i of "session", use the coordinates to locate which sessions the extension line intersects or is compatible with in the header content sub-information. If there are multiple located session information, the session information needs to be merged. Therefore, the session expression of this cell Cell_i is denoted as Cell_i_section. Among them, arrow A in the figure shows the vertical extension direction, and arrow B shows the horizontal extension direction. Finally, convert Cell_i_week into Chinese expression, and the value range is Monday, Tuesday, Wednesday, Thursday, Friday, Saturday, Sunday. Secondly, convert Cell_i_section into Chinese expression, and the expression form is from the first session to the second session. Finally, update the content text sub-information of Cell_i to complete the splicing of the content cell and the header cell.
[0075] In the embodiment of the present application, the position relationship between the content cell and the header cell is determined through the cell position information and multiple header position sub-information, and based on the position relationship between the content cell and the header cell, the content cell and the header cell are spliced to obtain the first table, ensuring that the position relationship between the content cell and the header cell in the first table matches the position relationship between the content area and the header area in the first image, and improving the accuracy of the reconstructed table.
[0076] In some embodiments of the present application, the number of text content information is multiple, and the content position information corresponds to the text content information one by one; generating at least one content cell based on the text content information includes: identifying the semantic relevance between multiple text content information; dividing the multiple text content information into at least one content cell based on the semantic relevance; generating at least one content cell according to the at least one content cell.
[0077] In the embodiments of the present application, when multiple text content information is recognized, it is necessary to identify the semantic relevance between the multiple text content information, regard the text content information with semantic relevance as a group of text information, and generate a content cell according to a group of text information with semantic relevance, ensuring that the text content information in each content cell has semantic relevance and improving the accuracy of content area recognition in the first image.
[0078] Specifically, use a semantic aggregation model and a pixel smoothing model to parse each cell area in the first image. Identify the correlation between each table text information in the first image through the semantic aggregation model. First, compare the current table text information with the nearest row of table text information below, and determine whether the two paragraphs of table text information are relevant. If they are relevant, add the two rows of table text information to the relevant set, and continue to compare the nearest row of table text information below all the table text information in the relevant set, and so on in a loop until it is determined to be irrelevant and then stop. The table text information in several relevant sets with relevance is used as text content information, and each group of relevant sets corresponds to the text information in a content cell.
[0079] In some embodiments of the present application, before generating the first table by splicing at least one content cell with a header cell based on the cell position information of the content cell, the table generation method includes: obtaining the edge pixel coordinates of the content cell in the content area; moving the edge pixel coordinates along a second preset direction, and obtaining the pixel difference between the positions of the edge pixel coordinates before and after the movement; in the case where the pixel difference is greater than the difference threshold, determining the edge pixel coordinates at the current position as the cell position information.
[0080] In the embodiments of the present application, after the content cells are divided based on semantic relevance, it is necessary to calibrate the cell position information of the content cells. Specifically, the edge pixel coordinates of the content cell in the content area are determined based on the content position information of the text content information at the edge within the content cell. After obtaining the edge pixel coordinates, the edge pixel coordinates are moved in the second preset direction, and the pixel difference between the pixel values at the positions before and after the movement of the edge pixel coordinates is determined during the movement. If the pixel difference is greater than the difference threshold, it is determined that the actual edge position of the content cell has been found, and the moved edge pixel coordinates are determined as the cell position information.
[0081] It should be noted that the content cell has four edge pixel coordinates, and each edge pixel coordinate has two second preset extension directions. Specifically, the content cell includes two left and right sides and two upper and lower sides. The second preset extension directions of the two left and right sides include extending to the left and extending to the right, and the second preset extension directions of the two upper and lower sides include extending upward and extending downward.
[0082] Exemplarily, using the trained object detection model (YOLO, You Only Look Once) based on deep learning and the image pixels, aggregation is performed based on the position and spatial layout of the content cells, so that adjacent and logically related texts in the image can be recognized as a whole. First, the trained YOLO model is used to select the content cells. Then, for each box, there is an edge pixel coordinate representation (x1, x2, y1, y2). Pixel-level expansion and reduction are performed on the edge pixel coordinates of the box. For the two horizontal lines of each box, movement needs to be performed in the upper and lower directions. If the pixel values of the horizontal line after movement and the horizontal line before movement are within the allowable tolerance range, the movement needs to continue. If it is outside the allowable tolerance range, the movement stops, where the allowable tolerance is the difference threshold in the above embodiments. Similarly, the two vertical lines of each box need to be moved in the left and right directions. Finally, after the horizontal and vertical lines are moved, the box is redrawn, and the edge pixel coordinates are (n_x1, n_x2, n_y1, n_y2), where n_x1 is the horizontal axis value after the left vertical line is moved, n_x2 is the horizontal axis value after the right vertical line is moved, n_y1 is the vertical axis value after the upper horizontal line is moved, and n_y2 is the vertical axis value after the lower horizontal line is moved, so as to obtain the cell position information of each content cell.
[0083] In the embodiments of the present application, after the content cells are determined, the actual position of the content cells is calibrated based on the edge pixel coordinates of the content cells, so as to obtain the cell position information corresponding to the content cells, further improving the matching degree between the generated content cells and the content area in the first image.
[0084] In some embodiments of the present application, after splicing at least one content cell with a header cell based on the cell position information of the content cell to generate a first table, the following steps are further included: extracting the text content in the content cell; inputting the text content into a content recognition model to extract the first target content in the text content; typesetting the first target content according to a target format to generate a second target content; and replacing the content cell in the first table with the second target content to generate a second table.
[0085] In the embodiments of the present application, the text content in the content cell of the first table is the same as the text content in the first image. By extracting the first target content in the text content, typesetting the first target content according to the target format to obtain the second target content, and then replacing the first target content with the second target content, the content reconstruction of the content cell is completed, so that the text in each content cell is a standardized data structure, which is convenient for subsequent table processing such as merging cells in the second table.
[0086] It should be noted that when extracting the first target content, the first target content can be displayed to facilitate the user to view and adjust the first target content. For example Figure 5d As shown, exemplarily, according to the classification information of the first target content 505, the first target content 505 is displayed, and the user can edit the first target content. After the editing is completed, the user clicks the "Next" button 506 to trigger the action of reconstructing the content cell.
[0087] Exemplarily, the first table is a course schedule table, and the first target content is extracted through the content recognition model of the content cell. After the information of each content cell in the first table is spliced, the content cells in the first table need to be input to the content recognition model respectively. Among them, the prompt parameters include "week", "start class number", "end class number", "course name", "teacher", "location", "number of weeks of class". Specifically, for example, the training situation of the course information recognition model is as follows:
[0088] The prompt words include: You are a course schedule assistant, and your task is to extract the course name, teacher name, week, start class number, end class number, number of weeks, and classroom from the information provided by the user. The course output format is: [{"class":"course name","teacher":"teacher name","week":"week (number from 1 to 7)","start":"start class number (number)","end":"end class number (number)","round":"number of weeks (converted format)","room":"classroom"}], note: you only need to output the parameters with values.
[0089] Input parameters include: Modern Management, Teacher Z, Weeks 1 - 17, Friday, Periods 5 - 6, Room A401.
[0090] Output parameters include: <|class|>Modern Management; <|teacher|>Teacher Z; <|week|>5; <|start|>5; <|end|>6; <|round|>1 - 17; <|room|>Room A401.
[0091] The loss functions include:
[0092] Totalloss = alpha × (class_celoss + teacher_celoss + round_celoss + room_celoss) + beta × (week_regloss + start_regloss + end_regloss);
[0093] Among them, class_celoss is the loss for identifying the course name, which is a common cross - entropy loss function. teacher_celoss is the loss function for identifying the teacher, round_celoss is the loss function for identifying the class weeks, room_celoss is the loss function for identifying the classroom, week_regloss is the loss function for identifying the day of the week, start_regloss is the loss function for identifying the start period, end_regloss is the loss function for identifying the end period. alpha is the influence factor controlling the weight of the cross - entropy loss, and beta is the influence factor controlling the weight of the regression loss. It should be noted that the above loss functions are all common cross - entropy loss functions.
[0094] For the table generation method provided in the embodiments of the present application, the execution subject can be a table generation device. In the embodiments of the present application, taking the table generation device executing the table generation method as an example, the table generation method provided in the embodiments of the present application is described.
[0095] Figure 8 Shows the second flow diagram of the table generation method provided in some embodiments of the present application, as Figure 8 shown, the table generation method includes:
[0096] Step 801, turn on the table recognition and reconstruction function;
[0097] Step 802, determine whether it is in the full - automatic mode. If the determination is yes, execute Step 803; if the determination is no, execute Step 805;
[0098] Step 803, identify and parse the content area to obtain content cells;
[0099] Step 804, identify and parse the table header area to obtain table header cells;
[0100] Step 805, determine whether it is in semi-automatic mode. If it is, execute Step 806; if not, execute Step 808;
[0101] Step 806, select and parse the content area to obtain content cells;
[0102] Step 807, select and parse the table header area to obtain table header cells;
[0103] Step 808, determine whether it is in manual mode. If it is, execute Step 809; if not, end;
[0104] Step 809, after clicking on the table coordinates, select and parse the content area to obtain content cells;
[0105] Step 810, splice the content cells and the table header cells to form a first table;
[0106] Step 811, identify the first target content of the first table;
[0107] Step 812, determine whether the manual operation is completed. If not, return to execute Step 809; if so, execute Step 813;
[0108] Step 813, replace the first target content with the second target content.
[0109] In some embodiments of the present application, a table generation device is provided. Figure 9 The schematic block diagram of the table generation device provided in some embodiments of the present application is shown. As Figure 9 shown, the table generation device 900 includes:
[0110] An acquisition module 902, configured to acquire a first image including a table, where the first image includes a table header area and a content area;
[0111] A generation module 904, configured to generate at least one content cell based on text content information, where the text content information is content information determined by performing image recognition on the content area, and the semantic information in the same content cell is mutually related;
[0112] A splicing module 906, configured to splice at least one content cell and a table header cell based on the cell position information of the content cell to generate a first table, where the table header cell corresponds to the table header area, and the cell position information is position information determined according to the content position information of the text content information.
[0113] In the embodiments of the present application, the electronic device can identify the content area in the first image, generate content cells within the content area based on the semantic information of the recognized text content information, and the content included in each content cell has semantic relevance. Moreover, based on the content position information of the text content information, the cell position information of each content cell is determined, so that the content cells can be spliced with the header cells based on the cell position information, ensuring smooth semantics inside the content cells, accurate division of the content cells, and improving the matching accuracy of the display position of the content cells and the position of the header cells, thereby improving the accuracy of generating a table based on the first image.
[0114] In some embodiments of the present application, the table generation device 900 further includes:
[0115] A first recognition module, configured to recognize the table text information and text position information in the first image, and the table text information and the text position information are in one-to-one correspondence;
[0116] A first determination module, configured to determine the text content information and header content information in the table text information, and determine the content position information and header position information in the text position information;
[0117] The generation module 904 is further configured to generate header cells based on the header content information and the header position information.
[0118] In the embodiments of the present application, when the user needs to perform table reconstruction on the table content in the entire first image, the user can select the fully automatic recognition and reconstruction table function. At this time, the electronic device recognizes all the table text information in the first image and the corresponding text position information, and further screens the text content information and header content information in the table text information to facilitate subsequent automatic construction of content cells and header cells, and the operation of splicing the content cells and header cells based on the position relationship.
[0119] In some embodiments of the present application, the table generation device 900 further includes:
[0120] A second recognition module, configured to recognize the text content information and content position information in the content area in response to a first selection input for at least part of the content area, where the text content information is at least part of the information selected by the first selection input within the content area;
[0121] The second recognition module is further configured to recognize the header content information and header position information in the header area in response to a second selection input for at least part of the header area, where the header content information is at least part of the information selected by the second selection input within the header area, and the header position information is the position of the header content information in the first image;
[0122] The generating module 904 is further configured to generate header cells based on the header content information and the header position information.
[0123] In the embodiments of the present application, when the user needs to perform table reconstruction on the table content in at least part of the first image, the user can select the semi-automatic recognition and reconstruction table function. The user selects the text content information and the content position information in at least part of the content area to be recognized through the first selection input, and selects at least part of the header content information and the header position information to be recognized through the second selection input, facilitating subsequent automatic construction of content cells and header cells based on the selected text content information and header content information, and the operation of splicing the content cells and the header cells based on the positional relationship.
[0124] In some embodiments of the present application, the obtaining module 902 is further configured to obtain an expression template library corresponding to the header content information;
[0125] The table generating device 900 further includes:
[0126] The searching module is configured to sequentially extract the header content sub-information in the header content information along a first preset direction, and search for the header content sub-information in the expression template library;
[0127] The second determining module is further configured to, when the header content sub-information is found in the expression template library, determine a plurality of header position sub-information corresponding to the plurality of header content sub-information according to the first preset direction and the header position information;
[0128] The generating module 904 is further configured to generate header cells according to the plurality of header content sub-information and the plurality of header position sub-information.
[0129] In the embodiments of the present application, by performing recognition and analysis on the header content information, the extension direction of each header content sub-information in the header content information can be determined, so as to determine the corresponding header position sub-information, and the header cells can be constructed based on the header position sub-information, ensuring that the layout and content of the header cells match the header area in the first image.
[0130] In some embodiments of the present application, the splicing module 906 is further configured to splice at least one content cell and the header cell based on the cell position information and the plurality of header position sub-information to obtain a first table.
[0131] In the embodiments of the present application, the positional relationship between the content cell and the header cell is determined through the cell position information and multiple header position sub-information, and based on the positional relationship between the content cell and the header cell, the content cell and the header cell are spliced to obtain the first table, ensuring that the positional relationship between the content cell and the header cell in the first table matches the positional relationship between the content area and the header area in the first image, and improving the accuracy of the reconstructed table.
[0132] In some embodiments of the present application, the number of text content information is multiple, and the content position information corresponds to the text content information one by one;
[0133] The table generation device 900 further includes:
[0134] A third recognition module, configured to recognize the semantic relevance between multiple text content information;
[0135] A division module, configured to divide multiple text content information into at least one content cell based on the semantic relevance;
[0136] The generation module 904 is further configured to generate at least one content cell according to at least one content cell.
[0137] In the embodiments of the present application, when multiple text content information is recognized, it is necessary to recognize the semantic relevance between multiple text content information, regard the text content information with semantic relevance as a group of text information, and generate content cells according to a group of text information with semantic relevance, ensuring that the text content information in each content cell has semantic relevance and improving the accuracy of recognizing the content area in the first image.
[0138] In some embodiments of the present application, the acquisition module 902 is further configured to acquire the edge pixel coordinates of the content cell in the content area;
[0139] The acquisition module 902 is further configured to move the edge pixel coordinates along a second preset direction and acquire the pixel difference between the positions of the edge pixel coordinates before and after the movement;
[0140] The table generation device 900 further includes:
[0141] A third determination module, configured to determine the edge pixel coordinates at the current position as the cell position information when the pixel difference is greater than the difference threshold.
[0142] In the embodiments of the present application, after determining the content cell, the actual position of the content cell is calibrated based on the edge pixel coordinates of the content cell, so as to obtain the cell position information corresponding to the content cell, further improving the matching degree between the generated content cell and the content area in the first image.
[0143] In some embodiments of the present application, the table generation device 900 further includes:
[0144] An extraction module, configured to extract the text content in the content cell;
[0145] The generation module 904 is further configured to typeset the first target content according to a target format to generate a second target content;
[0146] The generation module 904 is further configured to replace the content cells in the first table with the second target content to generate a second table.
[0147] In the embodiments of the present application, the text content in the content cells of the first table is the same as the text content in the first image. By extracting the first target content in the text content, typesetting the first target content according to the target format to obtain the second target content, and then replacing the first target content with the second target content, the content reconstruction of the content cells is completed, so that the text in each content cell is a standardized data structure, which is convenient for subsequent table processing such as merging cells of the second table.
[0148] The table generation device in the embodiments of the present application may be an electronic device or a component in an electronic device, such as an integrated circuit or a chip. The electronic device may be a terminal or other devices other than the terminal. Exemplarily, the electronic device may be a mobile phone, a tablet computer, a laptop computer, a handheld computer, a vehicle-mounted electronic device, a Mobile Internet Device (MID), an augmented reality (AR) / virtual reality (VR) device, a robot, a wearable device, an ultra-mobile personal computer (UMPC), a netbook, or a personal digital assistant (PDA), etc., and may also be a server, a Network Attached Storage (NAS), a personal computer (PC), a television (TV), a teller machine, or a self-service machine, etc. The embodiments of the present application do not make specific limitations.
[0149] The table generation device in the embodiments of the present application may be a device with an operating system. The operating system may be an Android operating system, an iOS operating system, or other possible operating systems. The embodiments of the present application do not make specific limitations.
[0150] The table generation device provided by the embodiments of the present application can implement each process implemented by the above method embodiments. To avoid repetition, it will not be elaborated here.
[0151] Optionally, the embodiments of the present application further provide an electronic device, Figure 10 which shows a structural block diagram of the electronic device in some embodiments of the present application, as Figure 10 shown, the electronic device 1000 includes a processor 1002, a memory 1004, a program or instruction stored on the memory 1004 and executable on the processor 1002. When the program or instruction is executed by the processor 1002, it implements each process of the above table generation method embodiment and can achieve the same technical effect. To avoid repetition, it will not be elaborated here.
[0152] It should be noted that the electronic devices in the embodiments of the present application include the above-mentioned mobile electronic devices and non-mobile electronic devices.
[0153] Figure 11 It is a schematic diagram of the hardware structure of an electronic device for implementing the embodiments of the present application.
[0154] The electronic device 1100 includes but is not limited to: a radio frequency unit 1101, a network module 1102, an audio output unit 1103, an input unit 1104, a sensor 1105, a display unit 1106, a user input unit 1107, an interface unit 1108, a memory 1109, and a processor 1110 and other components.
[0155] Those skilled in the art can understand that the electronic device 1100 may further include a power supply (such as a battery) for supplying power to each component. The power supply can be logically connected to the processor 1110 through a power management system, so as to implement functions such as management of charging, discharging, and power consumption management through the power management system. Figure 11 The structure of the electronic device shown in does not constitute a limitation on the electronic device. The electronic device may include more or fewer components than shown, or combine some components, or have different component arrangements, which will not be elaborated here.
[0156] Among them, the processor 1110 is used to obtain a first image including a table, and the first image includes a header area and a content area;
[0157] The processor 1110 is used to generate at least one content cell based on the text content information, where the text content information is the content information determined by performing image recognition on the content area, and the semantic information in the same content cell is mutually related;
[0158] A processor 1110 is configured to splice at least one content cell with a header cell based on the cell position information of the content cell to generate a first table, where the header cell corresponds to a header area, and the cell position information is position information determined according to the content position information of the text content information.
[0159] In the embodiments of the present application, the electronic device can identify the content area in the first image, generate content cells in the content area based on the semantic information of the recognized text content information, each content cell includes content with semantic relevance, and also determine the cell position information of each content cell based on the content position information of the text content information, so that it can be spliced with the header cell based on the cell position information, ensuring smooth semantics inside the content cell, accurate division of the content cell, and improving the matching accuracy of the display position of the content cell and the position of the header cell, thereby improving the accuracy of generating a table based on the first image.
[0160] In some embodiments of the present application, the processor 1110 is configured to identify the table text information and text position information in the first image, and the table text information and the text position information are in one-to-one correspondence;
[0161] The processor 1110 is configured to determine the text content information and header content information in the table text information, and the content position information and header position information in the text position information;
[0162] The processor 1110 is configured to generate a header cell based on the header content information and the header position information.
[0163] In the embodiments of the present application, when the user needs to reconstruct the table content in the entire first image, the user can select the fully automatic recognition and reconstruction table function. At this time, the electronic device recognizes all the table text information in the first image and the corresponding text position information, and further screens the text content information and header content information in the table text information to facilitate subsequent automatic construction of content cells and header cells, and the operation of splicing the content cells and header cells based on the position relationship.
[0164] In some embodiments of the present application, the processor 1110 is configured to, in response to a first selection input for at least part of the content area, identify the text content information and content position information in the content area, where the text content information is at least part of the information selected by the first selection input in the content area;
[0165] A processor 1110, configured to identify header content information and header position information in a header area in response to a second selection input for at least a part of the header area, where the header content information is at least part of the information selected by the second selection input within the header area, and the header position information is the position of the header content information in the first image;
[0166] A processor 1110, configured to generate header cells based on the header content information and the header position information.
[0167] In the embodiments of the present application, when a user needs to perform table reconstruction on the table content in at least a part of the first image, the user can select the semi-automatic recognition and reconstruction table function. The user selects the text content information and content position information in at least a part of the content area to be recognized through a first selection input, and selects at least a part of the header content information and header position information to be recognized through a second selection input, which facilitates subsequent automatic construction of content cells and header cells according to the selected text content information and header content information, and the operation of splicing the content cells and header cells based on the positional relationship.
[0168] In some embodiments of the present application, a processor 1110 is configured to obtain an expression template library corresponding to the header content information;
[0169] A processor 1110 is configured to sequentially extract header content sub-information in the header content information along a first preset direction, and search for the header content sub-information in the expression template library;
[0170] A processor 1110 is configured to, when the header content sub-information is searched in the expression template library, determine multiple header position sub-information corresponding to the multiple header content sub-information according to the first preset direction and the header position information;
[0171] A processor 1110 is configured to generate header cells according to the multiple header content sub-information and the multiple header position sub-information.
[0172] In the embodiments of the present application, by performing identification and analysis on the header content information, the extension direction of each header content sub-information in the header content information can be determined, so as to determine the corresponding header position sub-information, and based on the header position sub-information, header cells can be constructed, ensuring that the layout and content of the header cells match the header area in the first image.
[0173] In some embodiments of the present application, a processor 1110 is configured to splice at least one content cell and the header cell based on the cell position information and the multiple header position sub-information to obtain a first table.
[0174] In the embodiments of the present application, the positional relationship between the content cell and the header cell is determined through the cell position information and multiple header position sub-information, and based on the positional relationship between the content cell and the header cell, the content cell and the header cell are spliced to obtain the first table, ensuring that the positional relationship between the content cell and the header cell in the first table matches the positional relationship between the content area and the header area in the first image, and improving the accuracy of reconstructing the table.
[0175] In some embodiments of the present application, the number of text content information is multiple, and the content position information corresponds to the text content information one by one;
[0176] The processor 1110 is configured to identify the semantic relevance between multiple text content information;
[0177] The processor 1110 is configured to divide multiple text content information into at least one content cell based on the semantic relevance;
[0178] The processor 1110 is configured to generate at least one content cell according to at least one content cell.
[0179] In the embodiments of the present application, when multiple text content information is recognized, it is necessary to recognize the semantic relevance between the multiple text content information, regard the text content information with semantic relevance as a group of text information, and generate a content cell according to a group of text information with semantic relevance, ensuring that the text content information in each content cell has semantic relevance and improving the accuracy of recognizing the content area in the first image.
[0180] In some embodiments of the present application, the processor 1110 is configured to obtain the edge pixel coordinates of the content cell in the content area;
[0181] The processor 1110 is configured to move the edge pixel coordinates along the second preset direction and obtain the pixel difference between the positions of the edge pixel coordinates before and after the movement;
[0182] The processor 1110 is configured to determine the edge pixel coordinates at the current position as the cell position information when the pixel difference is greater than the difference threshold.
[0183] In the embodiments of the present application, after the content cell is determined, the actual position of the content cell is calibrated based on the edge pixel coordinates of the content cell, so as to obtain the cell position information corresponding to the content cell, further improving the matching degree between the generated content cell and the content area in the first image.
[0184] In some embodiments of the present application, the processor 1110 is configured to extract the text content in the content cell;
[0185] A processor 1110 for inputting text content into a content recognition model and extracting first target content from the text content;
[0186] A processor 1110 for typesetting the first target content according to a target format to generate second target content;
[0187] A processor 1110 for replacing the content cells in the first table with the second target content to generate a second table.
[0188] In the embodiment of the present application, the text content in the content cells of the first table is the same as the text content in the first image. By extracting the first target content from the text content, typesetting the first target content according to the target format to obtain the second target content, and then replacing the first target content with the second target content, the content reconstruction of the content cells is completed, so that the text in each content cell is a standardized data structure, which is convenient for subsequent table processing such as merging cells of the second table.
[0189] It should be understood that in the embodiment of the present application, the input unit 1104 may include a Graphics Processing Unit (GPU) 11041 and a microphone 11042. The graphics processor 11041 processes the image data of static pictures or action files obtained by an image capture device (such as a camera) in the action file capture mode or the image capture mode. The display unit 1106 may include a display panel 11061, and the display panel 11061 may be configured in the form of a liquid crystal display, an organic light emitting diode, etc. The user input unit 1107 includes at least one of a touch panel 11071 and other input devices 11072. The touch panel 11071 is also called a touch screen. The touch panel 11071 may include two parts: a touch detection device and a touch controller. The other input devices 11072 may include, but are not limited to, a physical keyboard, function keys (such as volume control keys, switch keys, etc.), a trackball, a mouse, and a joystick, which will not be elaborated here.
[0190] The memory 1109 can be used to store software programs and various data. The memory 1109 may mainly include a first storage area for storing programs or instructions and a second storage area for storing data. Among them, the first storage area can store an operating system, application programs or instructions required for at least one function (such as a sound playback function, an image playback function, etc.). In addition, the memory 1109 can include volatile memory or non-volatile memory, or the memory 1109 can include both volatile and non-volatile memory. Among them, the non-volatile memory can be a Read-Only Memory (ROM), a Programmable ROM (PROM), an Erasable PROM (EPROM), an Electrically EPROM (EEPROM), or a flash memory. The volatile memory can be a Random Access Memory (RAM), a Static RAM (SRAM), a Dynamic RAM (DRAM), a Synchronous DRAM (SDRAM), a Double Data Rate SDRAM (DDR SDRAM), an Enhanced SDRAM (ESDRAM), a Synch link DRAM (SLDRAM), and a Direct Rambus RAM (DRRAM). The memory 1109 in the embodiments of the present application includes but is not limited to these and any other suitable types of memory.
[0191] The processor 1110 may include one or more processing units; optionally, the processor 1110 integrates an application processor and a modem processor. Among them, the application processor mainly processes operations related to the operating system, user interface, and application programs, etc., and the modem processor mainly processes wireless communication signals, such as a baseband processor. It can be understood that the above modem processor may not be integrated into the processor 1110.
[0192] The embodiments of the present application also provide a readable storage medium. A program or instructions are stored on the readable storage medium. When the program or instructions are executed by a processor, each process of the above method embodiments is implemented, and the same technical effects can be achieved. To avoid repetition, it will not be elaborated here.
[0193] Among them, the processor is the processor in the electronic device in the above embodiments. The readable storage medium includes computer-readable storage media, such as computer read-only memory ROM, random access memory RAM, magnetic disks, or optical discs, etc.
[0194] Another embodiment of the present application provides a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor. The processor is used to run programs or instructions to implement each process of the above embodiment of the table generation method, and can achieve the same technical effects. To avoid repetition, it will not be elaborated here.
[0195] It should be understood that the chip mentioned in the embodiments of the present application can also be referred to as a system-on-chip, system chip, chip system, or system-on-chip, etc.
[0196] The embodiments of the present application provide a computer program product. The program product is stored in a storage medium. The program product is executed by at least one processor to implement each process of the above embodiment of the table generation method, and can achieve the same technical effects. To avoid repetition, it will not be elaborated here.
[0197] It should be noted that in this article, the term "including", "comprising", or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, device, article, or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or also includes elements inherent to such process, device, article, or device. Without more limitations, the element defined by the statement "including one..." does not exclude the existence of another identical element in the process, device, article, or device including that element. In addition, it should be pointed out that the devices and the scope of the devices in the embodiments of the present application are not limited to performing functions in the order shown or discussed, and may also include performing functions in a substantially simultaneous manner or in the reverse order according to the functions involved. For example, the described devices may be executed in a different order than described, and various steps may be added, omitted, or combined. Additionally, the features described with reference to certain examples may be combined in other examples.
[0198] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-described embodiment devices can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases, the former is a better implementation method. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, can be embodied in the form of a computer software product. The computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disc), and includes several instructions for causing a terminal (which can be a mobile phone, computer, server, or network device, etc.) to execute the devices of the various embodiments of the present application.
[0199] The embodiments of the present application have been described above in conjunction with the accompanying drawings. However, the present application is not limited to the above specific embodiments. The above specific embodiments are merely illustrative rather than restrictive. Those of ordinary skill in the art, under the inspiration of the present application and without departing from the purpose of the present application and the scope protected by the claims, can also make many forms, all of which fall within the protection scope of the present application.
Claims
1. A method for generating a table, characterized in that, The table generation method includes: Obtain a first image including a table, where the first image includes a header area and a content area; Generate at least one content cell based on the text content information, where the text content information is the content information determined by performing image recognition on the content area, and the semantic information in the same content cell is mutually related; Splice the at least one content cell with a header cell based on the cell position information of the content cell to generate a first table, where the header cell corresponds to the header area, and the cell position information is the position information determined according to the content position information of the text content information.
2. The table generation method according to claim 1, wherein After obtaining the first image including the table, the table generation method further includes: Identify the table text information and text position information in the first image, where the table text information and the text position information are in one-to-one correspondence; Determine the text content information and header content information in the table text information, and determine the content position information and header position information in the text position information; Generate the header cell based on the header content information and the header position information.
3. The table generation method according to claim 1, wherein After obtaining the first image including the table, the table generation method further includes: In response to a first selection input for at least part of the content area, identify the text content information and the content position information in the content area, where the text content information is at least part of the information selected by the first selection input within the content area; In response to a second selection input for at least part of the header area, identify the header content information and header position information in the header area, where the header content information is at least part of the information selected by the second selection input within the header area, and the header position information is the position of the header content information in the first image; Generate the header cell based on the header content information and the header position information.
4. The table generation method according to claim 2 or 3, characterized in that The header content information includes multiple header content sub-information. Generating the header cell based on the header content information and the header position information includes: Obtain an expression template library corresponding to the header content information; Sequentially extract the header content sub-information in the header content information along a first preset direction, and search for the header content sub-information in the expression template library; When the header content sub-information is found in the expression template library, determine multiple header position sub-information corresponding to the multiple header content sub-information according to the first preset direction and the header position information; Generate the header cell according to the multiple header content sub-information and the multiple header position sub-information.
5. The table generation method according to claim 4, wherein Splicing the at least one content cell with the header cell based on the cell position information of the content cell to generate a first table includes: Splice the at least one content cell with the header cell based on the cell position information and the multiple header position sub-information to obtain the first table.
6. A table generation device, characterized in that, The table generation device includes: An acquisition module, configured to acquire a first image including a table, where the first image includes a header area and a content area; A generation module, configured to generate at least one content cell based on text content information, where the text content information is content information determined by performing image recognition on the content area, and the semantic information in the same content cell is mutually related; A splicing module, configured to splice the at least one content cell and a header cell based on the cell position information of the content cell to generate a first table, where the header cell corresponds to the header area, and the cell position information is position information determined according to the content position information of the text content information.
7. The table generation device according to claim 6, wherein It further includes: A first recognition module, configured to recognize table text information and text position information in the first image, where the table text information and the text position information are in one-to-one correspondence; A first determination module, configured to determine the text content information and header content information in the table text information, and determine the content position information and header position information in the text position information; The generation module is further configured to generate the header cell based on the header content information and the header position information.
8. The table generating device according to claim 6, wherein It further includes: A second recognition module, configured to, in response to a first selection input for at least part of the content area, recognize the text content information and the content position information in the content area, where the text content information is at least part of the information selected by the first selection input within the content area; The second recognition module is further configured to, in response to a second selection input for at least part of the header area, recognize the header content information and the header position information in the header area, where the header content information is at least part of the information selected by the second selection input within the header area, and the header position information is the position of the header content information in the first image; The generation module is further configured to generate the header cell based on the header content information and the header position information.
9. The table generating device according to claim 7 or 8, characterized in that, The header content information includes multiple header content sub-information; The acquisition module is further configured to acquire an expression template library corresponding to the header content information; The table generation device further includes: A search module, configured to sequentially extract the header content sub-information in the header content information in a first preset direction, and search for the header content sub-information in the expression template library; A second determination module, configured to, when the header content sub-information is found in the expression template library, determine multiple header position sub-information corresponding to the multiple header content sub-information according to the first preset direction and the header position information; The generation module is further configured to generate the header cell according to the multiple header content sub-information and the multiple header position sub-information.
10. The table generation device according to claim 9, wherein The splicing module is further configured to splice the at least one content cell and the header cell based on the cell position information and the multiple header position sub-information to obtain the first table.