Image processing method, apparatus, device, storage medium and program product
By obtaining the positional information of the table and text objects, generating the target outline and completing the outer edge of the table, the problem of incomplete extraction caused by missing table edges is solved, and the accuracy and completeness of table information extraction are improved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- INDUSTRIAL AND COMMERCIAL BANK OF CHINA
- Filing Date
- 2026-02-05
- Publication Date
- 2026-06-16
AI Technical Summary
In existing technologies, when the table edges in a table image are missing, the table structure in the image cannot be extracted accurately and completely, affecting the accuracy and completeness of subsequent information extraction.
By obtaining the positional information of the table and text objects in the image, a target outline is generated. Based on this outline, the outer edge of the table is completed, and a new image with a complete outer edge is generated.
This improves the accuracy and completeness of table structure extraction, ensuring the efficiency and accuracy of subsequent information extraction.
Smart Images

Figure CN122223734A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of financial technology or other related fields, and in particular to an image processing method, apparatus, device, storage medium and program product. Background Technology
[0002] Currently, in information processing scenarios for financial business, it is often necessary to perform Optical Character Recognition (OCR) on table images taken and uploaded by users to identify the information in the table, and then extract and process the information in the table according to the table structure.
[0003] However, in existing solutions, when the table in the image has missing borders, it is impossible to accurately and completely extract the table structure from the image, which in turn affects the accuracy and completeness of subsequent extraction of information from the table. Summary of the Invention
[0004] This application provides an image processing method, apparatus, device, storage medium, and program product to solve the technical problem of being unable to accurately and completely extract table structures from images.
[0005] In a first aspect, this application provides an image processing method, comprising:
[0006] A first image is acquired, the first image containing a first table, the first table being missing at least one outer edge of the target table; content recognition is performed on the first image to obtain first position information of the first table and second position information of the text object in the first image; based on the first position information and the second position information, a target contour line is obtained, the target contour line being used to indicate the minimum contour that accommodates the text object and the first table; based on the target contour line, a second image is generated, the second image containing the first table and the outer edge of the target table.
[0007] Secondly, this application provides an image processing apparatus, comprising:
[0008] The acquisition module is used to acquire a first image, the first image containing a first table, the first table being missing at least one outer edge of the target table;
[0009] The recognition module is used to perform content recognition on the first image to obtain the first position information of the first table and the second position information of the text objects in the first image.
[0010] The processing module is configured to obtain a target outline based on the first position information and the second position information, wherein the target outline is used to indicate the minimum outline that accommodates the text object and the first table.
[0011] A generation module is used to generate a second image based on the target outline, the second image containing the first table and the outer edge of the target table.
[0012] Thirdly, this application provides an electronic device, including: a processor and a memory communicatively connected to the processor; the memory stores computer-executable instructions; the processor executes the computer-executable instructions stored in the memory to implement the image processing method provided in any of the implementations of the first aspect above.
[0013] Fourthly, this application provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, are used to implement the image processing method provided in any of the implementations of the first aspect above.
[0014] Fifthly, this application provides a computer program product, including a computer program that, when executed by a processor, implements the image processing method provided by any of the implementations of the first aspect above.
[0015] The image processing method, apparatus, device, storage medium, and program product provided in this application acquire a first image containing a first table, the first table being missing at least one outer edge of a target table; perform content recognition on the first image to obtain first position information of the first table and second position information of text objects in the first image; obtain a target contour line based on the first and second position information, the target contour line indicating the minimum contour to accommodate the text objects and the first table; and generate a second image based on the target contour line, the second image containing the first table and the outer edge of the target table. By measuring the first image, identifying the first table missing at least one outer edge of a target table and the text objects in the first image, and then measuring the first and second position information corresponding to the first table and the text objects respectively, and determining the minimum contour to indicate the minimum contour to accommodate the text objects and the first table, i.e., the target contour line, based on the first and second position information, the outer edge of the target table in the first image is completed based on the target contour line to generate the second image, thereby overcoming the problem that the table structure cannot be accurately extracted due to the incomplete table structure in the table image, and improving the accuracy and completeness of extracting information within the table. Attached Figure Description
[0016] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0017] Figure 1 This is a schematic diagram of an application scenario provided by an embodiment of this application;
[0018] Figure 2 A flowchart illustrating the image processing method provided in the embodiments of this application. Figure 1 ;
[0019] Figure 3 This is a flowchart illustrating the steps for generating the fourth image after step S101.
[0020] Figure 4 This is a schematic diagram illustrating a process for generating a second image, as provided in an embodiment of this application.
[0021] Figure 5 A flowchart illustrating the image processing method provided in the embodiments of this application. Figure 2 ;
[0022] Figure 6 This is a structural block diagram of the image processing apparatus provided in the embodiments of this application;
[0023] Figure 7 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application;
[0024] Figure 8 This is a schematic diagram of the hardware structure of an electronic device provided in an embodiment of this application.
[0025] The accompanying drawings illustrate specific embodiments of this application, which will be described in more detail below. These drawings and descriptions are not intended to limit the scope of the concept in any way, but rather to illustrate the concept of this application to those skilled in the art through reference to particular embodiments. Detailed Implementation
[0026] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims.
[0027] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, storage, use, processing, transmission, provision, disclosure, and application of the relevant data all comply with the relevant laws, regulations, and standards of the relevant countries and regions, have taken necessary confidentiality measures, do not violate public order and good morals, and provide corresponding operation access points for users to choose to authorize or refuse.
[0028] It should be noted that the image processing method, electronic device, storage medium and program product provided in this application can be used in the field of fintech, or in any field other than fintech. The application field of the image processing method, electronic device, storage medium and program product in this application is not limited.
[0029] The specific application scenario of this application is the extraction of information from table images. Figure 1 This is a schematic diagram of an application scenario provided in an embodiment of this application, such as... Figure 1 As shown, for example, paper forms can be converted into table images by taking pictures or scanning with a device (such as a mobile phone). The information in the table images can then be identified and extracted, including extracting the table structure (such as data columns for names, phone numbers, etc.) and the corresponding information within the table structure. After that, the information can be further processed, such as storing it in a database or processing business transactions.
[0030] For the aforementioned application scenarios, existing solutions often fail to accurately and completely extract the table structure from images when the table edges are missing. For example... Figure 1 The missing top edge (indicated by a dotted line) of the table will cause the position of text such as names and phone numbers in the table to change, which will affect the accuracy and completeness of subsequent extraction of information from the table.
[0031] The image processing method, electronic device, storage medium, and program product provided in this application are intended to solve the above-mentioned technical problems of the prior art.
[0032] The technical solution of this application and how the technical solution of this application solves the above-mentioned technical problems are described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments. The embodiments of this application will now be described with reference to the accompanying drawings.
[0033] refer to Figure 2 , Figure 2A flowchart illustrating the image processing method provided in the embodiments of this application. Figure 1 The method of this embodiment can be applied to a terminal device or a server. In one possible implementation, for a terminal device executing the method provided in this embodiment, the terminal device can implement the image processing method provided in this embodiment by executing program code deployed locally and / or externally. In another possible implementation, a server can be used to deploy functional services implemented based on the image processing method provided in this embodiment, and the terminal device can implement the image processing method provided in this embodiment by accessing the server and calling the corresponding functional services. For example, the image processing method provided in this embodiment includes:
[0034] Step S101: Obtain a first image containing a first table, the first table being missing at least one outer edge of the target table.
[0035] For example, refer to Figure 1 The illustrated application scenario uses a terminal device (such as a mobile phone) as an example. First, the terminal device obtains a first image by taking a picture or downloading it from an external source. The first image contains a table missing at least one outer edge, i.e., the first table. The missing outer edge of the first table is the outer edge of the target table, such as the left outer edge or the top outer edge of the table. The missing outer edge of the target table in the first table may be lost during the shooting or cropping process due to unreasonable shooting angles, cropping positions, or other reasons; this will not be elaborated upon further.
[0036] Furthermore, in one possible implementation, after step S101, the method may further include: performing grayscale processing on the first image to generate a third image.
[0037] For example, after obtaining the first image, the first image is converted to grayscale to remove color interference, thereby improving the contrast of the table lines in the first image and thus improving the accuracy of subsequent extraction of line segment objects.
[0038] In another possible implementation, after step S101, the method may further include: obtaining the tilt angle of the first table in the first image relative to the plane where the first image is located, and calibrating the first table based on the tilt angle to generate a fourth image, wherein the first table in the fourth image is located on the plane where the fourth image is located.
[0039] For example, during the scanning or capturing of the first image, the table in the first image may be tilted due to factors such as device angle and shooting environment. A tilted table image can lead to inaccurate subsequent table line extraction results, affecting not only the overall structure recognition of the table but also significantly impacting the accuracy of text recognition. Therefore, before extracting table lines, the first image needs to be tilted to align the table borders with the image coordinate axes. Specifically, this involves moving the pixels in the first image to reconstruct the first table, placing it on the plane of the fourth image, thereby improving the accuracy of text recognition. The tilt angle can be determined based on the angle between the two parallel outer edges of the first table (e.g., the left and right outer edges). The specific implementation method for reconstructing the image based on the tilt angle to achieve tilt correction will not be elaborated here.
[0040] It is understandable that the two optimization schemes for the first image described above can be executed individually, and subsequent processing steps can be performed based on the resulting optimized image (the third or fourth image). Alternatively, they can be executed sequentially and simultaneously, for example, first converting the first image to grayscale, and then correcting its tilt angle. Specifically, after step S101, a step to generate the fourth image is also included, such as... Figure 3 As shown, it includes:
[0041] Step S101-1: Perform grayscale processing on the first image to generate the third image.
[0042] Step S101-2: Obtain the tilt angle of the first table in the third image relative to the plane where the third image is located.
[0043] Step S101-3: Calibrate the first table based on the tilt angle to generate a fourth image, wherein the first table in the fourth image is located in the plane of the fourth image.
[0044] Accordingly, after performing the above steps, the specific implementation of step S102 includes:
[0045] Step S102A: Perform content recognition on the fourth image to obtain the first position information of the first table and the second position information of the text objects in the fourth image.
[0046] In this embodiment, by first converting the first image to grayscale and then performing tilt correction, the image clarity can be improved, color interference can be reduced, and the tilt correction effect can be improved, thereby improving the accuracy of subsequent extraction of the first table and text objects.
[0047] Step S102: Perform content recognition on the first image to obtain the first position information of the first table and the second position information of the text objects in the first image.
[0048] After obtaining the first image, content recognition is performed on the first image by calling an image processing model or tool to identify the first table and the text objects within it, and to determine the first position information corresponding to the first table and the second position information of the text objects. The first position information describes the position of the first table. In one possible implementation, it describes the position of all line segments constituting the first table; in another, it describes the position of each cell within the first table; and in yet another, it describes the position of the smallest outline that accommodates the first table. The second position information describes the position of the text objects within the first table. These text objects can be characters, symbols, numbers, or other similar identifiers. In one possible implementation, the second position information describes the position of each text object in the first image; in another possible implementation, it describes the position of the entire set of text objects within the first image.
[0049] Step S103: Based on the first position information and the second position information, obtain the target outline. The target outline is used to indicate the minimum outline that accommodates the text object and the first table.
[0050] Step S104: Generate a second image based on the target contour line. The second image contains the first table and the outer edge of the target table.
[0051] Next, by performing a union operation on the first and second positional information based on their positional relationship, the target outline used to indicate the minimum outline of the text object and the first table can be obtained. Then, the target outline is mapped back to the first image. By comparing the target outline with the outer edges of each table in the first image, the missing outer edges of the first table, i.e., the target table outer edges, can be determined. Finally, this target table outer edge is added to the first image to obtain an image of the first table with complete outer edges, i.e., the second image.
[0052] For example, the specific implementation method of step S104 includes:
[0053] Step S1041: Compare the difference line segments between the target outline and the outer edge of the first table in the first image to determine the outer edge of the target table;
[0054] Step S1042: At the location of the outer edge of the target table in the first image, add the outer edge of the target table to generate the second image.
[0055] Figure 4 This is a schematic diagram illustrating a process for generating a second image according to an embodiment of this application. The following is in conjunction with... Figure 4 To further explain the above process, such as... Figure 4 As shown, exemplarily, firstly, image P1 is grayscaled and tilt-corrected to obtain image P2. Then, the first table and the text objects within the first table are extracted from image P2, and the first position information corresponding to the first table and the second position information corresponding to the text objects are obtained. For example, as shown in the figure, the first position information is represented by a rectangle L1, and the second position information is represented by a rectangle L2. Next, the intersection of rectangles L1 and L2 is calculated to obtain the target contour line L3. Finally, this target contour line is mapped onto the previous image P1, or onto the grayscaled and tilt-corrected image P2, and aligned with the outer border of the first table. The difference line segment between the target contour line and the outer edge of the first table is determined, i.e., the outer edge of the target table. Then, the outer edge of the target table is drawn, which can complete the outer edge of the target table, resulting in a second image P3 containing the first table with a complete outer edge, thus improving the efficiency and accuracy of information extraction from the image. If there are missing line segment objects in the first image, the missing line segment objects can be detected by comparison, and the internal line segments corresponding to the missing line segment objects in the first table can be filled in. The specific implementation method will be described in detail in the following embodiments.
[0056] The image processing method provided in this embodiment acquires a first image containing a first table, which is missing at least one outer edge of the target table. Content recognition is performed on the first image to obtain first position information of the first table and second position information of text objects in the first image. Based on the first and second position information, a target contour line is obtained, which indicates the minimum contour to accommodate the text objects and the first table. Based on the target contour line, a second image is generated, containing the first table and the outer edge of the target table. By measuring the first image, the first table missing at least one outer edge of the target table and the text objects in the first image are identified. The first and second position information corresponding to the first table and the text objects are then measured, and based on the first and second position information, the minimum contour to indicate the text objects and the first table, i.e., the target contour line, is determined. Finally, the outer edge of the target table in the first image is completed based on the target contour line to generate the second image. This overcomes the problem of inaccurate table structure extraction due to incomplete table structure in the table image, improving the accuracy and completeness of information extraction from the table.
[0057] Figure 5 A flowchart illustrating the image processing method provided in the embodiments of this application. Figure 2 ,exist Figure 2 Based on the illustrated embodiment, steps S102-S104 are further refined, and exemplarily include:
[0058] Step S201: Obtain a first image containing a first table, the first table being missing at least one outer edge of the target table.
[0059] Step S202: Identify line segment objects in the first image and obtain the line segment coordinates corresponding to each line segment object, wherein the line segment coordinates are used to characterize the position of the endpoints of the line segment objects.
[0060] Step S203: Obtain the first position information of the first table based on the minimum endpoint coordinates and maximum endpoint coordinates of the line segment coordinates corresponding to each line segment object.
[0061] For example, after obtaining and preprocessing the first image, line segment recognition is performed on the first image to obtain line segment objects in the first image. Based on the position of the line segment objects in the first image, the line segment coordinates corresponding to the line segment objects are generated. The line segment coordinates are used to represent the position of the endpoints of the line segment objects. For example, after performing line segment recognition on the first image, multiple line segment objects are obtained, including line segment Ls_1. The line segment coordinates of line segment Ls_1 are [(xa_1, ya_1), (xb_1, yb_1)], where (xa_1, ya_1) represents the x-coordinate and y-coordinate of endpoint a of line segment Ls_1; (xb_1, yb_1) represents the x-coordinate and y-coordinate of endpoint b of line segment Ls_1. Similarly, the line segment coordinates of line segment Ls_2 are [(xa_2, ya_2), (xb_2, yb_2)], and so on. Next, the coordinates of the minimum and maximum endpoints of each line segment are calculated. The minimum endpoint is the endpoint with the smallest x-coordinate and y-coordinate, and this minimum endpoint is set as the bottom-left endpoint of the rectangle. The maximum endpoint is the endpoint with the largest x-coordinate and y-coordinate, and this maximum endpoint is set as the top-right endpoint of the rectangle. The rectangle formed by these minimum and maximum endpoints constitutes the first position information. In other words, the first position information includes the coordinates of the bottom-left and top-right endpoints of a rectangle (bounding box). Using this first position information, the outer contour of the first table can be depicted.
[0062] Step S204: Identify the text objects in the first image and obtain the alignment bounding box of the text objects. The alignment bounding box represents the fixed-size outline of the text objects.
[0063] Step S205: Obtain the second position information of the text object based on the alignment bounding box of the text object.
[0064] For example, on the other hand, the terminal device identifies text objects in the first image and represents them using an axis-aligned bounding box (AABB) data structure. The AABB represents the fixed-size outline of the text object. Then, for example, the maximum endpoint (top right corner) and minimum endpoint (bottom left corner) of the AABB are used as second positional information. That is, the second positional information also includes the coordinates of the bottom left corner and the top right corner of a rectangle (bounding box). Through this second positional information, the overall outline of all text objects in the first image can be depicted. The specific implementation of recognizing text and line segments in the image described above will not be elaborated further here.
[0065] In this embodiment, the first and second location information obtained through the above method enable precise positioning of the first table and the text in the first table, thereby improving the accuracy of the first table in the subsequently generated second image.
[0066] Step S206: Obtain the first bounding box based on the first position information. The first bounding box represents the outline of the first table.
[0067] Step S207: Based on the second position information, obtain the second bounding box, which represents the overall outline of all text objects in the first image.
[0068] Step S208: Obtain the target outline based on the union of the first bounding box and the second bounding box. The target outline is a rectangle.
[0069] For example, after obtaining the first location information and the second location information, a corresponding first bounding box and a second bounding box are generated according to the regions described by the first location information and the second location information, respectively. After taking the union of the first bounding box and the second bounding box, a rectangular region that can contain the first bounding box and the second bounding box is generated, and the outline of the rectangular region is determined as the target outline.
[0070] In this embodiment, through the above steps, a first bounding box for covering the outline of the first table and a second bounding box for covering the entirety of all text objects in the first image are obtained. Then, based on the first bounding box and the second bounding box, a rectangular area covering the first bounding box and the second bounding box is constructed, and the outline of the rectangular area is determined as the target outline line, thereby realizing the prediction of the outer outline of the first table, so that the predicted first table can cover all text and improve the integrity of the first table.
[0071] For example, after step S208, the method further includes:
[0072] Step S209: By comparing the line segment objects with the first table, at least one missing line segment object is identified from the line segment objects. The missing line segment object is a line segment object that is not fully displayed.
[0073] Step S210: Generate a second image based on the target contour line and the missing line segment object.
[0074] For example, in some possible cases, when the first image has low image clarity due to shooting or scanning, the line segments constituting the table may not be displayed, i.e., there are missing line segment objects, thus affecting the structural integrity of the table. To address this issue, in this embodiment, after obtaining multiple line segment objects constituting the first table through image recognition of the first image in the previous steps, the line segment coordinates of the line segment objects are aligned with the first table in the first image to easily identify the incomplete line segment objects, i.e., the missing line segment objects. Then, based on the target outline and the missing line segment objects, a second image is generated by drawing together on the first image. Through the steps of this embodiment, the internal structure of the table can be completed, solving the problem of missing lines in internal cells, thereby further improving the structural integrity of the table.
[0075] In a more specific embodiment, for example, multiple line segment objects, such as line segment objects Ls_1, Ls_2, Ls_3, and Ls_4, are obtained by identifying the first image. Then, based on the line segment coordinates of the line segment objects, the first image is aligned, and compared with the overall structure of the first table in the first image. It can be determined that line segment object Ls_1 is a missing line segment object. Subsequently, during the generation of the second image, in addition to completing the outer edge of the target table based on the target contour line, the line segments corresponding to line segment object Ls_1 in the first table are also completed to Ls_1r, thereby ensuring the integrity of the internal line segments of the first table and improving the table's completeness and accuracy. This process of identifying missing line segment objects can be executed using a pre-trained image processing model, or it can be implemented by determining whether the line segment length of each line segment object is consistent with the line segment length of its adjacent parallel line segment object; the specific configuration can be adjusted as needed.
[0076] Corresponding to the image processing method in the above embodiments, Figure 6 This is a structural block diagram of an image processing apparatus provided in an embodiment of this application. The method described in the above embodiments can be executed by this image processing apparatus, which can be implemented by software and / or hardware, and can be integrated into an electronic device with certain data processing capabilities. The electronic device may include, but is not limited to, mobile terminals with big data processing capabilities, as well as fixed terminals with big data processing capabilities such as desktop computers and supercomputers.
[0077] For ease of explanation, only the parts relevant to the embodiments of this application are shown. (Refer to...) Figure 6 The image processing device 3 includes:
[0078] The acquisition module 31 is used to acquire a first image, which contains a first table, and the first table is missing at least one outer edge of the target table.
[0079] The recognition module 32 is used to perform content recognition on the first image to obtain the first position information of the first table and the second position information of the text objects in the first image.
[0080] Processing module 33 is used to obtain a target outline based on the first position information and the second position information. The target outline is used to indicate the minimum outline that accommodates the text object and the first table.
[0081] The generation module 34 is used to generate a second image based on the target contour line. The second image contains the first table and the outer edge of the target table.
[0082] According to one or more embodiments of this application, after acquiring the first image, the recognition module 34 is further configured to: perform grayscale processing on the first image to generate a third image; when the recognition module 34 performs content recognition on the first image to obtain the first position information of the first table and the second position information of the text object in the first image, it is specifically configured to: perform content recognition on the third image to obtain the first position information of the first table and the second position information of the text object in the third image.
[0083] According to one or more embodiments of this application, after acquiring the first image, the recognition module 34 is further configured to: acquire the tilt angle of the first table in the first image relative to the plane where the first image is located; calibrate the first table based on the tilt angle to generate a fourth image, wherein the first table in the fourth image is located on the plane where the fourth image is located; when the recognition module 34 performs content recognition on the first image to obtain the first position information of the first table and the second position information of the text object in the first image, it is specifically configured to: perform content recognition on the fourth image to obtain the first position information of the first table and the second position information of the text object in the fourth image.
[0084] According to one or more embodiments of this application, the recognition module 32 is specifically configured to: recognize line segment objects in the first image and obtain the line segment coordinates corresponding to each line segment object, wherein the line segment coordinates are used to characterize the position of the endpoints of the line segment objects; obtain the first position information of the first table based on the minimum endpoint coordinates and the maximum endpoint coordinates of the line segment coordinates corresponding to each line segment object; recognize text objects in the first image and obtain the alignment bounding box of the text objects, wherein the alignment bounding box characterizes the fixed-size outline of the text objects; and obtain the second position information of the text objects based on the alignment bounding box of the text objects.
[0085] According to one or more embodiments of this application, the processing module 33 is specifically used for: obtaining a first bounding box based on first position information, the first bounding box representing the outline of a first table; obtaining a second bounding box based on second position information, the second bounding box representing the outline of the whole composed of all text objects in the first image; and obtaining a target outline line based on the union of the first bounding box and the second bounding box, the target outline line being a rectangle.
[0086] According to one or more embodiments of this application, the processing module 33 is specifically used to: determine at least one missing line segment object from the line segment objects by comparing the line segment objects and the first table, wherein the missing line segment object is a line segment object that is not fully displayed; the generation module 34 is specifically used to: generate a second image based on the target outline and the missing line segment object.
[0087] According to one or more embodiments of this application, the generation module 34 is specifically used to: compare the difference line segments between the target outline and the outer edge of the first table in the first image to determine the outer edge of the target table; and add the outer edge of the target table at the location of the outer edge of the first table in the first image to generate a second image.
[0088] The acquisition module 31, recognition module 32, processing module 33, and generation module are connected sequentially. The image processing device 3 provided in this embodiment can execute the technical solution of the above method embodiment, and its implementation principle and technical effect are similar, so it will not be described again here.
[0089] Figure 7 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application, such as... Figure 7 As shown, the electronic device 4 includes:
[0090] Processor 41, and memory 42 communicatively connected to processor 41;
[0091] Memory 42 stores instructions executed by the computer;
[0092] The processor 41 executes computer execution instructions stored in the memory 42 to achieve, for example, Figures 2-5 The image processing method in the illustrated embodiment.
[0093] Optionally, the processor 41 and the memory 42 are connected via a bus 43.
[0094] For relevant instructions, please refer to the corresponding text. Figures 2-5 The relevant descriptions and effects of the steps in the corresponding embodiments are understood, and will not be elaborated on here.
[0095] This application provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, are used to implement this application. Figures 2-5 The image processing method provided in any of the corresponding embodiments.
[0096] This application provides a computer program product, including a computer program, which, when executed by a processor, implements this application. Figures 2-5 The image processing method provided in any of the corresponding embodiments.
[0097] To implement the above embodiments, this application also provides an electronic device.
[0098] refer to Figure 8 The diagram illustrates a structural schematic of an electronic device 900 suitable for implementing embodiments of this application. The electronic device 900 can be a terminal device or a server. The terminal device can include, but is not limited to, mobile terminals such as mobile phones, laptops, digital radio receivers, personal digital assistants (PDAs), tablet computers, portable media players (PMPs), and in-vehicle terminals (e.g., in-vehicle navigation terminals), as well as fixed terminals such as digital TVs and desktop computers. Figure 8 The electronic device shown is merely an example and should not impose any limitation on the functionality and scope of use of the embodiments of this application.
[0099] like Figure 8 As shown, the electronic device 900 may include a processing unit (e.g., a central processing unit, a graphics processing unit, etc.) 901, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 902 or a program loaded from a storage device 908 into a random access memory (RAM) 903. The RAM 903 also stores various programs and data required for the operation of the electronic device 900. The processing unit 901, ROM 902, and RAM 903 are interconnected via a bus 904. An input / output (I / O) interface 905 is also connected to the bus 904.
[0100] Typically, the following devices can be connected to I / O interface 905: input devices 906 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 907 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 908 including, for example, magnetic tapes, hard disks, etc.; and communication devices 909. Communication device 909 allows electronic device 900 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 8 An electronic device 900 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively.
[0101] Specifically, according to embodiments of this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 909, or installed from a storage device 908, or installed from a ROM 902. When the computer program is executed by the processing device 901, it performs the functions defined in the methods of the embodiments of this application.
[0102] It should be noted that the computer-readable medium described above in this application can be a computer-readable signal medium or a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this application, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In this application, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.
[0103] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.
[0104] The aforementioned computer-readable medium carries one or more programs, which, when executed by the electronic device, cause the electronic device to perform the methods shown in the above embodiments.
[0105] Computer program code for performing the operations of this application can be written in one or more programming languages or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, and C++, and conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a Local Area Network (LAN) or a Wide Area Network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0106] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0107] The units or modules described in the embodiments of this application can be implemented in software or hardware. The names of the units or modules do not necessarily limit the specific unit itself.
[0108] The functions described above in this document can be performed at least in part by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), system-on-a-chip (SoCs), complex programmable logic devices (CPLDs), and so on.
[0109] In the context of this application, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. Machine-readable media can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0110] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, as some steps may be performed in other orders or simultaneously according to this application. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are all optional embodiments, and the actions and modules involved are not necessarily essential to this application.
[0111] It should be further noted that although the steps in the flowchart are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowchart may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the sub-steps or stages of other steps.
[0112] It should be understood that the above-described device embodiments are merely illustrative, and the device of this application can also be implemented in other ways. For example, the division of units / modules in the above embodiments is only a logical functional division, and there may be other division methods in actual implementation. For example, multiple units, modules, or components may be combined, or integrated into another system, or some features may be ignored or not executed.
[0113] Furthermore, unless otherwise specified, the functional units / modules in the various embodiments of this application can be integrated into one unit / module, or each unit / module can exist physically separately, or two or more units / modules can be integrated together. The integrated units / modules described above can be implemented in hardware or as software program modules.
[0114] When integrated units / modules are implemented in hardware, the hardware can be digital circuits, analog circuits, etc. The physical implementation of the hardware structure includes, but is not limited to, transistors, memristors, etc. Unless otherwise specified, the processor can be any suitable hardware processor, such as a CPU, GPU, FPGA, DSP, and ASIC, etc. Unless otherwise specified, the storage unit can be any suitable magnetic or magneto-optical storage medium, such as Resistive Random Access Memory (RRAM), Dynamic Random Access Memory (DRAM), Static Random Access Memory (SRAM), Enhanced Dynamic Random Access Memory (EDRAM), High-Bandwidth Memory (HBM), Hybrid Memory Cube (HMC), etc.
[0115] If the integrated unit / module is implemented as a software program module and sold or used as an independent product, it can be stored in a computer-readable storage device (CMD). Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a memory and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned memory includes various media capable of storing program code, such as a USB flash drive, read-only memory (ROM), random access memory (RAM), portable hard drive, magnetic disk, or optical disk.
[0116] In the above embodiments, the descriptions of each embodiment have their own emphasis. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments. The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as the combination of these technical features does not contradict each other, it should be considered within the scope of this specification.
[0117] Other embodiments of this application will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of this application that follow the general principles of this application and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this application are indicated by the following claims.
[0118] It should be understood that this application is not limited to the precise structure described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this application is limited only by the appended claims.
Claims
1. An image processing method, characterized in that, Obtain a first image, which contains a first table, the first table being missing at least one outer edge of the target table; Content recognition is performed on the first image to obtain the first position information of the first table and the second position information of the text objects in the first image. Based on the first position information and the second position information, a target outline is obtained, which is used to indicate the minimum outline that accommodates the text object and the first table. Based on the target outline, a second image is generated, which contains the first table and the outer edge of the target table.
2. The method according to claim 1, characterized in that, After acquiring the first image, the process also includes: The first image is converted to grayscale to generate the third image; The step of performing content recognition on the first image to obtain the first position information of the first table and the second position information of the text objects in the first image includes: Content recognition is performed on the third image to obtain the first position information of the first table and the second position information of the text objects in the third image.
3. The method according to claim 1, characterized in that, After acquiring the first image, the process also includes: Obtain the tilt angle of the first table in the first image relative to the plane on which the first image is located; The first table is calibrated based on the tilt angle to generate a fourth image, wherein the first table in the fourth image is located in the plane of the fourth image; The step of performing content recognition on the first image to obtain the first position information of the first table and the second position information of the text objects in the first image includes: Content recognition is performed on the fourth image to obtain the first position information of the first table and the second position information of the text objects in the fourth image.
4. The method according to claim 1, characterized in that, The step of performing content recognition on the first image to obtain the first position information of the first table and the second position information of the text objects in the first image includes: Identify line segment objects in the first image and obtain the line segment coordinates corresponding to each line segment object, wherein the line segment coordinates are used to characterize the position of the endpoints of the line segment objects; Based on the minimum endpoint coordinates and maximum endpoint coordinates of the line segment coordinates corresponding to each line segment object, the first position information of the first table is obtained; Identify text objects in the first image and obtain the alignment bounding box of the text objects, wherein the alignment bounding box represents the fixed-size outline of the text objects; The second position information of the text object is obtained based on the alignment bounding box of the text object.
5. The method according to claim 4, characterized in that, The step of obtaining the target contour line based on the first position information and the second position information includes: Based on the first location information, a first bounding box is obtained, and the first bounding box represents the outline of the first table; Based on the second position information, a second bounding box is obtained, which represents the overall outline of all text objects in the first image. The target outline is obtained by the union of the first bounding box and the second bounding box, and the target outline is a rectangle.
6. The method according to claim 4, characterized in that, Also includes: By comparing the line segment objects with the first table, at least one missing line segment object is determined from the line segment objects. The missing line segment object is a line segment object that is not fully displayed. The step of generating a second image based on the target contour line includes: A second image is generated based on the target contour line and the missing line segment object.
7. The method according to claim 1, characterized in that, The step of generating a second image based on the target contour line includes: The outer edge of the target table is determined by comparing the line segments that differ between the target outline and the outer edge of the first table in the first image. At the location of the outer edge of the target table in the first image, add the outer edge of the target table to generate the second image.
8. An image processing apparatus, characterized in that, include: The acquisition module is used to acquire a first image, the first image containing a first table, the first table being missing at least one outer edge of the target table; The recognition module is used to perform content recognition on the first image to obtain the first position information of the first table and the second position information of the text objects in the first image. The processing module is configured to obtain a target outline based on the first position information and the second position information, wherein the target outline is used to indicate the minimum outline that accommodates the text object and the first table. A generation module is used to generate a second image based on the target outline, the second image containing the first table and the outer edge of the target table.
9. An electronic device, characterized in that, include: A processor, and a memory communicatively connected to the processor; The memory stores computer-executed instructions; The processor executes computer execution instructions stored in the memory to implement the method as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions, which, when executed by a processor, are used to implement the method as described in any one of claims 1 to 7.
11. A computer program product, characterized in that, Includes a computer program that, when executed by a processor, implements the method of any one of claims 1 to 7.