Image processing method and device, electronic device and computer readable storage medium
By identifying and processing the input images and generating tables, determining key object cells, extracting object content and recording content, the accuracy and efficiency of extracting table information of electronic version files in the prior art is solved, and the rapid and accurate identification and information extraction of complex tables are achieved.
Patent Information
- Application Number
- CN202111215743.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-10-19
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2041-10-19
AI Technical Summary
It is difficult for the prior art to accurately identify and extract table information in electronic version files, especially when the table structure is complex or the noise is high.
By obtaining the input image, the recognition process is performed to obtain multiple object blocks, the key object block is determined, the unit table is generated, and the key object cells are determined based on the key object blocks, and the corresponding object content and record content are extracted.
It realizes rapid and accurate identification and information extraction of complex tables, and improves the efficiency and accuracy of image processing.
Smart Images

Figure CN113963366B_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present disclosure relate to an image processing method, an image processing apparatus, an electronic device, and a non-transitory computer-readable storage medium. Background Art
[0002] With the development of electronic office platforms, users often scan or take photos of various documents to save them as electronic files. At the same time, they also hope to process the scanned or photographed electronic files accordingly to obtain relevant information in the electronic files. Tables are a common form of presenting information in a structured manner and are often used in various files. In the process of processing electronic files with tables, it is usually necessary to identify the tables in the electronic files and extract various key information in the tables. However, due to the particularity of tables, how to accurately identify tables and accurately extract key information in tables is a problem that needs to be solved urgently. Summary of the invention
[0003] At least one embodiment of the present disclosure provides an image processing method, comprising: acquiring an input image; performing recognition processing on the input image to obtain a plurality of object blocks, wherein each object block includes an object content; determining a plurality of key object blocks among the plurality of object blocks; generating a unit table based on a plurality of block positions respectively corresponding to the plurality of object blocks, wherein the unit table includes a plurality of cells, the plurality of cells include a plurality of object cells, each object cell includes an object block; determining N key object cells among the plurality of object cells based on the plurality of key object blocks, wherein each key object cell among the N key object cells includes a key object block, and in the input image, the N key object cells are arranged in the same row along a first direction; determining at least one record content corresponding to N object contents in N key object blocks among the N key object cells based on a plurality of object contents in the plurality of object blocks, wherein each record content includes at least one object content; outputting M object contents among the N object contents and / or outputting L record contents among the at least one record content, wherein N, M and L are positive integers, and N is greater than 1.
[0004] For example, in an image processing method provided by an embodiment of the present disclosure, each object block has an object attribute, and the object attribute corresponding to each key object block is any key object attribute in the key attribute group.
[0005] For example, in an image processing method provided by an embodiment of the present disclosure, outputting M object contents among the N object contents includes: determining selection output information; determining and outputting object contents among the N object contents corresponding to the selection output information, wherein the M object contents include object contents among the N object contents corresponding to the selection output information.
[0006] For example, in an image processing method provided by an embodiment of the present disclosure, the selection output information includes an image type corresponding to the input image, the key attribute group includes multiple key object attributes, and determining and outputting the object content corresponding to the selection output information among the N object contents includes: determining at least one key object attribute corresponding to the image type among the multiple key object attributes; determining at least one key object block whose object attribute among the N key object blocks is any one of the at least one key object attribute; and outputting the object content in the at least one key object block, wherein the object content corresponding to the selection output information is the object content in the at least one key object block.
[0007] For example, in an image processing method provided by an embodiment of the present disclosure, the selected output information includes information predefined by a user.
[0008] For example, in an image processing method provided by an embodiment of the present disclosure, outputting L record contents among the at least one record content includes: determining selection output information; determining and outputting the record content among the at least one record content corresponding to the selection output information, wherein the L record contents include the record content among the at least one record content corresponding to the selection output information.
[0009] For example, in an image processing method provided by an embodiment of the present disclosure, the selected output information includes an image type corresponding to the input image, the key attribute group includes multiple key object attributes, and determining and outputting the record content corresponding to the selected output information in the at least one record content includes: determining at least one key object attribute corresponding to the image type in the multiple key object attributes; determining at least one key object block whose object attribute in the N key object blocks is any one of the at least one key object attributes; and outputting the record content in the at least one record content corresponding to the object content in the at least one key object block, wherein the record content corresponding to the selected output information is the record content in the at least one record content corresponding to the object content in the at least one key object block.
[0010] For example, in an image processing method provided by an embodiment of the present disclosure, based on the multiple key object blocks, N key object cells among the multiple object cells are determined, including: based on the multiple key object blocks, multiple key object cells among the multiple object cells are determined, wherein the multiple key object cells are object cells among the multiple object cells that correspond one-to-one to the multiple key object blocks; based on the positions of the multiple key object cells, the N key object cells among the multiple key object cells are determined.
[0011] For example, in an image processing method provided by an embodiment of the present disclosure, the multiple cells also include multiple blank cells, each blank cell does not include an object block, and based on the multiple object contents in the multiple object blocks, at least one record content corresponding to the N object contents in the N key object blocks in the N key object cells is determined, including: for the i-th key object cell among the N key object cells: obtaining P cells located on the first side of the i-th key object cell in the second direction; in response to the P cells including at least one object cell, based on the P cells, determining the record content corresponding to the object content in the key object block in the i-th key object cell; in response to the P cells being blank cells, determining that the key object block in the i-th key object cell does not have corresponding record content.
[0012] For example, in an image processing method provided by an embodiment of the present disclosure, in the input image, the edge of any one of the P cells in the first direction does not exceed the edge of the i-th key object cell in the first direction.
[0013] For example, in an image processing method provided in an embodiment of the present disclosure, in response to the P cells including at least one object cell, based on the P cells, determining the record content corresponding to the object content in the key object block in the ith key object cell, including: in response to P being 1 and the P cells being object cells, taking the object content in the object block in the P cells as the record content corresponding to the object content in the key object block in the ith key object cell; in response to P being greater than 1: merging the P cells to obtain at least one merged cell corresponding to the ith key object cell, wherein each merged cell includes at least one object block, and the merged content corresponding to each merged cell is the object content in the at least one object block; based on at least one merged content respectively corresponding to the at least one merged cell, determining the record content corresponding to the object content in the key object block in the ith key object cell.
[0014] For example, in an image processing method provided by an embodiment of the present disclosure, the object content in each object block includes text, and in this case, the object block is a text block. The object attribute of each object block includes a part of speech, and the merging process includes: based on the positions of the P cells, determining at least one cell group, wherein each cell group includes at least one cell of the P cells, and in the case where any cell group includes multiple cells, the multiple cells in the any cell group are arranged in a row along the first direction; based on the at least one cell group, determining at least one intermediate merged cell corresponding to the at least one cell group one by one, wherein in the case where any cell group includes multiple cells, the multiple cells in the any cell group are merged as the intermediate merged cell corresponding to the any cell group, and in the case where any cell group includes one cell, one cell in the any cell group is directly used as the intermediate merged cell corresponding to the any cell group; in response to the at least one intermediate merged cell including multiple intermediate merged cells, in the case where the first intermediate merged cell and the second intermediate merged cell to be merged in the multiple intermediate merged cells meet the merging condition, merging the first intermediate merged cell and the second intermediate merged cell, and in response to the at least one intermediate merged cell including one intermediate merged cell, using the one intermediate merged cell as a merged cell.
[0015] For example, in an image processing method provided by an embodiment of the present disclosure, the merging condition includes: the first intermediate merged cell and the second intermediate merged cell both include object blocks, and the object content in the object block in the first intermediate merged cell and the object content in the object block in the second intermediate merged cell are the same or semantically continuous, and in the input image, the first intermediate merged cell and the second intermediate merged cell are arranged successively in the second direction; or, the first intermediate merged cell and / or the second intermediate merged cell do not include object blocks, and in the input image, the first intermediate merged cell and the second intermediate merged cell are arranged successively in the second direction.
[0016] For example, in an image processing method provided by an embodiment of the present disclosure, in the input image, the first direction and the second direction are perpendicular to each other.
[0017] For example, in an image processing method provided in an embodiment of the present disclosure, based on at least one merged content corresponding to the at least one merged cell, the record content corresponding to the object content in the key object block in the i-th key object cell is determined, including: obtaining a filtering rule; based on the filtering rule, filtering the at least one merged content to determine the record content corresponding to the object content in the key object block in the i-th key object cell.
[0018] For example, in an image processing method provided in an embodiment of the present disclosure, the input image is recognized and processed to obtain a plurality of object blocks, including: using an object recognition model to recognize and process the input image to obtain a plurality of initial object blocks; using an object classification model to classify the plurality of initial object blocks according to object attributes to obtain the plurality of object blocks.
[0019] For example, in an image processing method provided by an embodiment of the present disclosure, determining multiple key object blocks among the multiple object blocks includes: using an object classification model to classify the multiple object blocks according to object attributes to obtain the multiple object blocks and multiple object attributes corresponding one-to-one to the multiple object blocks; based on the multiple object attributes, determining the multiple key object blocks from the multiple object blocks.
[0020] For example, in an image processing method provided by an embodiment of the present disclosure, a unit table is generated based on a plurality of block positions respectively corresponding to the plurality of object blocks, including: determining a plurality of separation lines based on the plurality of block positions, wherein there is at least one separation line between every two object blocks; and separating the plurality of object blocks by means of the plurality of separation lines to form the unit table.
[0021] At least one embodiment of the present disclosure further provides an image processing device, comprising: an image acquisition module, configured to acquire an input image; a recognition module, configured to perform recognition processing on the input image to obtain a plurality of object blocks, wherein each object block includes an object content; an object block determination module, configured to determine a plurality of key object blocks among the plurality of object blocks; a table generation module, configured to generate a unit table based on a plurality of block positions respectively corresponding to the plurality of object blocks, wherein the unit table includes a plurality of cells, the plurality of cells include a plurality of object cells, each object cell includes an object block; and a cell determination module, configured to determine a plurality of key object blocks among the plurality of object cells based on the plurality of key object blocks. N key object cells, wherein each of the N key object cells includes a key object block, and in the input image, the N key object cells are arranged in the same row along a first direction; a record content determination module, configured to determine at least one record content corresponding to N object contents in N key object blocks in the N key object cells based on multiple object contents in the multiple object blocks, wherein each record content includes at least one object content; an output module, configured to output M object contents among the N object contents and / or output L record contents among the at least one record content, wherein N, M and L are positive integers, and N is greater than 1.
[0022] At least one embodiment of the present disclosure further provides an electronic device, comprising: a memory, which non-transitorily stores computer-executable instructions; and a processor, which is configured to execute the computer-executable instructions, wherein the computer-executable instructions, when executed by the processor, implement the image processing method according to any embodiment of the present disclosure.
[0023] At least one embodiment of the present disclosure further provides a non-transitory computer-readable storage medium, wherein the non-transitory computer-readable storage medium stores computer-executable instructions, and when the computer-executable instructions are executed by a processor, the image processing method according to any embodiment of the present disclosure is implemented. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] In order to more clearly illustrate the technical solutions of the embodiments of the present disclosure, the drawings of the embodiments will be briefly introduced below. Obviously, the drawings in the following description only relate to some embodiments of the present disclosure, but are not intended to limit the present disclosure.
[0025] Figure 1 A schematic flow chart of an image processing method provided for at least one embodiment of the present disclosure;
[0026] Figure 2A A schematic diagram of an input image provided by at least one embodiment of the present disclosure;
[0027] Figure 2B A schematic diagram of another input image provided for at least one embodiment of the present disclosure;
[0028] Figure 3 For Figure 2B A schematic diagram of a cell table generated after processing a plurality of object blocks in an input image shown;
[0029] Figure 4 A schematic diagram of another input image provided by an embodiment of the present disclosure;
[0030] Figure 5 A schematic diagram of an image processing device provided by at least one embodiment of the present disclosure;
[0031] Figure 6 A schematic diagram of an electronic device provided for at least one embodiment of the present disclosure;
[0032] Figure 7 A schematic diagram of a non-transitory computer-readable storage medium provided for at least one embodiment of the present disclosure. DETAILED DESCRIPTION
[0033] In order to make the purpose, technical solution and advantages of the embodiments of the present disclosure clearer, the technical solution of the embodiments of the present disclosure will be clearly and completely described below in conjunction with the drawings of the embodiments of the present disclosure. Obviously, the described embodiments are part of the embodiments of the present disclosure, not all of the embodiments. Based on the described embodiments of the present disclosure, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present disclosure.
[0034] Unless otherwise defined, the technical terms or scientific terms used in the present disclosure should be understood by people with ordinary skills in the field to which the present disclosure belongs. The "first", "second" and similar words used in the present disclosure do not indicate any order, quantity or importance, but are only used to distinguish different components. "Include" or "comprise" and similar words mean that the elements or objects appearing before the word cover the elements or objects listed after the word and their equivalents, without excluding other elements or objects. "Connect" or "connected" and similar words are not limited to physical or mechanical connections, but can include electrical connections, whether direct or indirect. "Up", "down", "left", "right" and the like are only used to indicate relative positional relationships. When the absolute position of the described object changes, the relative positional relationship may also change accordingly.
[0035] In order to keep the following description of the embodiments of the present disclosure clear and concise, the present disclosure omits detailed descriptions of some known functions and known components.
[0036] At least one embodiment of the present disclosure provides an image processing method. The image processing method includes: acquiring an input image; performing recognition processing on the input image to obtain multiple object blocks, wherein each object block includes an object content; determining multiple key object blocks in the multiple object blocks; generating a unit table based on multiple block positions corresponding to the multiple object blocks, wherein the unit table includes multiple cells, the multiple cells include multiple object cells, and each object cell includes an object block; based on the multiple key object blocks, determining N key object cells in the multiple object cells, wherein each key object cell in the N key object cells includes a key object block, and in the input image, the N key object cells are arranged in the same row along a first direction; based on the multiple object contents in the multiple object blocks, determining at least one record content corresponding to N object contents in N key object blocks in the N key object cells, wherein each record content includes at least one object content; outputting M object contents in the N object contents and / or outputting L record contents in at least one record content. N, M and L are positive integers, and N is greater than 1.
[0037] In the image processing method provided by the embodiment of the present disclosure, first, a unit table is generated based on the block position corresponding to the object block, then, N key object cells are quickly and accurately determined based on the position information of each object cell in the unit table, and finally, based on multiple object contents in multiple object blocks and N key object cells, the content to be output can be determined and output, for example, the content to be output is M object contents among N object contents and / or L record contents among at least one record content. Therefore, the image processing method provided by the embodiment of the present disclosure can quickly and accurately identify the key information in the input image, so as to conveniently and quickly extract and output the content to be output.
[0038] At least one embodiment of the present disclosure further provides an image processing apparatus, an electronic device, and a non-transitory computer-readable storage medium corresponding to the above-mentioned image processing method.
[0039] The image processing method provided in the embodiment of the present disclosure can be applied to the image processing device provided in the embodiment of the present disclosure, and the image processing device can be configured on an electronic device. The electronic device can be a personal computer, a mobile terminal, etc., and the mobile terminal can be a hardware device such as a mobile phone or a tablet computer.
[0040] The embodiments of the present disclosure are described in detail below in conjunction with the accompanying drawings, but the present disclosure is not limited to these specific embodiments. It should be noted that in the embodiments of the present disclosure, "multiple" means two or more, that is, at least two.
[0041] Figure 1 A schematic flowchart of an image processing method provided for at least one embodiment of the present disclosure. Figure 2A A schematic diagram of an input image provided for at least one embodiment of the present disclosure. Figure 2B A schematic diagram of another input image provided for at least one embodiment of the present disclosure.
[0042] like Figure 1 As shown, the image processing method provided by at least one embodiment of the present disclosure includes the following steps S10 to S16.
[0043] Step S10: Obtain an input image.
[0044] Step S11: Perform recognition processing on the input image to obtain a plurality of object blocks. For example, each object block includes an object content.
[0045] Step S12: Determine a plurality of key object blocks among a plurality of object blocks.
[0046] Step S13: Generate a unit table based on the multiple block positions corresponding to the multiple object blocks. For example, the unit table includes multiple cells, the multiple cells include multiple object cells, and each object cell includes one object block.
[0047] Step S14: Based on the multiple key object blocks, determine N key object cells among the multiple object cells. For example, each of the N key object cells includes a key object block, and in the input image, the N key object cells are arranged in the same row along the first direction. For example, in some embodiments, in the input image, the N key object cells are continuously arranged in the same row along the first direction; for example, in other embodiments, in the input image, the N key object cells are arranged in the same row along the first direction, and there is at least one non-key object cell in the row where the N key object cells are located, and the non-key object cell is a cell other than the N key object cells in the multiple cells.
[0048] Step S15: Based on the multiple object contents in the multiple object blocks, determine at least one record content corresponding to the N object contents in the N key object blocks in the N key object cells. For example, each record content includes at least one object content.
[0049] Step S16: output M object contents among the N object contents and / or output L record contents among the at least one record content.
[0050] For example, N, M, and L are positive integers, and N is greater than 1.
[0051] For example, for step S10, the input image may be an image obtained by a user scanning or photographing an object, such as a business card, a test paper, a test sheet, a document, an invoice, etc. The input image may include a form (e.g., a wired form and / or a wireless form).
[0052] For example, the shape of the input image can be a regular shape such as a rectangle or a square, or an irregular shape, and the shape and size of the input image can be set by the user according to actual conditions. For example, the input image can be an image taken by a digital camera or a mobile phone, or an image scanned by a scanner. For example, the input image can be an original image directly collected by a digital camera, a mobile phone or a scanner. In addition, in order to avoid the influence of the data quality and data imbalance of the original image on the recognition of the input image, the image processing method provided by the embodiment of the present disclosure can also include the operation of preprocessing the original image, that is, the input image can also be an image obtained after preprocessing the original image. For example, the preprocessing can eliminate irrelevant information or noise information in the original image, so as to better process the original image. The preprocessing can, for example, include processing such as bending correction, scaling, cropping, gamma correction, image enhancement or noise reduction filtering on the original image, so as to improve the accuracy and reliability of various operations in subsequent steps. The bending correction can include global correction and local correction. The global correction can correct the global offset of the object content in the original image, so as to avoid the problem of inaccurate subsequent recognition of the object content due to the skewness of the object content. In some embodiments, global correction can be implemented by using an algorithm in opencv based on the idea of Leptonica (Leptonica is an open source image processing and image analysis library); in other embodiments, global correction can also be implemented by using a machine learning (e.g., neural network) method. Since some details in the original image may not be adjusted after global correction is performed on the original image, some supplementary corrections can be performed on the details ignored in the global correction process through local correction, thereby reducing or preventing the loss of details caused by global correction, and improving the accuracy and reliability of the input image obtained after the correction processing of the original image.
[0053] For example, the input image can be a grayscale image or a color image.
[0054] For example, step S11 may include: using an object recognition model to perform recognition processing on the input image to obtain a plurality of initial object blocks; using an object classification model to perform classification processing on the plurality of initial object blocks according to object attributes to obtain a plurality of object blocks.
[0055] For example, each initial object block includes at least one object content. The shape of each initial object block can be a regular shape such as a rectangle, a square, or an irregular shape, as long as the initial object block can cover the corresponding object content. Similarly, the shape of each object block can be a regular shape such as a rectangle, a square, or an irregular shape, as long as the object block can cover the corresponding object content.
[0056] For example, the object content in each object block may include at least one character, at least one graphic (circle, rectangle, etc.), at least one symbol (colon, comma, period, percent sign, etc.) or at least one data, etc. The characters may include Chinese characters and / or foreign characters, and the characters may be printed characters and / or handwritten characters, etc. For example, the object content in each object block is arranged in a row along a first direction. In the description of the present disclosure, it is taken that the object content contained in the input image includes at least one of characters, numbers and symbols, and at this time, the object attribute may be a part of speech. For example, the first direction may be a row direction or a column direction of the character / number arrangement.
[0057] It should be noted that, in the embodiment of the present disclosure, all contents (text, numbers, symbols and / or graphics) in an object block are one object content.
[0058] For example, the object recognition model and the object classification model can be implemented using machine learning technology and run on a general computing device or a dedicated computing device. The object recognition model and the object classification model are both pre-trained neural network models. For example, the object recognition model and the object classification model can be implemented using a neural network such as a deep convolutional neural network (DEEP-CNN).
[0059] In the process of training the object recognition model, first, the sample image can be annotated to mark each sample object block in the sample image, and then the initial object recognition model is trained with the annotated sample image, and finally the object recognition model is trained. In the process of training the object classification model, first, the sample image can be annotated to mark the key sample object blocks with key object attributes (for example, headers) and the non-key sample object blocks without key object attributes (for example, their object attributes may include data attributes, etc.), and then the initial object classification model is trained with the annotated sample image, and finally the object classification model is trained.
[0060] It should be noted that in an embodiment of the present disclosure, after processing the input image, the input image and the identified key object cells in the input image can also be added as samples to the training set used to train the initial object classification model and the initial object recognition model, thereby expanding the number of samples in the training set and optimizing the network model.
[0061] For example, in some embodiments, the object classification model can be implemented based on named entity recognition (NER) technology. NER, also known as proper name recognition, is a basic task in natural language processing and has a wide range of applications. Named entities generally refer to entities with specific meanings or strong referentiality in texts, usually including names of people, places, names of organizations, dates and times, proper nouns, etc. In other embodiments, the object classification model can be a pre-trained text classification neural network model.
[0062] For example, the object classification model can classify multiple initial object blocks based on object attributes (e.g., part of speech), thereby obtaining multiple object blocks. Each object block has a corresponding object attribute, and the object attribute can be set in different ways. For example, the object attribute can include a table title, a table header, a table body, etc., and the multiple object blocks include an object block whose object attribute is a table title, an object block whose object attribute is a table header, an object block whose object attribute is a table body, and the like. The object attribute can also include a person's name, a place name, an organization name, a date and time, a proper noun, and the like. The multiple object blocks include an object block whose object attribute is a person's name, an object block whose object attribute is a place name, an object block whose object attribute is an organization name, an object block whose object attribute is a date and time, an object block whose object attribute is a proper noun, and the like. In an embodiment of the present disclosure, the object classification model can identify each initial object block to determine the object attribute of each initial object block, and classify the multiple initial object blocks based on the object attribute.
[0063] For example, each object block may correspond to at least one initial object block.
[0064] For example, the object classification model can also perform a segmentation operation. Since the distances between some object contents are very close in the input image, these object contents may be included in the same initial object block during the object recognition model's recognition process. The segmentation operation can split these object contents in the initial object block and divide them into different object blocks. For example, for the object content "name so-and-so", when the object recognition module recognizes it, "name so-and-so" is only located in one initial object block, and in the process of classification by the object classification model, this initial object block needs to be split into two object blocks: the object block corresponding to "name" and the object block corresponding to "so-and-so". The object attribute of the object block corresponding to "name" can be a proper noun, while the object attribute of the object block corresponding to "so-and-so" can be a name.
[0065] For example, the difference between the initial object block and the object block corresponding to each other is that the object block has an object attribute, and other properties of the initial object block and the object block corresponding to each other may be the same.
[0066] For example, in the input image, the first direction and the second direction may be substantially perpendicular to each other, such as Figure 2A As shown, in the input image 200, the first direction may be the X-axis direction of the image coordinate system OXY (e.g., the horizontal direction), and is the row direction of the text / number arrangement, and the second direction may be the Y-axis direction of the image coordinate system OXY (e.g., the vertical direction), and is the column direction of the text / number arrangement. The embodiments of the present disclosure are not limited to this. For another example, the first direction may be the column direction of the text / number arrangement (e.g., the vertical direction), and the second direction may be the row direction of the text / number arrangement (e.g., the horizontal direction). It should be noted that in other embodiments, the first direction and the second direction are not the horizontal direction and the vertical direction, but are directions having a certain angle with the horizontal direction and the vertical direction. According to actual conditions, the first direction and the second direction may be any appropriate directions.
[0067] For example, Figure 2A As shown, the input image 200 is located in the image coordinate system OXY, and the input image 200 includes a table area 201, which indicates the area where the table in the input image is located. The table area 201 includes a plurality of object blocks 202, each of which includes an object content, such as Figure 2A As shown, the object block 202 includes an object block 202A and an object block 202B. The object content included in the object block 202A is six numbers "391001". The six numbers "391001" are arranged along a first direction (i.e. Figure 2A The object content included in the object block 202B is ten Chinese characters "Motor vehicle road parking fee collection", and the ten Chinese characters "Motor vehicle road parking fee collection" are also arranged in a row along the first direction. In some other embodiments of the present disclosure, the object content included in the object block 202 can also be arranged in a row along the second direction or in multiple rows or columns along the first direction and the second direction respectively.
[0068] For example, Figure 2B As shown, the input image 300 includes a table area 301, and the table area 301 includes a plurality of object blocks 302, each of which includes an object content. Figure 2B As shown, the object content included in the object block 302 is "80.00", "80.00" includes four numbers (8, 0, 0, 0) and a symbol (.), and "80.00" is arranged in a row along the first direction.
[0069] For example, step S12 may include: using an object classification model to classify multiple object blocks according to object attributes to obtain multiple object blocks and multiple object attributes corresponding to the multiple object blocks; based on the multiple object attributes, determining multiple key object blocks from the multiple object blocks.
[0070] For example, in step S12, the plurality of key object blocks are object blocks whose object attributes among the plurality of object blocks are key object attributes in the key attribute group.
[0071] For example, a key attribute group includes multiple key object attributes, and the object attribute corresponding to each key object block is the key object attribute in the key attribute group. For example, in some embodiments, the key object attribute may include a header, and in this case, the object block whose object attribute is the header is the key object block. For another example, in some other embodiments, the key object attribute may include a proper noun, and the object block whose object attribute is a proper noun is the key object block. It should be noted that the key object attribute can be set according to actual application requirements, and the embodiments of the present disclosure do not specifically limit this.
[0072] For example, step S13 may include: determining a plurality of separation lines based on a plurality of block positions; and performing separation processing on a plurality of object blocks through the plurality of separation lines to form a unit table.
[0073] For example, among the plurality of object blocks, there is at least one separation line between every two object blocks.
[0074] For example, each separation line may extend along the first direction or the second direction but may not pass through any object block.
[0075] For example, multiple separator lines can be determined based on the block position of the object block. The separator line may be a solid line in the table or a line added in the blank area. In the unit table, each object block is ultimately located in an object cell. After multiple object blocks are separated by multiple separator lines, the coordinates of the object blocks may be represented by unit rows and unit columns in the unit table instead of pixels. Each object block has an integer row number and column number. Multiple object blocks will not overlap, and the object content in the object block will not be divided.
[0076] For example, various suitable methods may be used to generate a unit table. For example, in the above description of the present disclosure, the process of generating a table unit is: first, a separation line is generated (the separation line may extend along the first direction or the second direction, but will not pass through any object block), and each object block is separated, thereby generating a unit table. The embodiments of the present disclosure do not limit the specific process and method of generating a unit table.
[0077] For example, a plurality of dividing lines may be used to divide the object block, and the plurality of dividing lines may include a horizontal dividing line and a vertical dividing line, and the horizontal dividing line is parallel to the horizontal direction (eg, the first direction), that is, Figure 2A In the X-axis direction of the image coordinate system OXY shown in FIG. , the vertical dividing line is parallel to the vertical direction (eg, the second direction), that is, Figure 2A The Y-axis direction of the image coordinate system OXY is shown in FIG. Because the distribution of multiple object blocks is relatively scattered, it is necessary to cluster multiple object blocks into a small number of dividing lines to divide the left and right edges of the object blocks (such as Figure 2A The two edges in the X-axis direction shown in FIG. 1 are clustered to the coordinates of the vertical dividing line (the horizontal coordinate (X-axis coordinate) in the image coordinate system OXY), and the upper and lower edges of the object block (such as Figure 2A The two edges in the Y-axis direction shown in the figure are clustered to the coordinates of the horizontal dividing line (the vertical coordinate (Y-axis coordinate) in the image coordinate system OXY), and each object block can be separated according to the edge of the object block. After the division is completed, different object blocks are located in different object cells. In addition, after separation based on multiple dividing lines, blank cells (i.e., cells that do not include object blocks) may also be formed. The object cells and blank cells together constitute a unit table.
[0078] Figure 3 For Figure 2B FIG. 1 is a schematic diagram of a cell table generated after processing multiple object blocks in an input image.
[0079] like Figure 3 As shown, each complete rectangular grid is a cell, and the unit table includes multiple cells. The multiple cells can be arranged into multiple rows and columns along the first direction and the second direction. The sizes of the cells can be different, or the sizes of some cells can be the same. The multiple cells include multiple object cells, such as Figure 3 The object cells 311 shown in FIG. 1 include each object cell including an object block, and the plurality of cells also include a plurality of blank cells, such as Figure 3 The blank cells 310 shown in the figure each do not include any object block. Figure 3 In the figure, the shaded part in each rectangular grid represents an object block, and the shades of the shaded part represent object blocks with different object attributes. For example, the object attributes of the object block in object cell 311A are different from the object attributes of the object block in object cell 311B. In some examples, the object attributes of the object block in object cell 311A are key object attributes, and thus object cell 311A is a key object cell.
[0080] For example, each cell is located in a row and column according to its position, such as Figure 3As shown, the target cell 311A is located in the third row and the sixth column of the cell table, and the target cell 311B is located in the first row and the ninth column of the cell table.
[0081] Figure 4 A schematic diagram of another input image provided according to an embodiment of the present disclosure.
[0082] For example, in some embodiments, step S14 may include: determining multiple key object cells among multiple object cells based on multiple key object blocks; and determining N key object cells among multiple key object cells based on the positions of the multiple key object cells. In the embodiments of the present disclosure, the key object cells required by the user, i.e., N key object cells, may be determined based on the positions of the key object cells. The specific value of N may be determined based on actual conditions, and the embodiments of the present disclosure are not limited thereto.
[0083] For example, the multiple key object cells are object cells in the multiple object cells that correspond one-to-one to the multiple key object blocks.
[0084] For example, key object cells that are consecutive in position are found among multiple key object cells to obtain N key object cells. In the input image, the N key object cells are arranged in the same row along a first direction. If the first direction is the row direction of text arrangement, the N key object cells are located in the same row; if the first direction is the column direction of text arrangement, the N key object cells are located in the same column.
[0085] For example, in a cell table, in a row of cells where N key object cells are located, there may be at least one non-key object cell. If the number of non-key object cells in a row of cells where N key object cells are located is within a predetermined number threshold, then this row can still be considered as an object that needs to be processed, that is, the subsequent steps (step S15 and step S16) can be performed on the N key object cells. For example, the predetermined number threshold can be determined according to the number of columns in a row of cells where N key object cells are located. When the number of columns in a row of cells where N key object cells are located (that is, the number of cells in the row of cells) is greater, the predetermined number threshold can be greater. When there are more columns, it means that the number of non-key object cells that can be tolerated by the row of cells where N key object cells are located is greater. For example, when the row of cells where N key object cells are located includes 5 columns, the predetermined number threshold can be 1, that is, it can tolerate that 1 non-key object cell is included in the row of cells where N key object cells are located; when the row of cells where N key object cells are located includes 10 columns, the predetermined number threshold can be 2, that is, it can tolerate that 2 non-key object cells are included in the row of cells where N key object cells are located. For another example, the predetermined quantity threshold may also be set by the user according to actual conditions, and the present disclosure does not impose any specific limitation on this.
[0086] For example, in some embodiments, when an object block whose object attribute is a table header is a key object block, and the predetermined number threshold is 2, the number of non-key object cells in a row of cells where N key object cells are located is less than or equal to 2, then the row of cells where the N key object cells are located can be considered as the row where the table header is located.
[0087] It should be noted that the “non-key target cell” refers to a cell in the cell table that is not a key target cell.
[0088] For example, Figure 4 As shown, the multiple key object blocks include an object block with object content of "serial number", an object block with object content of "goods (service) name", an object block with object content of "specification model", an object block with object content of "unit", an object block with object content of "quantity", an object block with object content of "unit price", an object block with object content of "amount", an object block with object content of "tax rate" and an object block with "tax amount", and the object cells including the above multiple key object blocks are key object cells, in Figure 4In the figure, the key object cells are indicated by dashed rectangular boxes 401 to 409 (hereinafter 401 to 409 represent key object cells). In terms of position, the key object cells 401 to 409 are arranged in a row along a first direction (for example, continuously arranged), thereby determining that the key object cells 401 to 409 are the above-mentioned N key object cells, where N is 9.
[0089] For example, in some embodiments, step S15 may include: for the i-th key object cell among N key object cells: obtaining P cells located on the first side of the i-th key object cell in the second direction; in response to the P cells including at least one object cell, determining, based on the P cells, the record content corresponding to the object content in the key object block in the i-th key object cell; in response to the P cells being blank cells, determining that the key object block in the i-th key object cell does not have corresponding record content.
[0090] For example, each record content includes at least one object content.
[0091] For example, in the input image, the edge of any one of the P cells in the first direction does not exceed the edge of the i-th key object cell in the first direction.
[0092] For example, in the vertical direction, two sides of each key object cell that are opposite to each other are the upper side and the lower side, and in the horizontal direction, two sides of each key object cell that are opposite to each other are the left side and the right side.
[0093] For example, in some embodiments, Figure 4 As shown, if the first direction is the horizontal direction and the second direction is the vertical direction, then the first side of the i-th key object cell in the second direction may be the lower side of the i-th key object cell (for example, Figure 4 , cell 4092 is located at the lower side of cell 4091); in other embodiments, if the first direction is the vertical direction and the second direction is the horizontal direction, then the first side of the i-th key object cell in the second direction may be the right side of the i-th key object cell (for example, Figure 4 , key object cell 402 is located on the right side of key object cell 401).
[0094] For example, Figure 4As shown, in some examples, if the ith key object cell is the key object cell 409, at this time, the P cells located on the first side of the ith key object cell (i.e., the key object cell 409) in the second direction include a cell 4091 containing object content "917.59", a cell 4092 containing object content "-75.93", a cell 4093 containing object content "458.80", and a cell 4094 containing object content "-37.96", that is, at this time P is 4, and the P cells are all object cells, and the record content corresponding to the object content in the key object block in the ith key object cell can be determined based on the object content of the object block in the P cells.
[0095] For example, Figure 4 As shown, in other examples, if the ith key object cell is the key object cell 406, at this time, the P cells located on the first side of the ith key object cell (i.e., the key object cell 406) in the second direction include a cell 4061 containing object content "3529.20", a blank cell 4062, a cell 4063 containing object content "3529.20", and a blank cell 4064, that is, at this time P is still 4, but the P cells include two object cells and two blank cells, and the record content corresponding to the object content in the key object block in the ith key object cell can be determined based on the object content of the object blocks in the object cells in the P cells (i.e., the cell 4061 containing object content "3529.20" and the cell 4063 containing object content "3529.20").
[0096] For another example, if all P cells are blank cells, the key object block in the i-th key object cell has no corresponding record content.
[0097] For example, in step S15, in response to P cells including at least one object cell, based on the P cells, determining the record content corresponding to the object content in the key object block in the i-th key object cell includes: in response to P being 1 and the P cells being object cells, taking the object content in the object block in the P cells as the record content corresponding to the object content in the key object block in the i-th key object cell; in response to P being greater than 1: merging the P cells to obtain at least one merged cell corresponding to the i-th key object cell; and determining the record content corresponding to the object content in the key object block in the i-th key object cell based on at least one merged content respectively corresponding to at least one merged cell. For example, each object content may correspond to one record content or multiple record contents.
[0098] For example, each merged cell includes at least one object block, and the merged content corresponding to each merged cell is the object content in the at least one object block included in the merged cell.
[0099] Since adjacent cells among the P cells may contain the same object block, or the object contents in the object blocks of adjacent cells are semantically continuous, the adjacent cells need to be merged into one merged cell. Figure 4 The cells 4021 to 4023 shown need to be merged into one merged cell (described in detail below).
[0100] For example, the merging process includes: determining at least one cell group based on the positions of P cells; determining at least one intermediate merged cell corresponding one-to-one to the at least one cell group based on the at least one cell group; in response to at least one intermediate merged cell including multiple intermediate merged cells, merging the first intermediate merged cell and the second intermediate merged cell to be merged among the multiple intermediate merged cells if the first intermediate merged cell and the second intermediate merged cell to be merged meet the merging condition; in response to at least one intermediate merged cell including one intermediate merged cell, treating the one intermediate merged cell as a merged cell.
[0101] For example, each cell group includes at least one cell among the P cells, and in the case where any cell group includes a plurality of cells, the plurality of cells in any cell group are arranged in a row along the first direction.
[0102] For example, at least one cell group may be determined according to the positions of the P cells. For example, if K cells of the P cells are arranged in a row along the first direction, the K cells are considered as a cell group. If a row includes only one cell of the P cells, the one cell is considered as a cell group. K is a positive integer and is greater than 1.
[0103] For example, each cell group corresponds to an intermediate merged cell. When any cell group includes multiple cells, multiple cells in any cell group are merged as the intermediate merged cell corresponding to any cell group. That is to say, in each cell group, multiple cells arranged in a row along the first direction need to be merged into one intermediate merged cell. When any cell group includes one cell, one cell in any cell group is directly used as the intermediate merged cell corresponding to any cell group.
[0104] For example, the object content in each object block includes text, and the object attribute of each object block includes a part of speech. In some embodiments, the merging condition includes: the first intermediate merged cell and the second intermediate merged cell both include object blocks, and the object content in the object block in the first intermediate merged cell and the object content in the object block in the second intermediate merged cell are the same or semantically continuous, and in the input image, the first intermediate merged cell and the second intermediate merged cell are arranged in sequence in a second direction. In some other embodiments, the merging condition includes: the first intermediate merged cell and / or the second intermediate merged cell do not include object blocks, and in the input image, the first intermediate merged cell and the second intermediate merged cell are arranged in sequence in a second direction. In some other embodiments, the merging condition may also include: the spacing between the first intermediate merged cell and the second intermediate merged cell is less than a predetermined threshold. That is, if the spacing between the first intermediate merged cell and the second intermediate merged cell is less than the predetermined threshold, the first intermediate merged cell and the second intermediate merged cell may be merged. The predetermined threshold may be set according to actual conditions, and the embodiments of the present disclosure do not impose specific restrictions on this.
[0105] For example, merging needs to be done based on comprehensive consideration of all cells in the cell table. Figure 4 As shown, if cells 4021 to 4032 themselves are not sufficient to determine whether they need to be merged, the conditions of cells in other columns can be combined to determine whether to merge. For example, since the record content corresponding to key object cell 401 ("serial number") includes 4 groups of content, namely the number "1", number "2", number "3" and number "4" on the lower side of key object cell 401, it can be seen that the record content corresponding to the remaining key object cells 402 to 409 has at most 4 groups of content. Based on this and combined with the conditions of cells 4021 to 4032 themselves, cells 4021 to 4032 are merged.
[0106] It should be noted that in the embodiments of the present disclosure, the merging process also follows the following principles: multiple cells (cells or intermediate cells) including numbers will basically not be merged, multiple cells (cells or intermediate cells) including body text can be merged, and cells (cells or intermediate cells) including numbers and blank cells (cells or intermediate cells) can be merged. Whether to merge cells can be comprehensively judged based on object content, cell location, semantic model, spacing between cells, etc.
[0107] For example, Figure 4As shown, in one example, if the i-th key object cell is the key object cell 402, at this time, the P cells located on the first side of the i-th key object cell (ie, the key object cell 402) in the second direction include cells 4021 to 4032, that is, P is 12 at this time. The cells 4021-4032 need to be merged to obtain at least one merged cell corresponding to the i-th key object cell. For example, the object contents in the object blocks in cells 4021-4023 are semantically continuous, so that cells 4021-4023 can be merged into one merged cell, and the merged content in the merged cell includes the object contents in the object blocks in cells 4021-4023, that is, the merged content is "*Mobile communication device* [Super hit] Super-sensitive Leica triple-camera smart chip full-screen in-screen fingerprint version"; the object contents in the object blocks in cells 4024-4026 are semantically continuous, so that cells 4024-4026 can be merged into one merged cell, and the merged content in the merged cell includes the object contents in the object blocks in cells 4024-4026, that is, the merged content is "*Mobile communication device* [Super hit] Super-sensitive Leica triple-camera smart chip full-screen in-screen fingerprint version"; the object contents in the object blocks in cells 4027-4029 are semantically continuous, so that cells 4027-4029 can be merged into one merged cell, and the merged content in the merged cell includes the object contents in the object blocks in cells 4027-4029, that is, the merged content is "* mobile communication device * super-sensitive Leica triple-camera smart chip full-screen in-screen fingerprint version mobile phone 8GB+128GB"; the object contents in the object blocks in cells 4030-4032 are semantically continuous, so that cells 4030-4032 can be merged into one merged cell, and the merged content in the merged cell includes the object contents in the object blocks in cells 4030-4032, that is, the merged content is "* mobile communication device * super-sensitive Leica triple-camera smart chip full-screen in-screen fingerprint version mobile phone 8GB+128GB". Therefore, after the merge process, the four merged cells corresponding to the key object cell 402.
[0108] For example, in some embodiments, based on at least one merged content corresponding to at least one merged cell, the record content corresponding to the object content in the key object block in the i-th key object cell is determined, including: taking at least one merged content corresponding to at least one merged cell as the record content corresponding to the object content in the key object block in the i-th key object cell.
[0109] For example, Figure 4As shown, in some examples, if the i-th key object cell is key object cell 402, the record content corresponding to the object content in the key object block in the i-th key object cell is the four merged contents in the above four merged cells: "*Mobile communication device* [Super hit] Super-sensitive Leica triple camera smart chip full screen in-screen fingerprint version", "*Mobile communication device* [Super hit] Super-sensitive Leica triple camera smart chip full screen in-screen fingerprint version", "*Mobile communication device* Super-sensitive Leica triple camera smart chip full screen in-screen fingerprint version mobile phone 8GB+128GB", "*Mobile communication device* Super-sensitive Leica triple camera smart chip full screen in-screen fingerprint version mobile phone 8GB+128GB".
[0110] For example, in some embodiments, based on at least one merged content corresponding to at least one merged cell, determining the record content corresponding to the object content in the key object block in the i-th key object cell includes: obtaining a filtering rule; based on the filtering rule, filtering at least one merged content to determine the record content corresponding to the object content in the key object block in the i-th key object cell.
[0111] For example, the filtering rules can be set by the user according to the actual situation, and the record content that the user does not need can be filtered out based on the filtering rules. For example, for the key object cell 401, the object content included in the key object cell 401 is "sequence number". If the record content corresponding to the object content "sequence number" includes "total", it is determined based on the filtering rules that the record content corresponding to the object content "sequence number" must be a number, then the record content "total" needs to be filtered out, and is not used as the record content corresponding to the object content "sequence number".
[0112] For example, in some embodiments, in step S16, N object contents and at least one record content corresponding to the N object contents may be output, that is, N and M are equal at this time.
[0113] For example, in some other embodiments, the content that needs to be output among the N object contents and at least one recorded content corresponding to the N object contents may be determined based on the user's selection.
[0114] In step S16, outputting M of the N object contents includes: determining selected output information; determining and outputting the object contents corresponding to the selected output information among the N object contents. For example, the M object contents include the object contents corresponding to the selected output information among the N object contents.
[0115] In step S16, outputting L recorded contents in at least one recorded content includes: determining selected output information; determining and outputting recorded contents in at least one recorded content corresponding to the selected output information. For example, the L recorded contents include the recorded contents in at least one recorded content corresponding to the selected output information.
[0116] For example, in some embodiments, the output information is selected to include the image type corresponding to the input image, because in many cases, not all key object blocks in the N key object blocks in the N key object cells are what the user wants to pay attention to, and not all object contents in the N key object blocks and their corresponding record contents are needed. Therefore, the target data can be marked for input images of different image types. After sample training, the model can automatically identify the cells where the target data corresponding to the input images of different image types are located. Finally, when taking data, only the object content and record content corresponding to the cells where these target data are located can be extracted for output. For example, in some examples, the types of tables in input images of different image types (e.g., financial tables, invoice tables, etc.) are also different, and the image type of the input image can be determined based on the type of table in the input image.
[0117] In step S16, determining and outputting the object content corresponding to the selected output information among the N object contents includes: determining at least one key object attribute corresponding to the image type among the multiple key object attributes; determining at least one key object block whose object attribute among the N key object blocks is any one of the at least one key object attribute; and outputting the object content in the at least one key object block. For example, the object content corresponding to the selected output information is the object content in the at least one key object block.
[0118] In step S16, determining and outputting the record content corresponding to the selected output information in at least one record content includes: determining at least one key object attribute corresponding to the image type in a plurality of key object attributes; determining at least one key object block whose object attribute in N key object blocks is any one of the at least one key object attribute; and outputting the record content corresponding to the object content in the at least one key object block in at least one record content. For example, the record content corresponding to the selected output information is the record content corresponding to the object content in the at least one key object block in the at least one record content.
[0119] For example, in some other embodiments, the selected output information includes information predefined by the user, that is, the user can specify the content to be output according to the actual situation. The information predefined by the user can directly specify the object content and / or record content to be output, for example, Figure 4As shown, in some examples, the information pre-defined by the user can represent the output object content "goods (services) name" and "specification model" and their corresponding record content.
[0120] In step S16, determining and outputting the object content corresponding to the selected output information among the N object contents includes: determining at least one key object block corresponding to the information predefined by the user among the N key object blocks; and outputting the object content in the at least one key object block. For example, the object content corresponding to the selected output information is the object content in the at least one key object block.
[0121] In step S16, determining and outputting the recorded content corresponding to the selected output information in at least one recorded content includes: determining at least one key object block corresponding to the information predefined by the user in the N key object blocks; and outputting the recorded content corresponding to the object content in the at least one key object block in the at least one recorded content. For example, the recorded content corresponding to the selected output information is the recorded content corresponding to the object content in the at least one key object block in the at least one recorded content.
[0122] For example, Figure 4 As shown, the object contents in the key object blocks in the key object cells 401 to 409 are respectively “serial number”, “goods (services) name”, “specification model”, “unit”, “quantity”, “unit price”, “amount”, “tax rate” and “tax amount”. In some examples, the output information is selected as the information pre-defined by the user, and the user-pre-defined information may indicate that the object content “unit”, the object content “serial number”, the object content “tax rate” and the corresponding record contents are not outputted. That is to say, the user may not need the object content “unit”, the object content “serial number” and the object content “tax rate”, and thus the object content “unit”, the object content “serial number”, the object content “tax rate” and the corresponding record contents may not be outputted; in other examples, the output information is selected to include the image type corresponding to the input image, and for Figure 4 For the input image shown, selecting the output information may indicate not outputting the object content "unit", the object content "serial number", the object content "tax rate" and their corresponding record contents.
[0123] Based on this, according to the selected output information, at least one key object attribute corresponding to the selected output information among the multiple key object attributes can be determined; the object attribute among the N key object blocks is determined to be at least one key object block of any one of the at least one key object attributes, and the at least one key object block includes a key object block with object content "goods (services) name", a key object block with object content "specification model", a key object block with object content "quantity", a key object block with object content "unit price", a key object block with object content "amount", and a key object block with object content "tax amount", and finally output the object content "goods (services) name", "specification model", "quantity", "unit price", "amount", "tax amount" and / or the record content corresponding to these object contents.
[0124] At least one embodiment of the present disclosure further provides an image processing device. Figure 5 A schematic diagram of an image processing device provided by at least one embodiment of the present disclosure.
[0125] like Figure 5 As shown, the image processing device 500 may include: an image acquisition module 501, an identification module 502, an object block determination module 503, a table generation module 504, a cell determination module 505, a record content determination module 506, and an output module 507. For example, these modules (i.e., the image acquisition module 501, the identification module 502, the object block determination module 503, the table generation module 504, the cell determination module 505, the record content determination module 506, and the output module 507) may be implemented by hardware (e.g., circuit) modules, software modules, or any combination of the two. The following embodiments are the same and will not be repeated. For example, these modules may be implemented by a central processing unit (CPU), a graphics processing unit (GPU), a tensor processing unit (TPU), a field programmable gate array (FPGA), or other forms of processing units with data processing capabilities and / or instruction execution capabilities and corresponding computer instructions.
[0126] For example, the image acquisition module 501 is configured to acquire an input image. For example, the input image may include a table.
[0127] For example, the recognition module 502 is configured to perform recognition processing on the input image to obtain a plurality of object blocks. For example, each object block includes an object content, and the object content may include text, numbers, graphics, symbols, etc.
[0128] For example, the object block determination module 503 is configured to determine a plurality of key object blocks among a plurality of object blocks.
[0129] For example, the table generation module 504 is configured to generate a unit table based on a plurality of block positions corresponding to a plurality of object blocks. For example, the unit table includes a plurality of cells, the plurality of cells include a plurality of object cells, and each object cell includes an object block.
[0130] For example, the cell determination module 505 is configured to determine N key object cells among the plurality of object cells based on the plurality of key object blocks. For example, each of the N key object cells includes a key object block, and in the input image, the N key object cells are arranged in the same row along the first direction.
[0131] For example, the record content determination module 506 is configured to determine at least one record content corresponding to N object contents in N key object blocks in N key object cells based on multiple object contents in multiple object blocks. For example, each record content includes at least one object content.
[0132] For example, the output module 507 is configured to output M object contents among N object contents and / or output L record contents among at least one record content.
[0133] For example, N, M, and L are positive integers, and N is greater than 1.
[0134] For example, the image acquisition module 501, the recognition module 502, the object block determination module 503, the table generation module 504, the cell determination module 505, the record content determination module 506 and / or the output module 507 may include codes and programs stored in the memory; the processor may execute the codes and programs to implement some or all functions of the above-mentioned image acquisition module 501, the recognition module 502, the object block determination module 503, the table generation module 504, the cell determination module 505, the record content determination module 506 and / or the output module 507. For example, the image acquisition module 501, the recognition module 502, the object block determination module 503, the table generation module 504, the cell determination module 505, the record content determination module 506 and / or the output module 507 may be dedicated hardware devices for implementing the above-mentioned functions. For example, the image acquisition module 501, the recognition module 502, the object block determination module 503, the table generation module 504, the cell determination module 505, the record content determination module 506 and / or the output module 507 may be a circuit board or a combination of multiple circuit boards, for implementing the functions described above. In the embodiment of the present application, the circuit board or the combination of multiple circuit boards may include: (1) one or more processors; (2) one or more non-temporary memories connected to the processor; and (3) firmware stored in the memory and executable by the processor.
[0135] It should be noted that the image acquisition module 501 can be used to implement Figure 1 In step S10 shown in FIG. 1 , the identification module 502 may be used to implement Figure 1 In step S11 shown in FIG. 1 , the object block determination module 503 may be used to implement Figure 1 In step S12 shown in FIG. 1 , the table generation module 504 can be used to implement Figure 1 In step S13 shown in FIG. 1 , the cell determination module 505 can be used to implement Figure 1 In step S14, the record content determination module 506 can be used to implement Figure 1 In step S15, the output module 507 can be used to implement Figure 1 Step S16 shown. Therefore, the specific description of the functions that can be realized by the image acquisition module 501, the recognition module 502, the object block determination module 503, the table generation module 504, the cell determination module 505, the record content determination module 506 and the output module 507 can refer to the relevant description of steps S10 to S16 in the embodiment of the above-mentioned image processing method, and the repeated parts will not be repeated. In addition, the image processing device 500 can achieve technical effects similar to those of the above-mentioned image processing method, which will not be repeated here.
[0136] It should be noted that in the embodiments of the present disclosure, the image processing device 500 may include more or fewer circuits or modules, and the connection relationship and specific configuration between the various circuits or modules are not limited and can be determined according to actual needs.
[0137] At least one embodiment of the present disclosure further provides an electronic device, Figure 6 A schematic diagram of an electronic device provided according to at least one embodiment of the present disclosure.
[0138] For example, Figure 6 As shown, the electronic device 600 includes a memory 601 and a processor 602 .
[0139] For example, the memory 601 is used to store computer executable instructions non-transiently. When the processor 602 is used to execute the computer executable instructions, the image processing method according to any of the above embodiments is implemented. The specific implementation of each step of the image processing method and the related explanation content can be referred to the embodiment of the above image processing method, which will not be repeated here.
[0140] For example, other implementations of the image processing method implemented by the processor 602 executing the computer executable instructions stored in the memory 601 are the same as the implementations mentioned in the aforementioned method embodiment part, and will not be repeated here.
[0141] For example, in some embodiments, the electronic device 600 also includes a communication interface and a communication bus. The processor 602, the communication interface and the memory 601 communicate with each other through the communication bus, and the components such as the processor 602, the communication interface, the memory 601 can also communicate through a network connection. The present disclosure does not limit the type and function of the network. For example, the communication bus can be a peripheral component interconnect standard (PCI) bus or an extended industrial standard architecture (EISA) bus. The communication bus can be divided into an address bus, a data bus, a control bus, etc. The communication interface is used to realize communication between the electronic device 600 and other devices. It should be noted that Figure 6 The components of the electronic device 600 shown are merely exemplary and non-limiting. The electronic device may also have other components according to actual application requirements.
[0142] For example, the processor 602 and the memory 601 may be arranged on a server side (or a cloud side).
[0143] For example, the processor 602 can control other components in the electronic device 600 to perform the desired functions. The processor 602 can be a device with data processing capabilities and / or program execution capabilities, such as a central processing unit (CPU), a network processor (NP), a tensor processing unit (TPU), or a graphics processing unit (GPU); it can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components. The central processing unit (CPU) can be an X86 or ARM architecture, etc.
[0144] For example, the memory 601 may include any combination of one or more computer program products, and the computer program product may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. Volatile memory may include, for example, random access memory (RAM) and / or cache memory (cache), etc. Non-volatile memory may include, for example, read-only memory (ROM), hard disk, erasable programmable read-only memory (EPROM), portable compact disk read-only memory (CD-ROM), USB memory, flash memory, etc. One or more computer-executable instructions may be stored on the computer-readable storage medium, and the processor 602 may execute the computer-executable instructions to implement various functions of the electronic device 600. Various applications and various data may also be stored in the storage medium.
[0145] For example, in some embodiments, the electronic device 600 may further include an image acquisition component. The image acquisition component is used to acquire an image, for example, an input image. The memory 601 is also used to store the acquired input image. For example, the image acquisition component may be a camera of a smart phone, a camera of a tablet computer, a camera of a personal computer, a lens of a digital camera, or even a webcam. The image acquisition component may also be a scanner, etc.
[0146] For example, for a detailed description of the process of the electronic device 600 performing image processing, reference may be made to the relevant description in the embodiment of the image processing method, and repeated parts will not be repeated.
[0147] Figure 7 A schematic diagram of a non-transitory computer-readable storage medium provided for at least one embodiment of the present disclosure. Figure 7 As shown, the non-transitory computer-readable storage medium 700 may non-transitorily store one or more computer-executable instructions 701. For example, when the computer-executable instructions 701 are executed by a processor, one or more steps in the image processing method described above may be executed.
[0148] For example, the non-transitory computer-readable storage medium 700 may be applied to the above-mentioned electronic device, for example, the non-transitory computer-readable storage medium 700 may include a memory 601 in the electronic device 600. For the description of the non-transitory computer-readable storage medium 700, reference may be made to the description of the memory 601 in the embodiment of the electronic device 600, and the repeated parts will not be repeated.
[0149] There are a few points to note about this disclosure:
[0150] (1) The drawings of the embodiments of the present disclosure only relate to the structures related to the embodiments of the present disclosure, and other structures may refer to the general design.
[0151] (2) In the absence of conflict, the embodiments of the present disclosure and the features therein may be combined with each other to obtain new embodiments.
[0152] The above description is only a specific implementation of the present disclosure, but the protection scope of the present disclosure is not limited thereto. The protection scope of the present disclosure shall be based on the protection scope of the claims.
Claims
1. An image processing method, comprising: Get the input image; Performing recognition processing on the input image to obtain a plurality of object blocks, wherein each object block includes an object content; determining a plurality of key object blocks among the plurality of object blocks; Generate a unit table based on a plurality of block positions respectively corresponding to the plurality of object blocks, wherein the unit table includes a plurality of cells, the plurality of cells include a plurality of object cells, and each object cell includes an object block; Based on the multiple key object blocks, determine N key object cells among the multiple object cells, wherein each key object cell among the N key object cells includes a key object block, and in the input image, the N key object cells are arranged in the same row along a first direction; Based on the plurality of object contents in the plurality of object blocks, determining at least one record content corresponding to the N object contents in the N key object blocks in the N key object cells, wherein each record content includes at least one object content; outputting M object contents among the N object contents and / or outputting L record contents among the at least one record content, Wherein, N, M and L are positive integers, and N is greater than 1; Wherein, determining a plurality of key object blocks among the plurality of object blocks comprises: Using an object classification model to classify the plurality of object blocks according to object attributes, so as to obtain the plurality of object blocks and a plurality of object attributes corresponding to the plurality of object blocks one by one; Based on the plurality of object attributes, determining the plurality of key object blocks from the plurality of object blocks; The step of generating a unit table based on the plurality of block positions respectively corresponding to the plurality of object blocks includes: Determining a plurality of separation lines based on the plurality of block positions, wherein there is at least one separation line between every two object blocks; The plurality of object blocks are separated by the plurality of separation lines to form the unit table.
2. The image processing method according to claim 1, wherein: Each object block has an object attribute, and the object attribute corresponding to each key object block is any key object attribute in the key attribute group.
3. The image processing method according to claim 2, wherein: Outputting M object contents among the N object contents includes: Determine the selected output information; Determine and output the object content corresponding to the selected output information among the N object contents, wherein the M object contents include the object content corresponding to the selected output information among the N object contents.
4. The image processing method according to claim 3, wherein: The selected output information includes an image type corresponding to the input image, the key attribute group includes a plurality of key object attributes, Determining and outputting object content corresponding to the selected output information among the N object contents includes: determining at least one key object attribute among the plurality of key object attributes corresponding to the image type; Determine at least one key object block whose object attribute in the N key object blocks is any one of the at least one key object attribute; Outputting the object content in the at least one key object block, wherein the object content corresponding to the selected output information is the object content in the at least one key object block.
5. The image processing method according to claim 3, wherein: The selected output information includes information predefined by a user.
6. The image processing method according to any one of claims 2 to 5, wherein: Outputting L record contents of the at least one record content includes: Determine the selected output information; Determine and output the record content corresponding to the selected output information in the at least one record content, wherein the L record contents include the record content corresponding to the selected output information in the at least one record content.
7. The image processing method according to claim 6, wherein: The selected output information includes an image type corresponding to the input image, the key attribute group includes a plurality of key object attributes, Determining and outputting a record content corresponding to the selected output information in the at least one record content includes: determining at least one key object attribute among the plurality of key object attributes corresponding to the image type; Determine at least one key object block whose object attribute in the N key object blocks is any one of the at least one key object attribute; Output the record content in the at least one record content corresponding to the object content in the at least one key object block, wherein the record content corresponding to the selected output information is the record content in the at least one record content corresponding to the object content in the at least one key object block.
8. The image processing method according to any one of claims 1 to 5, wherein: Determining N key object cells among the plurality of object cells based on the plurality of key object blocks includes: Based on the multiple key object blocks, determining multiple key object cells among the multiple object cells, wherein the multiple key object cells are object cells among the multiple object cells that correspond one-to-one to the multiple key object blocks; Based on the positions of the plurality of key object cells, the N key object cells among the plurality of key object cells are determined.
9. The image processing method according to any one of claims 1 to 5, wherein: The plurality of cells further include a plurality of blank cells, each blank cell does not include an object block, Determining at least one record content corresponding to the N object contents in the N key object blocks in the N key object cells based on the multiple object contents in the multiple object blocks includes: For the i-th key object cell among the N key object cells: Acquire P cells located on a first side of the i-th key object cell in the second direction; In response to the P cells including at least one object cell, determining, based on the P cells, record content corresponding to the object content in the key object block in the i-th key object cell; In response to the P cells being blank cells, it is determined that the key object block in the i-th key object cell has no corresponding record content.
10. The image processing method according to claim 9, wherein: In the input image, an edge of any one of the P cells in the first direction does not exceed an edge of the i-th key object cell in the first direction.
11. The image processing method according to claim 9, wherein: In response to the P cells including at least one object cell, determining, based on the P cells, record content corresponding to the object content in the key object block in the i-th key object cell, comprises: In response to P being 1 and the P cells being object cells, taking the object contents in the object blocks in the P cells as the record contents corresponding to the object contents in the key object blocks in the i-th key object cell; In response to P being greater than 1: Merging the P cells to obtain at least one merged cell corresponding to the i-th key object cell, wherein each merged cell includes at least one object block, and the merged content corresponding to each merged cell is the object content in the at least one object block; Based on at least one merged content corresponding to each of the at least one merged cell, a record content corresponding to the object content in the key object block in the i-th key object cell is determined.
12. The image processing method according to claim 11, wherein: The object content in each object block includes text, and the object attributes of each object block include part of speech. The merging process includes: Based on the positions of the P cells, at least one cell group is determined, wherein each cell group includes at least one cell among the P cells, and in the case where any cell group includes a plurality of cells, the plurality of cells in the any cell group are arranged in a row along the first direction; Based on the at least one cell group, at least one intermediate merged cell corresponding to the at least one cell group is determined, wherein, in the case where any cell group includes multiple cells, the multiple cells in the any cell group are merged as the intermediate merged cell corresponding to the any cell group, and in the case where any cell group includes one cell, one cell in the any cell group is directly used as the intermediate merged cell corresponding to the any cell group; In response to the at least one intermediate merged cell including a plurality of intermediate merged cells, if a first intermediate merged cell and a second intermediate merged cell to be merged among the plurality of intermediate merged cells meet a merging condition, merging the first intermediate merged cell and the second intermediate merged cell, In response to the at least one intermediate merged cell including one intermediate merged cell, the one intermediate merged cell is treated as a merged cell.
13. The image processing method according to claim 12, wherein: The merger conditions include: The first intermediate merged cell and the second intermediate merged cell both include object blocks, and object content in the object block in the first intermediate merged cell is the same as or semantically continuous with object content in the object block in the second intermediate merged cell, and in the input image, the first intermediate merged cell and the second intermediate merged cell are arranged in sequence and continuously in the second direction; or, The first intermediate merged cell and / or the second intermediate merged cell do not include an object block, and in the input image, the first intermediate merged cell and the second intermediate merged cell are arranged in sequence and continuously in the second direction.
14. The image processing method according to claim 9, wherein: In the input image, the first direction and the second direction are perpendicular to each other.
15. The image processing method according to claim 11, wherein: Determining the record content corresponding to the object content in the key object block in the i-th key object cell based on at least one merged content corresponding to each of the at least one merged cell includes: Get filtering rules; Based on the filtering rule, the at least one merged content is filtered to determine the record content corresponding to the object content in the key object block in the i-th key object cell.
16. The image processing method according to any one of claims 2 to 5, wherein: The input image is subjected to recognition processing to obtain a plurality of object blocks, including: Using an object recognition model to perform recognition processing on the input image to obtain a plurality of initial object blocks; The object classification model is used to classify the multiple initial object blocks according to object attributes to obtain the multiple object blocks.
17. An image processing device, comprising: An image acquisition module configured to acquire an input image; A recognition module configured to perform recognition processing on the input image to obtain a plurality of object blocks, wherein each object block includes an object content; an object block determination module, configured to determine a plurality of key object blocks among the plurality of object blocks; A table generation module, configured to generate a unit table based on a plurality of block positions respectively corresponding to the plurality of object blocks, wherein the unit table includes a plurality of cells, the plurality of cells include a plurality of object cells, and each object cell includes an object block; a cell determination module configured to determine N key object cells among the plurality of object cells based on the plurality of key object blocks, wherein each of the N key object cells includes a key object block, and in the input image, the N key object cells are arranged in the same row along a first direction; a record content determination module configured to determine, based on a plurality of object contents in the plurality of object blocks, at least one record content corresponding to N object contents in N key object blocks in the N key object cells, wherein each record content includes at least one object content; an output module, configured to output M object contents among the N object contents and / or output L record contents among the at least one record content, Wherein, N, M and L are positive integers, and N is greater than 1; Wherein, when determining a plurality of key object blocks among the plurality of object blocks, the object block determining module is configured as follows: Using an object classification model to classify the plurality of object blocks according to object attributes, so as to obtain the plurality of object blocks and a plurality of object attributes corresponding to the plurality of object blocks one by one; Based on the plurality of object attributes, determining the plurality of key object blocks from the plurality of object blocks; Wherein, when generating a unit table based on a plurality of block positions respectively corresponding to the plurality of object blocks, the table generating module is configured as follows: Determining a plurality of separation lines based on the plurality of block positions, wherein there is at least one separation line between every two object blocks; The plurality of object blocks are separated by the plurality of separation lines to form the unit table.
18. An electronic device, comprising: A memory non-transitorily stores computer executable instructions; a processor configured to execute the computer executable instructions, Wherein, when the computer executable instructions are executed by the processor, the image processing method according to any one of claims 1-16 is implemented.
19. A non-transitory computer-readable storage medium, wherein: The non-transitory computer-readable storage medium stores computer-executable instructions, and when the computer-executable instructions are executed by a processor, the image processing method according to any one of claims 1-16 is implemented.
Citation Information
Patent Citations
Image processing method and device, electronic equipment and storage medium
CN112906532A
Image processing method and device, electronic equipment and storage medium
CN112926421A