Image processing method and device, equipment and storage medium

By identifying the three parts of a bank credit document and generating a mask image, the problem of low-quality images affecting OCR processing was solved, thereby improving image quality and processing efficiency.

CN120976026APending Publication Date: 2025-11-18AGRICULTURAL BANK OF CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511094844.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-06
Publication Date
2025-11-18

AI Technical Summary

Technical Problem

In the process of scanning bank credit documents, low-quality images affect the OCR image processing effect, leading to misreading or loss of key information, which affects the efficiency and accuracy of business processing.

Method used

By acquiring the original document image, a three-part image (foreground image, background image, and unknown area image) is determined. Based on the three-part image and the original document image, a mask image is determined and processed to obtain the target document image.

Benefits of technology

It improves image quality and enhances the efficiency and accuracy of target document image processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120976026A_ABST
    Figure CN120976026A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides an image processing method and device, equipment and a storage medium. The method comprises the following steps: acquiring an original document image; determining a tripartite graph corresponding to the original document image; wherein the tripartite graph comprises a foreground graph, a background graph and an unknown region graph; based on the tripartite graph and the original document image, determining a mask graph corresponding to the original document image; and processing the mask image to obtain a target document image. According to the embodiment of the invention, by determining the tripartite graph corresponding to the original document image, determining the mask graph corresponding to the original document image based on the tripartite graph and the original document image, and then processing the mask graph to obtain the target document image, the original document image can be effectively enhanced, the image quality is improved, and the user experience is improved. And thus, the efficiency and accuracy of target document image processing are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular to an image processing method, apparatus, device, and storage medium. Background Technology

[0002] In today's information age, bank credit documents, as crucial credentials for financial transactions, have become an inevitable trend in digital processing. Modern banking lending often requires image processing techniques such as Optical Character Recognition (OCR) to extract key information from document images. However, during the scanning and photographing of documents, limitations imposed by the shooting equipment, environment, and the document's inherent quality frequently result in low-quality scans. Such low-quality images negatively impact the effectiveness of subsequent OCR and other image processing, leading to misreading or loss of crucial information and affecting the efficiency and accuracy of business operations. Summary of the Invention

[0003] This application provides an image processing method, apparatus, device, and storage medium to improve image quality.

[0004] In a first aspect, embodiments of this application provide an image processing method, comprising: acquiring an original document image; determining a tripartite image corresponding to the original document image; wherein the tripartite image includes a foreground image, a background image, and an unknown region image; determining a mask image corresponding to the original document image based on the tripartite image and the original document image; and processing the mask image to obtain a target document image.

[0005] Secondly, embodiments of this application also provide an image processing apparatus, including: an original document image acquisition module for acquiring an original document image; a three-part image determination module for determining a three-part image corresponding to the original document image; wherein the three-part image includes a foreground image, a background image, and an unknown region image; a mask image determination module for determining a mask image corresponding to the original document image based on the three-part image and the original document image; and a processing module for processing the mask image to obtain a target document image.

[0006] Thirdly, embodiments of this application also provide an electronic device, the electronic device comprising: one or more processors; and a storage device for storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors implement the image processing method as described in embodiments of this application.

[0007] Fourthly, embodiments of this application also provide a storage medium containing computer-executable instructions, which, when executed by a computer processor, are used to perform the image processing method as described in embodiments of this application.

[0008] The technical solution of this application embodiment involves acquiring an original document image; determining a three-part image corresponding to the original document image; wherein the three-part image includes a foreground image, a background image, and an unknown region image; determining a mask image corresponding to the original document image based on the three-part image and the original document image; and processing the mask image to obtain a target document image. This application embodiment, by determining the three-part image corresponding to the original document image, determining the mask image corresponding to the original document image based on the three-part image and the original document image, and then processing the mask image to obtain the target document image, can effectively enhance the original document image, improve image quality, and thus improve the efficiency and accuracy of target document image processing. Attached Figure Description

[0009] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and the originals and elements are not necessarily drawn to scale.

[0010] Figure 1 This is a schematic flowchart of an image processing method provided in an embodiment of this application;

[0011] Figure 2 This is a schematic diagram of another image processing method provided in an embodiment of this application;

[0012] Figure 3 This is a schematic diagram of the structure of an image processing device provided in an embodiment of this application;

[0013] Figure 4 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0014] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.

[0015] It should be understood that the various steps described in the method embodiments of this disclosure may be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of this disclosure is not limited in this respect. The term "comprising" and its variations as used herein are open-ended, meaning "including but not limited to". It should be noted that the concepts of "first," "second," etc., mentioned in this disclosure are used only to distinguish different devices, modules, or units, and are not intended to limit the order of functions performed by these devices, modules, or units or their interdependencies. It should be noted that the modifications "a" and "a plurality" mentioned in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that unless explicitly indicated otherwise in the context, they should be understood as "one or more". It is understood that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of data) shall comply with the requirements of applicable laws, regulations, and relevant provisions.

[0016] Figure 1 This is a schematic flowchart of an image processing method provided in an embodiment of this application, applicable to the processing of document images. The method can be executed by an image processing device, which can be implemented in software and / or hardware, optionally through an electronic device, such as a mobile terminal, PC, or server. Figure 1 As shown, the method includes:

[0017] S110. Obtain the original document image.

[0018] The original document image can be in a format conforming to encoding standards, such as JPEG, PNG, or PSD. The content of the original document image can be any content that complies with laws, regulations, platform rules, and social ethics. In this embodiment, the original document image can be obtained by scanning a paper credit certificate.

[0019] S120. Determine the triangulation corresponding to the original document image.

[0020] The three-part diagram includes a foreground diagram, a background diagram, and an unknown area diagram.

[0021] In this embodiment, the tri-image corresponding to the original document image can be obtained in any way. For example, it can be manually drawn using image editing software, or it can be obtained using edge detection algorithms or deep learning methods. In this embodiment, after obtaining the foreground image and the background image, the image of the remaining pixels is the unknown region image.

[0022] It should be noted that before obtaining the triangulation corresponding to the original document image, the original document image can be converted to grayscale to obtain a grayscale image, and then the triangulation corresponding to the grayscale image can be determined. In the following embodiments, the original document images are all grayscale converted images.

[0023] Optionally, determining the foreground image corresponding to the original document image includes: sequentially performing image erosion and image dilation operations on the original document image to obtain a noisy image; and subtracting the original document image from the noisy image to obtain the foreground image corresponding to the original document image.

[0024] In this embodiment, an image opening operation is performed on the original document image, that is, an image erosion operation is first performed on the original document image, followed by an image dilation operation. After the opening operation, the text to be identified is eliminated, while the patchy smudges are retained. Therefore, the foreground image in the three-part image is obtained by subtracting the original document image from the image after the opening operation. Specifically, since the font in the original document image has relatively thin edges, and most of the noise in the original document image is patchy, such as ink droplets, the font in the original document image can be removed by the image erosion operation, and the edges of the smudges removed by the erosion operation can be restored by the image dilation operation, resulting in a document image containing only noise, denoted as the noise image. Subtracting the original document image from the noise image yields an image containing only font. The black pixels in the image containing only font are used as the foreground image in the three-part image.

[0025] Optionally, determining the background image corresponding to the original document image includes: determining the pixel with the highest grayscale value in the original document image and adding it to the background pixel set and adding it as a seed point to be processed to the seed point queue; cyclically executing the following steps: retrieving a seed point to be processed from the seed point queue; traversing the neighboring pixels of the seed point to be processed; if the difference between the grayscale value of the neighboring pixel and the grayscale value of the seed point to be processed is less than a set difference threshold, then adding the corresponding neighboring pixel to the background pixel set and adding it as a seed point to be processed to the seed point queue; when the seed point queue is empty, the loop ends, and the latest set of background pixels is used as the background image corresponding to the original document image.

[0026] In this embodiment, in the original document image, the white background is usually composed of pixels with the highest grayscale values. Therefore, we can start with the pixels with the highest grayscale values ​​and use a flood fill algorithm to obtain the background image of the tripartite image. The specific process is as follows: find the pixel with the highest grayscale value in the original document image and add it to the background pixel set and add it as a seed point to be processed to the seed point queue; the seed point queue is characterized by first-in, first-out.

[0027] The following steps are executed in a loop: take a seed point to be processed from the seed point queue as the starting point; traverse the neighboring pixels of the seed point to be processed; wherein, at this time, the neighboring pixels have not yet been added to the background pixel set, and the neighboring pixels can be 4-neighborhood (top, bottom, left, right of the seed point to be processed) or 8-neighborhood (top, bottom, left, right, top left, top right, bottom left, bottom right of the seed point to be processed).

[0028] If the difference between the grayscale value of the adjacent pixel and the grayscale value of the seed point to be processed is less than a set difference threshold, then the corresponding adjacent pixel is added to the background pixel set and added as a seed point to be processed to the seed point queue. In this embodiment, the set difference threshold is not limited; for example, it can be 5% of the grayscale value of the seed point to be processed.

[0029] When the queue of seed points to be processed is empty, the loop ends, and the latest set of background pixels is used as the background image corresponding to the original document image.

[0030] In this embodiment, the background image is obtained by using a flood filling algorithm, which can quickly and accurately identify the background image.

[0031] S130. Determine the mask image corresponding to the original document image based on the three-part image and the original document image.

[0032] The mask image can be an image with a transparency mask value.

[0033] In this embodiment, a matting algorithm can be used to determine the mask image corresponding to the original document image based on the three-part image and the original document image. The matting algorithm can be a supervised deep learning algorithm or an unsupervised deep learning algorithm. Unlike supervised deep learning algorithms, unsupervised deep learning algorithms do not require the labeling information of the input dataset. Instead, they discover the potential structure and patterns of the data through operations such as clustering, dimensionality reduction, and association rule mining. There are three main types of unsupervised matting algorithms: propagation-based matting algorithms, sampling-based matting algorithms, and evolutionary optimization-based matting algorithms. In this embodiment, a sampling-based matting algorithm is preferred.

[0034] This embodiment employs an unsupervised image matting algorithm, which does not require labeled data and can effectively process the original document image while ensuring the confidentiality of user data.

[0035] S140. Process the mask image to obtain the target document image.

[0036] In this embodiment, the mask image is binarized to obtain the target document image.

[0037] The technical solution of this application embodiment involves acquiring an original document image; determining a three-part image corresponding to the original document image; wherein the three-part image includes a foreground image, a background image, and an unknown region image; determining a mask image corresponding to the original document image based on the three-part image and the original document image; and processing the mask image to obtain a target document image. This application embodiment, by determining the three-part image corresponding to the original document image, determining the mask image corresponding to the original document image based on the three-part image and the original document image, and then processing the mask image to obtain the target document image, can effectively enhance the original document image, improve image quality, and thus improve the efficiency and accuracy of target document image processing.

[0038] Figure 2 This is a schematic flowchart of another image processing method provided in an embodiment of this application. The embodiments of this application are specific modifications based on the above-described embodiments of the invention. See also... Figure 2 The method provided in this application specifically includes the following steps:

[0039] S201. Obtain the original document image.

[0040] S202. Determine the trisection corresponding to the original document image.

[0041] The three-part diagram includes a foreground diagram, a background diagram, and an unknown area diagram.

[0042] S203. For each unknown pixel in the unknown region map, obtain the observed color value based on the original document image.

[0043] It should be noted that for each unknown pixel in the unknown region map, steps S203, S204, S205, S206, and S207 need to be performed. After all unknown pixels have been processed, step S208 is performed.

[0044] The observed color value is the actual observed color value, which can be a vector in a color space such as RGB or Lab. The observed color value is directly derived from the color data of the unknown pixels corresponding to the original document image, so the observed color value corresponding to the unknown pixels can be directly read from the original document image.

[0045] S204. In the foreground image, sample the foreground color of the unknown pixel.

[0046] In this embodiment, the coordinate distance and texture feature similarity between each foreground pixel and the unknown pixel in the foreground image can be calculated. Then, the Pareto front of each foreground pixel can be determined based on the coordinate distance and texture feature similarity. The Pareto fronts of all foreground pixels in the foreground image can be used as the foreground color samples of the unknown pixel.

[0047] Optionally, in the foreground image, sampling the foreground color sample of the unknown pixel includes: determining a first pixel coordinate distance between each foreground pixel in the foreground image and the unknown pixel; determining a first texture feature similarity between each foreground pixel in the foreground image and the unknown pixel; determining the Pareto front of all foreground pixels based on the first pixel coordinate distance and the first texture feature similarity; and using the Pareto front of all foreground pixels in the foreground image as the foreground color sample of the unknown pixel.

[0048] For example, the first pixel coordinate distance between each foreground pixel in the foreground image and the unknown pixel can be calculated using the following formula:

[0049]

[0050] Where g1 is the coordinate distance of the first pixel, z is the unknown pixel, and F is the foreground image. S represents the foreground pixel in the foreground image. z For the spatial coordinates of the unknown pixel, Foreground pixels The spatial coordinates are ||*||2, which is the L2 norm.

[0051] The first texture feature similarity between each foreground pixel in the foreground image and the unknown pixel can be calculated using the following formula:

[0052]

[0053] Where g2 is the first texture feature similarity, T z For the local texture feature vector of the unknown pixel, Foreground pixels The local texture feature vector.

[0054] In this embodiment, the Pareto front of the foreground pixel can be obtained in the following way:

[0055]

[0056] Where P is the Pareto front of the foreground pixels, and x and y are different foreground pixels in F. In this formula, the Pareto front means that if a solution x in the solution space has no other solution whose g1 and g2 values ​​are both better than x, then solution x is called a Pareto solution, and the set of all Pareto solutions is the Pareto front of the foreground pixels (each foreground pixel corresponds to one Pareto front).

[0057] S205. In the background image, sample the background color of the unknown pixel.

[0058] In this embodiment, the coordinate distance and texture feature similarity between each background pixel in the background image and the unknown pixel can be calculated. Then, the Pareto front of each background pixel can be determined based on the coordinate distance and texture feature similarity. The Pareto fronts of all background pixels in the background image can be used as the background color samples of the unknown pixel.

[0059] Optionally, in the background image, sampling the background color sample of the unknown pixel includes: determining the second pixel coordinate distance between each background pixel in the background image and the unknown pixel; determining the second texture feature similarity between each background pixel in the background image and the unknown pixel; determining the Pareto front of all background pixels based on the second pixel coordinate distance and the second texture feature similarity; and using the Pareto front of all background pixels in the background image as the background color sample of the unknown pixel.

[0060] In this embodiment, the background color sample can be obtained in the same way as the foreground color sample.

[0061] For example, the second pixel coordinate distance between each background pixel in the background image and the unknown pixel can be calculated using the following formula:

[0062]

[0063] Where g3 is the distance to the second pixel coordinates, z is the unknown pixel, and B is the background image. S represents the background pixels in the background image. z For the spatial coordinates of the unknown pixel, background pixels The spatial coordinates are ||*||2, which is the L2 norm.

[0064] The similarity of the second texture feature between each background pixel in the background image and the unknown pixel can be calculated using the following formula:

[0065]

[0066] Where g4 is the second texture feature similarity, T z For the local texture feature vector of the unknown pixel, background pixels The local texture feature vector.

[0067] In this embodiment, the Pareto front of the background pixels can be obtained in the following way:

[0068]

[0069] Where P1 is the Pareto front of the background pixels, and x1 and y1 are different background pixels in B. In this formula, the Pareto front means that if a solution x1 in the solution space has no other solution whose g3 and g4 values ​​are better than x1, then the solution x1 is called a Pareto solution, and the set of all Pareto solutions is the Pareto front of the background pixels (each background pixel corresponds to one Pareto front).

[0070] In this embodiment, by using the Pareto fronts of all foreground pixels in the foreground image as the foreground color samples of the unknown pixel and the Pareto fronts of all background pixels in the background image as the background color samples of the unknown pixel, the subsequent evaluation function can quickly find the optimal sample pair combination among the foreground color samples and the background color samples, thereby improving search efficiency.

[0071] S206. Determine the optimal sample pair combination corresponding to the unknown pixel based on the evaluation function constructed based on the foreground color sample and the background color sample.

[0072] Here, a sample pair combination is a combination of foreground pixels from the foreground color sample and background pixels from the background color sample. The optimal sample pair combination can be the optimal solution for the evaluation function.

[0073] Optionally, the evaluation function is constructed as follows: the evaluation function is determined based on the spatial proximity term and texture similarity term of the foreground color sample and the background color sample; wherein, the spatial proximity term is obtained based on the spatial coordinates of the unknown pixel, the foreground color sample and the background color sample; and the texture similarity term is obtained based on the local texture feature vectors of the unknown pixel, the foreground color sample and the background color sample.

[0074] For example, the specific formula for the evaluation function is as follows:

[0075]

[0076] Where z represents any unknown pixel in the unknown region map, U1, F1, and B1 represent the unknown region map, the foreground color sample, and the background color sample, respectively, and g z (x z ) represents the foreground pixel in the foreground color sample. Background pixels in background color sample For unknown pixel x z The evaluation function.

[0077]

[0078] in Indicates spatially adjacent terms, To represent texture similarity items, the two are represented as follows:

[0079]

[0080] Among them, S z , Representing the unknown pixel z and the foreground pixel in the foreground color sample, respectively. Background pixels in background color sample The spatial coordinates, ||*||2 represent the L2 norm.

[0081]

[0082] Where T z , Representing the unknown pixel z and the foreground pixel in the foreground color sample, respectively. Background pixels in background color sample The local texture feature vector.

[0083] At this point, an evaluation function can be found for each unknown pixel. The optimal combination of sample pairs.

[0084] S207. Determine the transparency masking value of the unknown pixel based on the optimal sample pair combination and the observed color value.

[0085] For example, the formula for calculating the transparency mask value of an unknown pixel is:

[0086]

[0087] Where ||*|| represents the modulus of the orientation quantity, α z I represents the transparency mask value for the unknown pixel Z. z For the observed color value of the unknown pixel Z, F1 z B1 represents the color of the foreground pixel in the optimal sample pair combination for the unknown pixel Z. z The color of the background pixel in the optimal sample pair combination for the unknown pixel Z.

[0088] S208. Determine the mask image corresponding to the original document image based on the transparency mask value of each foreground pixel in the foreground image, the transparency mask value of each background pixel in the background image, and the transparency mask value of each unknown pixel in the unknown region image.

[0089] It should be noted that each foreground pixel in the foreground image is a known foreground pixel (white), and the opacity mask value of each foreground pixel in the foreground image is 1. Each background pixel in the background image is a known background pixel (black), and the opacity mask value of each background pixel in the background image is 0. The opacity mask value of each unknown pixel (gray) is between [0,1].

[0090] S209. Process the mask image to obtain the target document image.

[0091] In this embodiment, the following steps are taken: First, an original document image is acquired. Second, a tripartite image corresponding to the original document image is determined. Third, for each unknown pixel in the unknown region image, an observed color value is obtained based on the original document image. Fourth, in the foreground image, a foreground color sample of the unknown pixel is sampled. Fifth, in the background image, a background color sample of the unknown pixel is sampled. Sixth, an evaluation function constructed based on the foreground and background color samples is used to determine the optimal sample pair combination corresponding to the unknown pixel. Seventh, based on the optimal sample pair combination and the observed color value, a transparency masking value for the unknown pixel is determined. Eighth, a masking image corresponding to the original document image is determined based on the transparency masking values ​​of each foreground pixel in the foreground image, each background pixel in the background image, and each unknown pixel in the unknown region image. Finally, the masking image is processed to obtain the target document image. In this embodiment, by sampling the foreground color samples and background color samples of the unknown pixel, determining the optimal sample pair combination corresponding to the unknown pixel based on the evaluation function constructed based on the foreground color samples and the background color samples, and determining the transparency mask value of the unknown pixel based on the optimal sample pair combination and the observed color value, the transparency mask value of the unknown pixel can be obtained quickly and accurately, thereby quickly obtaining the mask image and improving the accuracy of the mask image.

[0092] Figure 3 This is a schematic diagram of the structure of an image processing device provided in an embodiment of this application, as shown below. Figure 3 As shown, the device includes: an original document image acquisition module 310, used to acquire an original document image;

[0093] The tripartite determination module 320 is used to determine the tripartite corresponding to the original document image; wherein the tripartite includes a foreground image, a background image, and an unknown region image;

[0094] The mask image determination module 330 is used to determine the mask image corresponding to the original document image based on the three-part image and the original document image.

[0095] The processing module 340 is used to process the mask image to obtain the target document image.

[0096] The technical solution of this application embodiment involves acquiring an original document image through an original document image acquisition module; determining a three-part image corresponding to the original document image through a three-part image determination module; wherein the three-part image includes a foreground image, a background image, and an unknown region image; determining a mask image corresponding to the original document image through a mask image determination module based on the three-part image and the original document image; and processing the mask image through a processing module to obtain a target document image. This application embodiment, by determining the three-part image corresponding to the original document image, determining the mask image corresponding to the original document image based on the three-part image and the original document image, and then processing the mask image to obtain the target document image, can effectively enhance the original document image, improve image quality, and thus improve the efficiency and accuracy of target document image processing.

[0097] Optionally, the three-part image determination module is specifically used to: sequentially perform image erosion and image dilation operations on the original document image to obtain a noisy image; and subtract the original document image from the noisy image to obtain the foreground image corresponding to the original document image.

[0098] Optionally, the tripartite determination module is further configured to: determine the pixel with the highest grayscale value in the original document image and add it to the background pixel set and add it as a seed point to be processed to the seed point queue; cyclically execute the following steps: retrieve a seed point to be processed from the seed point queue; traverse the adjacent pixels of the seed point to be processed; if the difference between the grayscale value of the adjacent pixel and the grayscale value of the seed point to be processed is less than a set difference threshold, add the corresponding adjacent pixel to the background pixel set and add it as a seed point to be processed to the seed point queue; when the seed point queue is empty, the loop ends, and the latest background pixel set is used as the background image corresponding to the original document image.

[0099] Optionally, the masking image determination module is specifically used for: for each unknown pixel in the unknown region image, obtaining an observed color value based on the original document image; sampling a foreground color sample of the unknown pixel in the foreground image; sampling a background color sample of the unknown pixel in the background image; determining the optimal sample pair combination corresponding to the unknown pixel based on an evaluation function constructed based on the foreground color sample and the background color sample; determining the transparency masking value of the unknown pixel based on the optimal sample pair combination and the observed color value; and determining the masking image corresponding to the original document image based on the transparency masking value of each foreground pixel in the foreground image, the transparency masking value of each background pixel in the background image, and the transparency masking value of each unknown pixel in the unknown region image.

[0100] Optionally, the mask image determination module is further configured to: determine the first pixel coordinate distance between each foreground pixel in the foreground image and the unknown pixel; determine the first texture feature similarity between each foreground pixel in the foreground image and the unknown pixel; determine the Pareto front of all foreground pixels based on the first pixel coordinate distance and the first texture feature similarity; and use the Pareto front of all foreground pixels in the foreground image as the foreground color sample of the unknown pixel.

[0101] Optionally, the mask image determination module is further configured to: determine the second pixel coordinate distance between each background pixel in the background image and the unknown pixel; determine the second texture feature similarity between each background pixel in the background image and the unknown pixel; determine the Pareto front of all background pixels based on the second pixel coordinate distance and the second texture feature similarity; and use the Pareto front of all background pixels in the background image as the background color sample of the unknown pixel.

[0102] Optionally, the above apparatus further includes an evaluation function construction module, which is specifically used to: determine the evaluation function based on the spatial proximity term and texture similarity term of the foreground color sample and the background color sample; wherein, the spatial proximity term is obtained based on the spatial coordinates of the unknown pixel, the foreground color sample and the background color sample; and the texture similarity term is obtained based on the local texture feature vectors of the unknown pixel, the foreground color sample and the background color sample.

[0103] The image processing apparatus provided in this application embodiment can execute the image processing method provided in any embodiment of this disclosure, and has the corresponding functional modules and beneficial effects for executing the method.

[0104] Figure 4 A schematic diagram of an electronic device 10, which can be used to implement embodiments of this application, is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (such as helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the application described and / or claimed herein.

[0105] like Figure 4As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12 or a random access memory (RAM) 13, communicatively connected to the at least one processor 11. The memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes based on the computer program stored in the ROM 12 or loaded from storage unit 18 into the RAM 13. The RAM 13 may also store various programs and data required for the operation of the electronic device 10. The processor 11, ROM 12, and RAM 13 are interconnected via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.

[0106] Multiple components in electronic device 10 are connected to input / output (I / O) interface 15, including: input unit 16, such as keyboard, mouse, etc.; output unit 17, such as various types of monitors, speakers, etc.; storage unit 18, such as disk, optical disk, etc.; and communication unit 19, such as network card, modem, wireless transceiver, etc. Communication unit 19 allows electronic device 10 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0107] Processor 11 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Processor 11 performs the various methods and processes described above, such as image processing.

[0108] In some embodiments, the method image processing may be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or mounted on electronic device 10 via read-only memory (ROM) 12 and / or communication unit 19. When the computer program is loaded into random access memory (RAM) 13 and executed by processor 11, one or more steps of the method image processing described above may be performed. Alternatively, in other embodiments, processor 11 may be configured to perform method image processing by any other suitable means (e.g., by means of firmware).

[0109] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0110] Computer programs used to implement the methods of this application may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0111] In the context of this application, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. Alternatively, a computer-readable storage medium can be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0112] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0113] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or computing systems that include middleware components (e.g., application servers), or computing systems that include frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.

[0114] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.

[0115] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the image processing method provided in any embodiment of this application.

[0116] In the implementation of the computer program product, computer program code for performing the operations of this application can be written in one or more programming languages ​​or a combination thereof. Programming languages ​​include object-oriented programming languages ​​such as Java, Smalltalk, and C++, as well as conventional procedural programming languages ​​such as C or similar languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0117] Note that the above description is merely a preferred embodiment and the technical principles employed in this application. Those skilled in the art will understand that this application is not limited to the specific embodiments described herein, and various obvious changes, readjustments, and substitutions can be made without departing from the scope of protection of this application. Therefore, although this application has been described in detail through the above embodiments, this application is not limited to the above embodiments. Many other equivalent embodiments may be included without departing from the concept of this application, and the scope of this application is determined by the scope of the appended claims.

Claims

1. An image processing method, characterized in that, include: Obtain the original document image; Determine the three-part image corresponding to the original document image; wherein the three-part image includes a foreground image, a background image, and an unknown region image; Based on the trisection and the original document image, determine the mask image corresponding to the original document image; The mask image is processed to obtain the target document image.

2. The method according to claim 1, characterized in that, Determining the foreground image corresponding to the original document image includes: The original document image is subjected to image erosion and image dilation operations in sequence to obtain a noisy image; The foreground image corresponding to the original document image is obtained by subtracting the original document image from the noisy image.

3. The method according to claim 1, characterized in that, Determining the background image corresponding to the original document image includes: The pixel with the highest grayscale value in the original document image is identified and added to the background pixel set and added as a seed point to be processed to the seed point queue. Repeat the following steps: Take one of the seed points to be processed from the queue of seed points to be processed; Traverse the neighboring pixels of the seed point to be processed; If the difference between the gray value of the adjacent pixel and the gray value of the seed point to be processed is less than a set difference threshold, then the corresponding adjacent pixel is added to the background pixel set and added to the seed point queue as a seed point to be processed. When the queue of seed points to be processed is empty, the loop ends, and the latest set of background pixels is used as the background image corresponding to the original document image.

4. The method according to claim 1, characterized in that, Determining the mask image corresponding to the original document image based on the triangulation and the original document image includes: For each unknown pixel in the unknown region map, the observed color value is obtained based on the original document image; In the foreground image, foreground color samples of the unknown pixels are sampled; In the background image, sample the background color of the unknown pixel; The evaluation function constructed based on the foreground color sample and the background color sample determines the optimal sample pair combination corresponding to the unknown pixel; The transparency masking value of the unknown pixel is determined based on the optimal sample pair combination and the observed color value. The mask image corresponding to the original document image is determined based on the transparency mask value of each foreground pixel in the foreground image, the transparency mask value of each background pixel in the background image, and the transparency mask value of each unknown pixel in the unknown region image.

5. The method according to claim 4, characterized in that, In the foreground image, sampling the foreground color sample of the unknown pixel includes: Determine the first pixel coordinate distance between each foreground pixel in the foreground image and the unknown pixel; Determine the first texture feature similarity between each foreground pixel in the foreground image and the unknown pixel; The Pareto front of all foreground pixels is determined based on the first pixel coordinate distance and the first texture feature similarity. The Pareto front of all foreground pixels in the foreground image is used as the foreground color sample of the unknown pixel.

6. The method according to claim 4, characterized in that, In the background image, sampling the background color samples of the unknown pixels includes: Determine the second pixel coordinate distance between each background pixel in the background image and the unknown pixel; Determine the similarity of the second texture feature between each background pixel in the background image and the unknown pixel; The Pareto front of all background pixels is determined based on the second pixel coordinate distance and the second texture feature similarity. The Pareto front of all background pixels in the background image is used as the background color sample for the unknown pixel.

7. The method according to claim 4, characterized in that, The evaluation function is constructed as follows: The evaluation function is determined based on the spatial proximity term and texture similarity term of the foreground color sample and the background color sample; wherein the spatial proximity term is obtained based on the spatial coordinates of the unknown pixel, the foreground color sample and the background color sample; and the texture similarity term is obtained based on the local texture feature vectors of the unknown pixel, the foreground color sample and the background color sample.

8. An image processing apparatus, characterized in that, include: The original document image acquisition module is used to acquire original document images; The tripartite determination module is used to determine the tripartite corresponding to the original document image; wherein the tripartite includes a foreground image, a background image, and an unknown region image; The mask image determination module is used to determine the mask image corresponding to the original document image based on the three-part image and the original document image. The processing module is used to process the mask image to obtain the target document image.

9. An electronic device, characterized in that, The electronic device includes: One or more processors; Storage device for storing one or more programs. When the one or more programs are executed by the one or more processors, the one or more processors implement the image processing method as described in any one of claims 1-7.

10. A storage medium comprising computer-executable instructions, which, when executed by a computer processor, are used to perform the image processing method as described in any one of claims 1-7.