Background image generation method, electronic device and computer storage medium
By detecting target content and masking the real images generated based on physical files, generating a fusion mask and erasing the image area, the problem of difficulty in generating realistic scanned image backgrounds of existing scanning software is solved, and the training effect of low-cost and efficient background image generation and machine learning model is improved.
Patent Information
- Application Number
- CN202311355147.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-10-18
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2043-10-18
AI Technical Summary
When generating realistic scan image backgrounds, existing scanning software requires professional rendering software. The threshold is high and the cost is high, and the generated background effect is not realistic enough, resulting in poor performance of machine learning models on real data.
By acquiring the real image generated based on the physical file, performing target content detection and mask processing, generating a fusion mask, and erasing the image area to obtain a background image, reducing the cost and complexity of generating a background image.
It realizes the generation of realistic background images without complex parameter learning and debugging, which reduces costs and is consistent with real samples in data distribution, improving the background effect of generating scanned images and the training effect of machine learning models.
Smart Images

Figure CN117422734B_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the field of computer technology, and in particular to a background image generation method, an electronic device, and a computer storage medium. Background Art
[0002] With the development of computer technology, people often need to electronically scan documents in their daily lives and work, such as electronic books, invoice reimbursement, scanning and printing of work documents, scanning of certificate materials, and other application scenarios.
[0003] Because professional scanners are expensive, scanning software came into being. Through scanning software, users can achieve convenient and low-cost document scanning anytime and anywhere. With the widespread use of scanning software, scanning scenarios are becoming more and more diverse, which requires scanning software to be able to iterate and update continuously to adapt to new scene requirements. However, the iterative update of scanning software requires a large number of training samples in various scanning scenarios as a prerequisite. To this end, in an existing method, professional rendering software (such as unity3D, etc.) is used to build a virtual lighting environment, which can include different light sources, different obstructions, different texture backgrounds, etc., and by changing the position, strength, color temperature, etc. of the light source and then matching different obstructions and paper textures to simulate the scene of the user scanning the file, so that the background of the generated scanned image is more realistic, thereby making the scanned image more realistic. However, in this method, firstly, the threshold for using this type of professional rendering software is high, and it takes a long time to debug and learn to achieve a certain simulation effect, which is time-consuming and results in a high overall implementation cost. Secondly, there is a large deviation between the data distribution of the image background simulated by this type of professional rendering software and the actual sample distribution, which makes the background effect of the generated scanned image poor. As a result, although the machine learning model that implements software scanning works well on simulated training samples, it works poorly on real data. Summary of the invention
[0004] In view of this, an embodiment of the present application provides a background image generation solution to at least partially solve the above-mentioned problem.
[0005] According to a first aspect of an embodiment of the present application, a background image generation method is provided, including: obtaining a real image generated based on a physical file; performing a plurality of different types of target content detection on the real image to obtain a corresponding plurality of target content image areas and a corresponding plurality of target content masks; fusing the plurality of target content masks to obtain a fused mask; erasing the image area corresponding to the fused mask from the real image, and obtaining a background image corresponding to the real image based on the erasing result.
[0006] According to the second aspect of an embodiment of the present application, there is provided an electronic device, comprising: a processor, a memory, a communication interface and a communication bus, wherein the processor, the memory and the communication interface communicate with each other via the communication bus; the memory is used to store at least one executable instruction, and the executable instruction enables the processor to perform an operation corresponding to the method described in the first aspect.
[0007] According to a third aspect of an embodiment of the present application, a computer storage medium is provided, on which a computer program is stored. When the program is executed by a processor, the method described in the first aspect is implemented.
[0008] According to the solution provided by the embodiment of the present application, target content detection and mask processing can be performed based on a real image containing file information, and a fusion mask can be generated based on multiple target content masks, and image regions can be erased based on the fusion mask, thereby obtaining a background image corresponding to the real image. Therefore, on the one hand, there is no need for image processing personnel to perform complex parameter learning and debugging, and the corresponding background image can be generated based on the real image, which reduces the cost of background image generation; on the other hand, the background image is generated based on the real image, so the data distribution is consistent with the real sample distribution, which not only makes the background effect of the generated scanned image better; and later, when the scanned image is synthesized based on the background image as a training sample, the training effect of the machine learning model that implements software scanning can also be better. BRIEF DESCRIPTION OF THE DRAWINGS
[0009] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in the embodiments of the present application. For ordinary technicians in this field, other drawings can also be obtained based on these drawings.
[0010] Figure 1 A schematic diagram of an exemplary system applicable to the embodiment scheme of the present application.
[0011] Figure 2 The present invention is a flowchart of a method for generating a background image according to an embodiment of the present application.
[0012] Figure 3 This is an optional sub-step flowchart of step S104 of the present application.
[0013] Figure 4 This is an optional sub-step flowchart of step S106 of the present application.
[0014] Figure 5 This is an optional sub-step flowchart of step S108 of the present application.
[0015] Figure 6 A schematic diagram of an example scenario in an embodiment of the present application.
[0016] Fig. 7A A schematic diagram of an example scenario of synthesizing a scanned document sample image in an embodiment of the present application.
[0017] Figure 7B A schematic diagram of another scenario example of synthesizing a scanned document sample image in an embodiment of the present application.
[0018] Figure 8 It is a schematic diagram of the structure of an electronic device according to an embodiment of the present application. DETAILED DESCRIPTION
[0019] In order to enable those skilled in the art to better understand the technical solutions in the embodiments of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments in the embodiments of the present application should fall within the scope of protection of the embodiments of the present application.
[0020] The specific implementation of the embodiment of the present application is further explained below in conjunction with the accompanying drawings of the embodiment of the present application.
[0021] Figure 1 An exemplary system applicable to the embodiment of the present application is shown. Figure 1 As shown, the system 100 may include a cloud service end 102, a communication network 104 and / or one or more user devices 106. Figure 1 It should be noted that the solution of the embodiment of the present application can be completed by the cloud service end 102 and the user device 106 in collaboration, or it can be completed by the user device 106 independently.
[0022] The cloud server 102 may be any suitable device for storing information, data, programs and / or any other suitable type of content, including but not limited to distributed storage system devices, server clusters, computing cloud server clusters, etc. In some embodiments, the cloud server 102 may perform any suitable function. For example, when the cloud server 102 and the user device 106 collaborate to complete the solution of the embodiment of the present application, in some embodiments, the cloud server 102 may receive a real image generated based on a physical file sent by the user device 106, and generate a background image corresponding to the real image. As an optional example, in some embodiments, the cloud server 102 may first perform a plurality of different types of target content detection on the real image, obtain a plurality of corresponding target content image areas and a plurality of corresponding target content masks; then, the plurality of target content masks may be fused to obtain a fused mask; then, from the real image, the image area corresponding to the fused mask is erased, and the background image corresponding to the real image is obtained according to the erasing result. As another example, in some optional embodiments, the cloud server 102 may send the generated background image to the user device 106. In some optional embodiments, the cloud server 102 can synthesize a scanned file sample image based on the background image (for example, synthesizing a scanned file sample image based on the background image and a preset file content image; or synthesizing a scanned file sample image based on the background image and a preset file content text), and train a machine learning model for generating scanned files based on the scanned file sample image.
[0023] In some embodiments, the communication network 104 can be any suitable combination of one or more wired and / or wireless networks. For example, the communication network 104 can include any one or more of the following: the Internet, an intranet, a wide area network (WAN), a local area network (LAN), a wireless network, a digital subscriber line (DSL) network, a frame relay network, an asynchronous transfer mode (ATM) network, a virtual private network (VPN) and / or any other suitable communication network. The user device 106 can be connected to the communication network 104 via one or more communication links (e.g., communication link 112), and the communication network 104 can be linked to the cloud service end 102 via one or more communication links (e.g., communication link 114). The communication link can be any communication link suitable for transmitting data between the user device 106 and the cloud service end 102, such as a network link, a dial-up link, a wireless link, a hard-wired link, any other suitable communication link, or any suitable combination of such links.
[0024] The user device 106 may include any one or more user devices suitable for presenting images, interacting with users, etc. When the solution of the embodiment of the present application is independently completed by the user device 106, the user device 106 may be used to generate a background image corresponding to the real image based on the real image generated based on the physical file. As an optional example, in some embodiments, the user device 106 may first perform a plurality of different types of target content detection on the real image to obtain a plurality of corresponding target content image areas and a plurality of corresponding target content masks; then, the plurality of target content masks may be fused to obtain a fused mask; then, from the real image, the image area corresponding to the fused mask is erased, and the background image corresponding to the real image is obtained according to the erasure result. In some optional embodiments, the user device 106 may synthesize a scanned file sample image based on the background image (for example, synthesize a scanned file sample image based on the background image and a preset file content image; or synthesize a scanned file sample image based on the background image and a preset file content text). Then, the user device 106 may send the scanned file sample image to the cloud service end 102 to train the machine learning model for generating scanned files in the cloud service end 102. In some embodiments, the user device 106 may include any suitable type of device. For example, in some embodiments, user device 106 may include a mobile device, a tablet computer, a laptop computer, a desktop computer, and / or any other suitable type of user device.
[0025] Based on the above system, an embodiment of the present application provides a background image generation solution, which is described below through multiple embodiments.
[0026] Figure 2 A flowchart of a background image generation method according to an embodiment of the present application is provided. According to a first aspect of the present application, a background image generation method is provided, referring to Figure 2 As shown, the method includes steps S102, S104, S106 and S108, specifically:
[0027] S102: Acquire a real image generated based on the physical file.
[0028] In this application, a physical file may refer to a physical file for scanning. A physical file may be of any type, for example, a physical file may include but is not limited to books, invoices, contracts, business licenses, work documents, certificate materials, identity cards, driver's licenses, etc.
[0029] The real image in the present application has the characteristics of a real sample, and its data distribution is consistent with that of the real sample, which makes it possible to generate a background image based on the real image, and the data distribution of the obtained background image is consistent with that of the real sample.
[0030] Optionally, the real image generated based on the physical document in step S102 of the present application may be an image obtained by pre-shooting the physical document with a camera. For example, taking the physical document as an invoice as an example, the real image generated based on the invoice may be obtained in the following manner: the user places the invoice to be scanned on a support (such as a desktop, etc.) or holds it in the hand, and then uses the camera on the mobile terminal to shoot the invoice from the direction facing the invoice, thereby obtaining the corresponding real image. It should be understood that this is only an example and does not constitute any limitation in the present application.
[0031] It should be understood that the real image generated based on the physical file in step S102 can be obtained by photographing the real image when needed and then directly obtaining it, or it can be obtained by photographing the real image in advance and then storing the real image in a predetermined storage space (including but not limited to storage media such as disks, hard disks, memories, or databases, etc.), and directly obtaining it from the storage space when needed. There is no limitation on this in this application.
[0032] Optionally, the real image may also include images with defects after being processed by the model (for example, these images may be from the historical process of users (including online users and / or offline users) using the scanning software, and the images with defects after being processed by the machine learning model used to implement software scanning are continuously accumulated). The required real images can be manually selected from the images with defects after being processed by the model and stored in the storage space, or the required real images can be picked out from the images with defects after being processed by the model according to the rules by the algorithm and stored in the storage space. For example, defects may be objects in the image, including but not limited to shadows, folds, transparent words, fingers, clips, etc. These objects can be used as part of the background of the real image, belong to the characteristics of the real sample, and their data distribution is consistent with the real sample. By generating a background image through such a real image, and subsequently using the background image to synthesize a sample image of a scanned file to train the machine learning model used to generate the scanned file, the capability boundary of the machine learning model can be expanded.
[0033] S104: Performing a plurality of different types of target content detection on the real image to obtain a plurality of corresponding target content image regions and a plurality of corresponding target content masks.
[0034] The real image generated based on the physical file obtained in step S102 includes multiple different types of target content, which are usually information content with practical significance. Therefore, in step S104 of the present application, multiple different types of target content detection can be performed on the real image to obtain corresponding multiple target content image areas and corresponding multiple target content masks, so that the subsequent steps can use multiple target content masks to perform data processing to obtain a background image that better meets the needs.
[0035] In the present application, the target content can be determined based on the corresponding real image. In some optional embodiments, the target content in the present application includes at least one of the following types: text type content, chart type content, color block pattern type content, and seal type content. Figure 3 In the flowchart shown, step S104 includes sub-steps S1041, S1042 and S1043, specifically:
[0036] S1041: Perform at least two of the following target content detections on the real image: detection on text, detection on charts, detection on color block patterns, and detection on seals.
[0037] In the present application, any feasible target detection algorithm can be used to achieve the detection of text, graphics, color block patterns and seals.
[0038] In the present application, by detecting at least two different types of target content in real images, such as text, graphics, color block patterns, and seals, the background image generation method in the present application can better adapt to the needs of document scanning in real life.
[0039] In addition, the present application does not specifically limit the text, charts, color block patterns, and seals. For example, the following description can be referred to for schematic understanding: the text may include but is not limited to Chinese characters, Arabic numerals, English text, etc.; the chart may include but is not limited to bar charts, pie charts, line charts, tables, etc.; the color block pattern may include any regular color block pattern or irregular color block pattern; the seal may include but is not limited to regular shape seals such as circular seals and rectangular seals, or may be a seal of other shapes.
[0040] It should also be noted that, although the present application performs at least two different types of target content detection on real images, including text, charts, color block patterns, and seals, this does not mean that multiple different types of target content can be detected from the real image. This depends on the actual situation of the target image, and the present application does not impose any restrictions on this. For example, if there is only text-type content in the real image, but no charts, color block patterns, seals, etc., although multiple different types of target detection will be performed, the detection results obtained may only have the results corresponding to the text-type content. Of course, this is just an example and does not serve as any limitation to the present application.
[0041] In addition, it should be noted that, in the present application, “multiple”, “plurality” and other numbers related to “plurality” mean two or more.
[0042] S1042: According to the detection result, a plurality of image regions corresponding to the detected plurality of types of target contents are obtained.
[0043] Specifically, according to the detection result of the target content detection obtained in sub-step S1041, multiple image areas corresponding to the detected multiple types of target content can be determined from the real image, and these image areas include at least part of the content in the real image except the background.
[0044] Since sub-step S1041 detects at least two different types of target contents of text, chart, color block pattern, and seal on the real image, after the corresponding content is detected, it can be appropriately processed to obtain the corresponding image area. For example, in some optional embodiments, sub-step S1042 includes at least one of the following sub-steps S1042A, S1042B, S1042C, and S1042D, specifically:
[0045] S1042A: If it is determined based on the detection result that text type content is detected, text recognition is performed on the text type content to determine the image areas corresponding to the respective characters based on the text recognition result.
[0046] Optionally, any suitable text recognition algorithm may be used to recognize text type content, so as to determine the image regions corresponding to each text according to the text recognition result. For example, the text recognition algorithm includes but is not limited to OCR (Optical Character Recognition) algorithm.
[0047] S1042B: If it is determined according to the detection result that chart type content is detected, image segmentation is performed on the chart type content to obtain an image region corresponding to the chart type content.
[0048] Optionally, any suitable image segmentation algorithm may be used to segment the chart type content to obtain the image region corresponding to the chart type content. For example, the image segmentation algorithm includes but is not limited to a semantic segmentation algorithm, an instance segmentation algorithm, and the like.
[0049] S1042C: If it is determined based on the detection result that a color block pattern type content is detected, image segmentation is performed on the color block pattern type content to obtain an image area corresponding to the color block pattern content.
[0050] Optionally, any suitable image segmentation algorithm may be used to segment the color block pattern type content to obtain an image region corresponding to the color block pattern type content. For example, the image segmentation algorithm includes but is not limited to a semantic segmentation algorithm, an instance segmentation algorithm, and the like.
[0051] S1042D: If it is determined based on the detection result that seal type content is detected, text recognition and frame recognition are performed on the seal type content to determine the image area corresponding to the seal based on the recognition result.
[0052] Optionally, any suitable text recognition algorithm and any suitable frame recognition algorithm may be used to recognize the seal type content to obtain the image area corresponding to the seal type content. For example, the text recognition algorithm includes but is not limited to an OCR (Optical Character Recognition) algorithm.
[0053] Based on this, in the present application, through the optional implementation of sub-steps S1042A to S1042D, for different target contents determined according to the detection results, different methods are used to adaptively determine the image area corresponding to the target content, so that the determination result of the image area can be accurately and effectively obtained.
[0054] S1043: Masking the multiple image regions to obtain corresponding multiple target content masks.
[0055] After obtaining multiple target regions in sub-step S1042, mask processing can be performed on multiple image regions to obtain corresponding multiple target content masks. The specific method of obtaining the mask can be implemented by a person skilled in the art in any appropriate manner, including but not limited to obtaining the mask by a machine learning model for masking the target object in the image. In addition, optionally, when the machine learning model outputs the mask corresponding to the image region, it can also output the confidence corresponding to each mask at the same time.
[0056] Based on this, in the present application, through the optional implementation of the above-mentioned sub-steps S1041 to S1043, it is possible to accurately and effectively implement various types of target content detection on real images to obtain corresponding multiple target content image areas and corresponding multiple target content masks, thereby facilitating subsequent use of them for data processing to obtain a background image that better meets the needs.
[0057] S106: Fusing multiple target content masks to obtain a fused mask.
[0058] After obtaining multiple target content masks, the multiple target content masks are fused (such as fused by merging, etc.) to obtain a fused mask that can better meet the data processing requirements. By fusing the masks, on the one hand, the effect of generating background images in subsequent steps can be better improved, and on the other hand, it is also convenient to reduce the amount of calculation when generating background images subsequently.
[0059] The specific implementation of step S106 is not limited in this application, and it only needs to meet the needs. Figure 4 In the flowchart shown, step S106 includes sub-steps S1061 and S1062, specifically:
[0060] S1061: Filter the multiple target content masks according to the confidence of each target content mask in the multiple target content masks.
[0061] For example, as mentioned above, when the machine learning model outputs the mask corresponding to the image area, it may also output the confidence corresponding to the mask. In this step, according to the confidence of each target content mask among multiple target content masks, the target content masks with lower confidence can be filtered out to ensure the accuracy of the target content mask.
[0062] The specific implementation of sub-step S1061 is not limited in this application. In some optional embodiments, filtering can be performed by setting a confidence threshold. For example, step S1061 can be specifically performed by determining whether the confidence of each target content mask in the multiple target content masks is greater than the confidence threshold. If so, it is retained, and if not, it is not retained, thereby achieving the purpose of filtering the multiple target content masks. The confidence threshold can be set as needed, for example, it can be set to 70%, 80%, etc.
[0063] Alternatively, sub-step S1061 may also set other conditions for filtering as needed, and this application does not make any specific restrictions here.
[0064] S1062: Fusing the filtered target content masks to obtain a fused mask.
[0065] Exemplarily, the target content masks may be fused in a manner such as mask merging.
[0066] Based on this, in the present application, through the optional implementation of the above-mentioned sub-steps S1061 to S1062, multiple target content masks can be effectively fused to obtain a fused mask. On the one hand, fusing multiple target content masks into a whole can be more convenient for subsequent processing, such as erasing processing, to improve processing speed; on the other hand, it can also better avoid erasing excessive background areas in the real image when the image area corresponding to the fused mask is subsequently erased, thereby better improving the effect of generating the background image in the subsequent steps and reducing the amount of calculation when generating the background image subsequently.
[0067] S108: Erasing the image area corresponding to the fusion mask from the real image, and obtaining a background image corresponding to the real image according to the erasing result.
[0068] After the fusion mask is obtained in step S106, the image area corresponding to the fusion mask can be erased from the real image, so that only the background remains in the real image after erasure, so that the background image corresponding to the real image can be obtained according to the erasing result.
[0069] Based on this, the background image generation method of the above steps S102 to S108 in this application can detect and mask the target content based on the real image containing the file information, generate a fusion mask based on multiple target content masks, and erase the image area based on the fusion mask, so as to obtain the background image corresponding to the real image. Therefore, on the one hand, there is no need for image processing personnel to perform complex parameter learning and debugging, and the corresponding background image can be generated based on the real image, which reduces the cost of background image generation; on the other hand, the background image is generated based on the real image, so the data distribution is consistent with the real sample distribution, which can not only make the background effect of the generated scanned image better; and, later, when the scanned image is synthesized based on the background image as a training sample, it can also make the training effect of the machine learning model for software scanning better.
[0070] In the present application, the specific implementation of step S108 is not limited, as long as the corresponding requirements can be met. For example, when implementing step S108 and its optional implementations, the content of the image area corresponding to the fusion mask can be erased from the real image to obtain an erasure result, and the erased area can be filled with the content of the adjacent area to obtain a background image corresponding to the real image according to the erasure result. For example, erasure can be implemented by including but not limited to traditional image filling technology, image_inpainting technology based on artificial intelligence (AI), or by comparison, and the present application does not make any limitation on this.
[0071] In some optional embodiments, after obtaining the fusion mask in step S106, the background image generation method further includes: performing morphological processing on the fusion mask according to morphological processing parameters; "erasing the image area corresponding to the fusion mask from the real image" in step S108 includes: erasing the image area corresponding to the fusion mask that has been morphologically processed from the real image.
[0072] Optionally, by performing morphological processing on the fused mask and erasing the morphologically processed fused mask from the real image, the effect of obtaining a background image corresponding to the real image according to such an erasing result is better.
[0073] Optionally, the fusion mask is subjected to morphological processing, which may include at least one of corrosion processing and dilation processing. Different morphological parameters may be used to perform different types of morphological processing according to the needs of different real images, which is not particularly limited in the present application.
[0074] In some optional embodiments, referring to Figure 5 In the flowchart shown, step S108 includes sub-steps S1081 and S1082, specifically:
[0075] S1081: Erasing the image area corresponding to the fusion mask from the real image to obtain an erased image;
[0076] S1082: Calculate the information entropy of the erased image. If the information entropy meets a preset information entropy threshold, use the erased image as the background image corresponding to the real image.
[0077] Specifically, in sub-step S1082, information entropy can be used to indicate the richness of the content in the image. A higher value of information entropy indicates a higher amount of information and richer content, and vice versa. Thus, the information entropy of the erased image can be calculated to determine whether the image area corresponding to the fusion mask in the real image is erased cleanly.
[0078] In the present application, a preset information entropy threshold can be used to determine whether the information entropy of the erased image meets the requirements. If the information entropy of the erased image meets the preset information entropy threshold, it can be considered that the image area corresponding to the fusion mask in the real image has been erased cleanly and can be used as the background image corresponding to the real image.
[0079] Optionally, if the information entropy of the erased image is lower than a preset information entropy threshold, the information entropy of the erased image satisfies the preset information entropy threshold, whereas if the information entropy of the erased image is higher than the preset information entropy threshold, the information entropy of the erased image does not satisfy the preset information entropy threshold. When the information entropy of the erased image satisfies the preset information entropy threshold, the erased image can be used as the background image corresponding to the real image.
[0080] Based on this, in the present application, through the optional implementation of sub-steps S1081 to S1082, it can be ensured that the image area corresponding to the fusion mask in the real image is erased cleanly, and the erased image obtained after erasing is used as the background image of the real image, which can make the background effect of the generated scanned image better, and subsequently when the scanned image is synthesized based on the background image as a training sample, it can also make the training effect of the machine learning model that realizes software scanning better.
[0081] In the present application, the preset information entropy threshold can be set according to actual needs. In some optional embodiments, the preset information entropy threshold is determined by: determining a preset number of historical images similar to the real image currently being processed, and determining the preset information entropy threshold according to the minimum value of the information entropy of the background image corresponding to the historical image. In this way, the preset information entropy threshold can be determined more reasonably, so that the result of determining whether the information entropy of the erased image meets the preset information entropy threshold is more reasonable.
[0082] Specifically, the historical image is the real image in history, and the real image currently being processed is only different in time. Optionally, the historical image mentioned in this application is similar to the real image currently being processed, which may refer to at least one of the type and content layout of the historical image being similar. For example, the type of the historical image is a real image generated based on an invoice, and the real image currently being processed is also a real image generated based on an invoice, then the two are considered similar; for another example, the type of the historical image is a real image generated based on a business license, and the real image currently being processed is also a real image generated based on a business license, then the two are considered similar; for another example, the historical image is a real image generated based on a page of a book with text type content and chart type content, and the real image currently being processed is a real image generated based on a page of a contract with text type content and chart type content, then the two are also considered similar. It should be understood that these examples are only examples for ease of understanding and are not intended to limit the present application in any way. In addition, the preset number can be selected as needed to meet the needs.
[0083] An example is given for ease of understanding: assuming that the preset number is 100, 100 historical images similar to the real image currently being processed can be determined, and the contents of the 100 historical images except the background are respectively erased, and the corresponding information entropy values are calculated to obtain 100 information entropy values, and the minimum information entropy value is determined from the 100 information entropy values, and is used as the preset information entropy threshold. It should be understood that this example is only an example for ease of understanding and does not serve as any limitation to the present application.
[0084] In some optional embodiments, if the information entropy of the erased image does not meet a preset information entropy threshold, the morphological processing parameters are adjusted to perform morphological processing on the fusion mask, and the fusion mask after the morphological processing is completed is used to return to step S108 and re-execute the step of "erasing the image area corresponding to the fusion mask from the real image".
[0085] Specifically, when the information entropy of the erased image does not meet the preset information entropy threshold, the morphological processing parameters are adjusted to perform morphological processing on the fusion mask, and the fusion mask after the morphological processing is completed is used to return to the real image, erase the image area corresponding to the fusion mask and re-execute, so as to facilitate the adjustment of the information entropy of the erased image that does not meet the preset information entropy threshold to meet the preset information entropy threshold, so as to ensure that the image area corresponding to the fusion mask in the real image is erased cleanly, and the erased image obtained by erasing is used as the background image of the real image, which can make the background effect of the generated scanned image better, and when the scanned image is subsequently synthesized based on the background image as a training sample, it can also make the training effect of the machine learning model that realizes software scanning better.
[0086] Figure 6 This is a schematic diagram of an example of a scenario in an embodiment of the present application. Figure 6 The background image generation method of the present application is generally understood as shown in FIG. Figure 6 A real image generated based on a physical file is shown in Figure 6 The real image includes text type content, icon type content, color block pattern type content and seal type content. The background image generated by the background image generation method in this application is output (such as Figure 6 As shown in the right interface of the figure, the above-mentioned text type content, icon type content, color block pattern type content and seal type content are erased. It should be understood that Figure 6 The examples in the description are not intended to limit the present application in any way.
[0087] In some optional embodiments, the background image generation method further includes: synthesizing a scanned document sample image based on the background image and a preset document content image. In this way, the present application can effectively synthesize a scanned document sample image for easy use.
[0088] Specifically, the preset file content image can be an existing image that can represent the file content. The preset file content image can also include an image representing the file content segmented from an image (the image mentioned here can include but is not limited to a real image). For example, the file content can include at least one of text content, chart content, color block pattern content, seal content, etc. The preset file content image can be stored in a suitable storage space after being segmented, so that it can be obtained when needed. Optionally, the preset file content image can be a solid color background or a transparent background, which can be more convenient to synthesize with the background image generated in step S108 to obtain a scanned file sample image. In an optional manner, the above-mentioned solid color background can be a white background, which is closer to the actual scanned file background and is also easier to synthesize.
[0089] Reference Fig. 7A The example scenario shown in the figure is understood, which shows an exemplary process of synthesizing a scanned document sample image based on a background image and a preset document content image. It should be understood that Fig. 7A The examples shown in the accompanying drawings are not intended to limit the embodiments of the present application.
[0090] Optionally, a scanned document sample image may be synthesized by overlaying or pasting a preset document content image onto a suitable area of the background image.
[0091] Alternatively, in some other optional embodiments, the background image generation method further comprises: synthesizing a scanned file sample image based on the background image and a preset file content text. In this way, the present application can also effectively synthesize a scanned file sample image for easy use.
[0092] Specifically, the preset file content text can be text content generated as needed. It should be noted that the difference between the file content text and the above-mentioned file content image is that the file content image is an existing image, while the file content text is generated as needed. For example, if the user wants the middle area of the scanned file sample image to include the word "invoice", the user can input the text content of the word "invoice" through the software, and adjust the appropriate text format (such as font, font size, etc.), and adjust such "invoice" text content to the middle area of the background image manually or automatically, thereby realizing the synthesis of a scanned file sample image that meets the needs based on the background image and the preset file content text; it should be understood that this example does not serve as any limitation to the present application.
[0093] Reference Figure 7B The example scenario shown in the figure is understood, which shows an exemplary process of synthesizing a scanned document sample image based on a background image and a preset document content text. It should be understood that Figure 7B The examples shown in the accompanying drawings are not intended to limit the embodiments of the present application.
[0094] In some optional embodiments, the background image generation method further includes: training a machine learning model for generating scanned documents based on sample images of the scanned documents.
[0095] The scanned document sample images obtained by the technical solution of the present application are used as training samples. The background images of the training samples are generated based on real images, and thus the data distribution is consistent with the real sample distribution. Therefore, the background effect of the training samples is better, and the training effect of the machine learning model used to generate scanned documents can be better. In addition, since the background images and scanned document sample images (i.e., training samples) can be efficiently generated through the technical solution in the present application, it is also convenient to help the machine learning model to iterate quickly and in a targeted manner.
[0096] In summary, the solution provided in the embodiments of the present application can detect and mask target content based on a real image containing file information, generate a fusion mask based on multiple target content masks, and erase the image area based on the fusion mask, so as to obtain a background image corresponding to the real image. Therefore, on the one hand, there is no need for image processing personnel to perform complex parameter learning and debugging, and the corresponding background image can be generated based on the real image, which reduces the cost of background image generation; on the other hand, the background image is generated based on the real image, so the data distribution is consistent with the real sample distribution, which can not only make the background effect of the generated scanned image better; and, later, when the scanned image is synthesized based on the background image as a training sample, it can also make the training effect of the machine learning model that realizes software scanning better.
[0097] It can be understood that the above description of the background image generation method is only an exemplary description of the present application and does not constitute any limitation to the present application.
[0098] According to a second aspect of the embodiment of the present application, an electronic device is provided, comprising: a processor, a memory, a communication interface and a communication bus, wherein the processor, the memory and the communication interface communicate with each other via the communication bus; the memory is used to store at least one executable instruction, wherein the executable instruction enables the processor to perform an operation corresponding to the method described in the first aspect. Figure 8 , shows a schematic diagram of the structure of an electronic device according to an embodiment of the present application. The specific embodiment of the present application does not limit the specific implementation of the electronic device.
[0099] like Figure 8As shown, the electronic device 800 may include: a processor (processor) 802 , a communication interface (Communications Interface) 804 , a memory (memory) 806 , and a communication bus 808 .
[0100] in:
[0101] The processor 802 , the communication interface 804 , and the memory 806 communicate with each other via a communication bus 808 .
[0102] The communication interface 804 is used to communicate with other electronic devices or servers.
[0103] The processor 802 is used to execute the program 810, and specifically can execute the relevant steps in the above-mentioned background image generation method embodiment.
[0104] Specifically, the program 810 may include program codes, which include computer operation instructions.
[0105] The processor 802 may be a CPU, a GPU (Graphic Processing Unit), an ASIC (Application Specific Integrated Circuit), or one or more integrated circuits configured to implement the embodiments of the present application. The one or more processors included in the smart device may be processors of the same type, such as one or more CPUs; or may be processors of different types, such as one or more CPUs and one or more ASICs.
[0106] The memory 806 is used to store the program 810. The memory 806 may include a high-speed RAM memory, and may also include a non-volatile memory (non-volatile memory), such as at least one disk memory.
[0107] The program 810 may include multiple computer instructions. Specifically, the program 810 may enable the processor 802 to execute operations corresponding to the background image generation method described in any of the aforementioned method embodiments through the multiple computer instructions.
[0108] The specific implementation of each step in program 810 can refer to the corresponding description of the corresponding steps and units in the above method embodiment, and has corresponding beneficial effects, which will not be repeated here. Those skilled in the art can clearly understand that for the convenience and simplicity of description, the specific working process of the above-described devices and modules can refer to the corresponding process description in the above method embodiment, which will not be repeated here.
[0109] According to the third aspect of the embodiments of the present application, the embodiments of the present application further provide a computer storage medium on which a computer program is stored, and when the program is executed by a processor, the method described in any of the foregoing multiple method embodiments is implemented. The computer storage medium includes, but is not limited to: a compact disc read-only memory (CD-ROM), a random access memory (RAM), a floppy disk, a hard disk, or a magneto-optical disk, etc.
[0110] An embodiment of the present application also provides a computer program product, including computer instructions, which instruct a computing device to execute operations corresponding to any background image generation method in the above-mentioned multiple method embodiments.
[0111] The electronic device 800 / computer storage medium / computer program product embodiment in the embodiment of the present application has been described in detail in the aforementioned background image generation method embodiment, so its related content and beneficial effects can be understood by referring to the above-mentioned method embodiment and will not be repeated here.
[0112] In addition, it should be noted that the user-related information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to sample data used to train the model, data used for analysis, stored data, displayed data, etc.) involved in the embodiments of the present application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data need to comply with the relevant laws, regulations and standards of the relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse.
[0113] It should be pointed out that, according to the needs of implementation, the various components / steps described in the embodiments of the present application can be split into more components / steps, or two or more components / steps or partial operations of components / steps can be combined into new components / steps to achieve the purpose of the embodiments of the present application.
[0114] The above-mentioned method according to the embodiment of the present application can be implemented in hardware, firmware, or implemented as software or computer code that can be stored in a recording medium (such as a CD-ROM, RAM, floppy disk, hard disk or magneto-optical disk), or implemented as a computer code originally stored in a remote recording medium or a non-temporary machine-readable medium downloaded through a network and stored in a local recording medium, so that the method described herein can be stored in such software processing on a recording medium using a general-purpose computer, a dedicated processor or programmable or dedicated hardware (such as an application-specific integrated circuit (ASIC) or a field programmable gate array (FPGA)). It can be understood that a computer, a processor, a microprocessor controller or programmable hardware includes a storage component (e.g., a random access memory (RAM), a read-only memory (ROM), a flash memory, etc.) that can store or receive software or computer code, and when the software or computer code is accessed and executed by a computer, a processor or hardware, the method described herein is implemented. In addition, when a general-purpose computer accesses the code for implementing the method shown here, the execution of the code converts the general-purpose computer into a dedicated computer for executing the method shown here.
[0115] Those of ordinary skill in the art will appreciate that the units and method steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for specific applications, but such implementation should not be considered to be beyond the scope of the embodiments of the present application.
[0116] The above implementation methods are only used to illustrate the embodiments of the present application, and are not limitations on the embodiments of the present application. Ordinary technicians in the relevant technical field can make various changes and modifications without departing from the spirit and scope of the embodiments of the present application. Therefore, all equivalent technical solutions also belong to the scope of the embodiments of the present application. The scope of patent protection of the embodiments of the present application should be limited by the claims.
Claims
1. A background image generation method, include: Get real images generated based on physical files; Performing a plurality of different types of target content detection on the real image to obtain a plurality of corresponding target content image regions and a plurality of corresponding target content masks; Fusing the multiple target content masks to obtain a fused mask; Erasing the image area corresponding to the fusion mask from the real image, and obtaining a background image corresponding to the real image according to the erasing result; Wherein, erasing the image area corresponding to the fusion mask from the real image, and obtaining the background image corresponding to the real image according to the erasing result, includes: Erasing the image area corresponding to the fusion mask from the real image to obtain an erased image; Calculate the information entropy of the erased image. If the information entropy meets a preset information entropy threshold, use the erased image as the background image corresponding to the real image. The preset information entropy threshold is determined by determining a preset number of historical images similar to the real image, and determining the minimum value among the information entropies of the background images corresponding to the historical images.
2. The method according to claim 1, in, The step of fusing the multiple target content masks to obtain a fused mask includes: filtering the multiple target content masks according to the confidence of each target content mask in the multiple target content masks; The filtered target content masks are fused to obtain a fused mask.
3. The method according to claim 2, in, After obtaining the fusion mask, the method further comprises: performing morphological processing on the fusion mask according to morphological processing parameters; Erasing the image region corresponding to the fusion mask from the real image includes: erasing the image region corresponding to the fusion mask that has been subjected to morphological processing from the real image.
4. The method according to any one of claims 1 to 3, in, The method further comprises: If the information entropy does not meet the preset information entropy threshold, the morphological processing parameters are adjusted to perform morphological processing on the fusion mask, and the fusion mask after the morphological processing is completed is used to return to the step of erasing the image area corresponding to the fusion mask from the real image and re-execute it.
5. The method according to any one of claims 1 to 3, in, The preset information entropy threshold is determined in the following way: A preset number of historical images similar to the real image currently being processed are determined, and the preset information entropy threshold is determined according to a minimum value among the information entropies of the background images corresponding to the historical images.
6. The method according to claim 1, in, The target content includes at least one of the following types: text type content, chart type content, color block pattern type content, and seal type content; The performing a plurality of different types of target content detection on the real image to obtain a plurality of corresponding target content image regions and a plurality of corresponding target content masks includes: Performing at least two of the following target content detections on the real image: detection on text, detection on charts, detection on color block patterns, and detection on seals; According to the detection results, a plurality of image regions corresponding to the detected plurality of types of target contents are obtained; The multiple image regions are masked to obtain corresponding multiple target content masks.
7. The method according to claim 6, in, The obtaining, according to the detection result, a plurality of image regions corresponding to the detected plurality of types of target contents comprises at least one of the following: If it is determined according to the detection result that text type content is detected, then text recognition is performed on the text type content to determine the image areas corresponding to the respective characters according to the text recognition result; If it is determined according to the detection result that a chart type content is detected, performing image segmentation on the chart type content to obtain an image area corresponding to the chart type content; If it is determined according to the detection result that a color block pattern type content is detected, image segmentation is performed on the color block pattern type content to obtain an image area corresponding to the color block pattern content; If it is determined according to the detection result that seal type content is detected, text recognition and frame recognition are performed on the seal type content to determine the image area corresponding to the seal according to the recognition result.
8. The method according to any one of claims 1 to 3, in, The method further comprises: synthesizing a scanned document sample image based on the background image and a preset document content image; or, A scanned document sample image is synthesized based on the background image and the preset document content text.
9. The method according to claim 8, in, The method further comprises: Based on the scanned document sample images, a machine learning model for generating scanned documents is trained.
10. An electronic device, include: A processor, a memory, a communication interface and a communication bus, wherein the processor, the memory and the communication interface communicate with each other via the communication bus; The memory is used to store at least one executable instruction, and the executable instruction enables the processor to perform an operation corresponding to the method according to any one of claims 1 to 9.
11. A computer storage medium having a computer program stored thereon, wherein when the program is executed by a processor, the method according to any one of claims 1 to 9 is implemented.
Citation Information
Patent Citations
Comprehensive evaluation method and evaluation system for scanned image quality
CN105261013A
Video processing method, system, device and medium
CN113362365A
Image generation method and device, equipment and storage medium
CN115131464A