Training sample generation method and electronic device
By using pixel change trend chart to light and shadow processing on the scanned document image, a simulated image with shadow and/or reflective areas is generated, which solves the problems of low efficiency and poor flexibility in the production of training samples in the prior art, and achieves efficient and flexible training sample generation, which improves the stability and effect of model training.
Patent Information
- Application Number
- PCT/CN2024/130579
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-18
- Filing Date
- 2024-11-07
- Publication Date
- 2025-06-26
AI Technical Summary
The prior art requires complex parameter adjustment and manual operation when generating training samples for machine learning model training, which are inefficient and poorly flexible, and are prone to introduce random color changes, which makes it difficult for model training to converge.
By obtaining a pixel change trend chart, the pixel change between the light and shadow region and other regions is determined, based on this, the original scanned document image is subjected to light and shadow processing, and a simulated scanned document image with shadow and/or reflective regions is generated, and the training sample is then generated.
It reduces the burden of manual operation, improves the efficiency of light and shadow processing, provides higher flexibility and diversity, avoids model training problems caused by color changes, and the generated training samples are of good quality.
Smart Images

Figure CN2024130579_26062025_PF_FP_ABST
Abstract
Description
Training sample generation method and electronic device
[0001] This application claims priority to the Chinese patent application filed with the China Patent Office on December 18, 2023, with application number 202311750611.3 and application name “Training Sample Generation Method and Electronic Device,” all contents of which are incorporated by reference into this application. Technical Field
[0002] The embodiments of the present application relate to the field of computer technology, and in particular to a training sample generation method and electronic device. Background Art
[0003] With the development of computer technology, people often need to electronically scan documents in their daily lives and work, such as electronic books, invoice reimbursement, work document scanning and printing, document material scanning, and other application scenarios.
[0004] Because professional scanners are expensive, scanning software has emerged. Scanning software allows users to conveniently and cost-effectively scan documents anytime, anywhere. With the widespread use of scanning software, scanning scenarios are becoming increasingly diverse. For example, in some scenarios, the effects of a filter enhancing the contrast of scanned images are unnecessary; instead, shadows and reflections can be removed from the captured scanned image. To achieve this, training samples tailored to these needs are needed to iteratively update the machine learning models that implement software scanning. One existing approach involves first collecting shadow-free, reflection-free images with a document background color. Then, specialized rendering software (such as Unity3D) is used to simulate shadows and reflections for these images. However, this approach is inefficient due to the complex parameter adjustments involved. Furthermore, the lighting effects achieved through parameter adjustments are limited to the current image, resulting in limited flexibility. Furthermore, the lighting effects simulated by the rendering software parameters can alter the document background color to a certain extent, introducing random color variations in the resulting training sample pairs, making subsequent model training difficult to converge.
[0005] Summary of the Invention
[0006] In view of this, an embodiment of the present application provides a training sample generation solution to at least partially solve the above-mentioned problem.
[0007] According to a first aspect of an embodiment of the present application, a training sample generation method is provided, comprising: obtaining a pixel change trend graph indicating pixel changes between a light and shadow area in an image and other areas other than the light and shadow area, wherein the light and shadow area includes a shadow area and / or a reflective area in the image; performing light and shadow processing on an original scanned document image based on the pixel change trend graph to obtain a corresponding simulated scanned document image having a shadow area and / or a reflective area; and generating a training sample pair based on the original scanned document image and the simulated scanned document image.
[0008] According to the second aspect of an embodiment of the present application, an electronic device is provided, comprising: a processor, a memory, a communication interface and a communication bus, wherein the processor, the memory and the communication interface communicate with each other through the communication bus; the memory is used to store at least one executable instruction, and the executable instruction enables the processor to perform an operation corresponding to the method described in the first aspect.
[0009] According to a third aspect of the embodiments of the present application, a computer storage medium is provided, on which a computer program is stored. When the program is executed by a processor, the method described in the first aspect is implemented.
[0010] According to a fourth aspect of an embodiment of the present application, a computer program product is provided, comprising computer instructions, which instruct a computing device to perform operations corresponding to the method described in the first aspect.
[0011] According to the solution provided by the embodiment of the present application, the pixel change trend graph can be used to determine the pixel change between the shadow area and the non-shadow area (i.e., other areas) under certain lighting conditions, and / or the pixel change between the reflective area and the non-reflective area (i.e., other areas). Based on this, a light and shadow processing can be performed on a certain original scanned document image. The original scanned document image can be an image without shadows and reflections. In this case, the original scanned document image is light and shadow processed using the pixel change trend graph to add shadows and / or reflections to some areas of the image, thereby generating a simulated scanned document image with shadow areas and / or reflective areas. Furthermore, by combining the original scanned document image and the simulated scanned document image, a training sample pair can be generated for training a machine learning model that implements software scanning. Therefore, on the one hand, there is no need for manual long-term learning and complex parameter adjustment of rendering software, which not only reduces the burden of manual operation, but also improves the efficiency of light and shadow processing of images; on the other hand, the pixel change trend graph describes the difference and change trend between shadows and / or reflections and other parts, which can be applied to original scanned document images in various situations and has high flexibility. If the pixel change trend graph includes multiple, different light and shadow processing can be performed on the original scanned document image according to needs to obtain more simulated scanned document images and expand the number of training samples; on the other hand, the pixel change trend graph will not change the background color of the original scanned document image, and will not bring about pixel deviation and random color changes. Therefore, the generated training sample pair is of better quality, which can avoid the problem of difficult convergence of model training in traditional methods. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments recorded in the embodiments of the present application. For ordinary technicians in this field, other drawings can also be obtained based on these drawings.
[0013] FIG1 is a schematic diagram of an exemplary system applicable to the embodiment of the present application.
[0014] FIG2 is a flowchart of the steps of a method for generating training samples according to an embodiment of the present application.
[0015] FIG3 is a flowchart of an optional sub-step for generating a pixel change trend graph of the present application.
[0016] FIG4 is a flowchart of an optional sub-step of sub-step S1022 of the present application.
[0017] FIG5 is a schematic diagram of an example scenario in an embodiment of the present application.
[0018] FIG6 is a schematic structural diagram of an electronic device according to an embodiment of the present application. DETAILED DESCRIPTION
[0019] In order to enable those skilled in the art to better understand the technical solutions in the embodiments of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the embodiments of the present application, all other embodiments obtained by ordinary technicians in this field should fall within the scope of protection of the embodiments of the present application.
[0020] The specific implementation of the embodiment of the present application is further explained below in conjunction with the accompanying drawings of the embodiment of the present application.
[0021] Figure 1 illustrates an exemplary system applicable to the embodiments of the present application. As shown in Figure 1 , the system 100 may include a cloud service 102, a communication network 104, and / or one or more user devices 106, with Figure 1 illustrating multiple user devices. It should be noted that the embodiments of the present application can be independently implemented by the cloud service 102, or by a user device 106 with high-performance hardware and software.
[0022] The cloud server 102 can be any suitable device for storing information, data, programs, and / or any other suitable type of content, including but not limited to a distributed storage system, a server cluster, a computing cloud server cluster, and the like. In some embodiments, the cloud server 102 can perform any suitable function. For example, when the cloud server 102 independently implements the solution of the embodiments of the present application, in some embodiments, the cloud server 102 can generate training sample pairs for training a machine learning model for implementing software scanning. As an optional example, in some embodiments, the cloud server 102 can first obtain a pixel change trend graph indicating pixel changes between light and shadow areas and other areas in an image, where the light and shadow areas include shadow areas and / or reflective areas in the image; then, based on the pixel change trend graph, the cloud server 102 can perform light and shadow processing on the original scanned document image to obtain a corresponding simulated scanned document image having shadow areas and / or reflective areas; and then, based on the original scanned document image and the simulated scanned document image, generate training sample pairs. In some embodiments, after generating the training sample pairs, the cloud server 102 can use the training sample pairs to train the machine learning model for implementing software scanning. After the training is completed, the cloud service end 102 can receive the scanned image with shadows and / or reflections sent by the user device 106, and process it through the machine learning model to generate a cleaner scanned image, and then return it to the user device 106.
[0023] In some embodiments, the communication network 104 can be any suitable combination of one or more wired and / or wireless networks. For example, the communication network 104 can include any one or more of the following: the Internet, an intranet, a wide area network (WAN), a local area network (LAN), a wireless network, a digital subscriber line (DSL) network, a frame relay network, an asynchronous transfer mode (ATM) network, a virtual private network (VPN), and / or any other suitable communication network. The user device 106 can be connected to the communication network 104 via one or more communication links (e.g., communication link 112), and the communication network 104 can be linked to the cloud service end 102 via one or more communication links (e.g., communication link 114). The communication link can be any communication link suitable for transmitting data between the user device 106 and the cloud service end 102, such as a network link, a dial-up link, a wireless link, a hard-wired link, any other suitable communication link, or any suitable combination of such links.
[0024] User device 106 may include any one or more user devices suitable for presenting images, interacting with users, etc. When the user device 106 independently implements the solution of the embodiments of the present application, the user device 106 may generate training sample pairs for training the machine learning model for implementing software scanning. As an optional example, in some embodiments, the user device 106 may first obtain a pixel change trend graph indicating pixel changes between light and shadow areas and other areas in the image, where the light and shadow areas include shadow areas and / or reflective areas in the image; then, based on the pixel change trend graph, light and shadow processing may be performed on the original scanned document image to obtain a corresponding simulated scanned document image having shadow areas and / or reflective areas; and then, based on the original scanned document image and the simulated scanned document image, a training sample pair is generated. In some optional embodiments, the user device 106 may send the generated training sample pairs to the cloud server 102 to train the machine learning model used to generate the software scanning in the cloud server 102. In some embodiments, the user device 106 may include any suitable type of device. For example, in some embodiments, user device 106 may include a mobile device, a tablet computer, a laptop computer, a desktop computer, and / or any other suitable type of user device.
[0025] Based on the above system, an embodiment of the present application provides a training sample generation solution, which is described below through multiple embodiments.
[0026] FIG2 is a flowchart of the steps of a training sample generation method according to an embodiment of the present application. According to the first aspect of the present application, a training sample generation method is provided. As shown in FIG2 , the method includes steps S102, S104, and S106. Specifically:
[0027] S102: Obtaining a pixel change trend graph for indicating pixel changes between a light and shadow area and other areas except the light and shadow area in an image, wherein the light and shadow area includes a shadow area and / or a reflective area in the image.
[0028] In the present application, a pixel change trend graph can be obtained first, and the pixel change trend graph can be generated based on an image having both light and shadow areas and other areas (i.e., areas other than the light and shadow areas, which may include target content with actual information content, or background content, etc.). The pixel change trend graph can indicate the pixel changes between the light and shadow areas in the image and other areas other than the light and shadow areas. Through the pixel change trend graph, the pixel changes between the shadow areas and the non-shadow areas (i.e., other areas) under certain lighting conditions, and / or the pixel changes between the reflective areas and the non-reflective areas (i.e., other areas) can be determined. Therefore, in the subsequent step S104, through the pixel change trend graph, it is convenient to process the original scanned document image to obtain a suitable simulated scanned document image based on the pixel changes between the light and shadow areas in the image indicated by the pixel change trend graph and other areas other than the light and shadow areas.
[0029] The present application does not limit the method for generating the pixel change trend graph. For example, in some optional embodiments, referring to the flowchart in FIG3 , the pixel change trend graph can be generated by the optional implementation of the following sub-steps S1021 and S1022, specifically:
[0030] S1021: Perform light and shadow processing on the pure color image to obtain a simulated image with shadow areas and / or reflective areas.
[0031] Optionally, a pure color image may be obtained first, and its image size may be determined as required. The pure color image may be a pure color image of any color, and may be determined according to actual requirements.
[0032] In a preferred embodiment, the pure color image can be a pure white image. This has the following advantages: on the one hand, after light and shadow processing, the pure white image produces better shadow areas and / or reflective areas; on the other hand, for software scanning of documents, many actual documents are on white paper. Therefore, using a pure white image as a pure color image to perform light and shadow processing to generate a simulation image, and then generating a pixel change trend graph based on the simulation image, can make the training samples obtained step by step based on the pixel change trend graph better fit the actual state of the actual document, and better meet the requirements for training the machine learning model used to implement software scanning.
[0033] It should be noted that a pure color image generally refers to an image in which each pixel has the same pixel value. However, in this application, if necessary, an image in which each pixel has a pixel value within a predetermined value ± a predetermined value may be considered a pure color image. It should be understood that the predetermined value may be set to 5 or any other value as needed.
[0034] For example, taking a pure white image as an example, a pure color image generally refers to a pure color image in which the pixel value of each pixel is 255 (that is, the predetermined pixel value is 255). In this application, when the needs are met, a pure white image may refer to an image in which the pixel value of each pixel is greater than or equal to 250 and less than or equal to 255 (that is, the predetermined value is 5).
[0035] Optionally, after obtaining a solid color image, light and shadow processing can be performed on the solid color image. For example, light simulation can be performed on the solid color image to achieve light and shadow processing, thereby obtaining a simulated image with shadow areas and / or reflective areas. The light simulation can be performed using image rendering software, which may include but is not limited to at least one of Unity3D, Blender, and Gmic.
[0036] The specific implementation method of step S1021 is not limited in this application. For example, in some optional methods, step S1021 includes: performing light and shadow processing on the pure color image including light source type setting and / or light source position setting to obtain a simulated image with shadow areas and / or reflective areas.
[0037] Based on this, in this application, by performing light and shadow processing on the pure color image including light source type setting and / or light source position setting, a simulation image can be obtained conveniently, quickly and effectively, so as to facilitate data processing through the simulation image in subsequent steps.
[0038] For example, the light source type setting can be to set the type of light source. For example, the light source type can be changed by adjusting at least one of the type of light source (such as sunlight, incandescent lamp, LED lamp, flashlight light, etc.), the color of the light source (for example, white light, red light, yellow light, green light, etc.), the light intensity of the light source, etc. Through the light source type setting, different shadow areas and / or reflective areas can be displayed on the solid color image. Light and shadow effects.
[0039] For another example, the light source position setting can be used to set the position of the light source relative to the image. This can be used to create different shadow and / or reflective lighting effects on a solid color image. For example, if the light source is set 1 meter above the image, the light source illuminates from directly above; if the light source is set 3 meters from the left front of the image, the light source illuminates from the left front, and so on.
[0040] It should be understood that the above examples of light source type settings and light source position settings are only examples and do not constitute any limitation to the embodiments of the present application.
[0041] Optionally, lighting simulation can be performed on a solid color image using image rendering software, and lighting simulation can be performed by setting the light source type and / or light source position to achieve light and shadow processing, so that under the influence of different types and / or different positions of light sources, different shadow areas and / or reflective areas can be generated on the solid color image, generating different simulation images to improve richness.
[0042] That is, in the present application, by changing the light source type and / or light source position, a plurality of different simulation images with different shadow areas and / or reflection areas can be obtained quickly and easily for the same pure color image.
[0043] In some optional embodiments, the light and shadow processing further includes at least one of the following: background texture setting, background color setting, occluder position setting, and occluder type setting.
[0044] By using or combining these multiple optional light and shadow processing methods, the methods of obtaining simulation images in this application are enriched, and simulation images can be obtained conveniently, quickly and effectively, so that data processing can be performed through the simulation images in subsequent steps.
[0045] For example, background texture setting may refer to setting a texture pattern of a shadow area and / or a reflective area on the background of a solid color image. These texture patterns may be pre-stored after being segmented from an actual shadowed image and / or reflective image, and may be directly obtained when needed, or they may be texture patterns preset in the image rendering software, as long as the requirements are met. Different background texture settings may be performed on the solid color image as needed (for example, different texture patterns may be set, including but not limited to moiré textures, etc.), thereby obtaining simulated images with different light and shadow effects. Alternatively, background texture setting may also be achieved through image rendering software.
[0046] For another example, background color setting may refer to setting the background color of a solid color image. Since the background of a solid color image is a solid color, this setting may be equivalent to modifying the color of at least a portion of the solid color image. For example, the background color of a pure white image may be set, and the background color may be at least partially adjusted to gray, so that the light and shadow effects of being blocked may be simulated to a certain extent (which may be regarded as a shadow effect, i.e., as a shadow area); for another example, the background color of a pure white image may be set, and the background color may be at least partially adjusted to red, so that the light and shadow effects of being illuminated by a red light may be simulated to a certain extent (which may be regarded as a reflective effect, i.e., as a reflective area). It should be understood that the above examples are merely examples and do not constitute any limitation on the embodiments of the present application. Alternatively, the background color setting may also be implemented by image rendering software.
[0047] For another example, the position setting of the obstruction can be to set the position of the obstacle relative to the image. The position of the obstruction will have a great impact on the shadow and / or reflection, so by setting the position of the obstruction, different shadow areas and / or light and shadow effects of the reflection area can be displayed on the solid color image. For example, when the light source is set directly above the image, the position of the obstruction is set to the upper left, lower right, or outside the image, etc., which may result in different shadow areas and / or light and shadow effects of the reflection area. It should be understood that the above examples are only examples and do not constitute any limitation to the embodiments of the present application. Optionally, the background color setting can also be achieved through image rendering software.
[0048] For another example, the occlusion type setting may be to set the type of occlusion, for example, by adjusting at least one of the occlusion type (e.g., a cup, a ball, a finger, etc.), the occlusion size, and the occlusion shape (e.g., a cube, a cylinder, a sphere, or a regular or irregular shape). By setting the occlusion type, different shadow and / or reflective lighting effects can be displayed on a solid color image.
[0049] It should be noted that the above light source type setting, light source position setting, background texture setting, background color setting, obstruction position setting, obstruction type setting and other light and shadow processing methods can be used in combination as needed, and this application does not limit this.
[0050] S1022: Obtain a pixel change trend graph based on the pixel mean values of other areas except the shadow area and the reflective area in the simulation image, and the pixel values of the shadow area and / or the reflective area.
[0051] After obtaining the simulation image, the pixel mean of other areas in the simulation image except the shadow area and the reflective area can be calculated, and then combined with the pixel values of the shadow area and / or reflective area in the simulation image to obtain a pixel change trend graph indicating the pixel changes between the light and shadow areas in the image and other areas except the light and shadow areas.
[0052] Based on this, in this application, through the optional implementation method of sub-steps S1021 to S1022, a pixel change trend graph can be obtained conveniently, quickly and effectively for subsequent processing.
[0053] In some optional embodiments, referring to the flowchart in FIG4 , sub-step S1022 includes the following sub-steps S1022A and S1022B, specifically:
[0054] S1022A: Calculate and obtain the pixel mean of other areas of the simulation image except the shadow area and the reflection area.
[0055] Any pixel mean calculation method can be used to calculate the pixel mean of other areas in the simulation image except the shadow area and the reflective area. For example, the pixel mean can be calculated by summing the pixel values of each pixel in the other area and then averaging the sum to obtain the pixel mean of the other area.
[0056] S1022B: Use the pixel values of the shadow area and / or the reflective area, subtract the pixel mean, and obtain a corresponding pixel change trend graph.
[0057] In the present application, after obtaining the pixel mean, for each pixel in the shadow area and / or the reflective area, the pixel mean is subtracted from the pixel value of each pixel to obtain the corresponding pixel change value.
[0058] It should be understood that when the simulation image only includes shadow areas or reflective areas, it is only necessary to subtract the pixel mean from the pixel values of the shadow areas or reflective areas; if the simulation image includes both shadow areas and reflective areas, the pixel mean is subtracted from the pixel values of both the shadow areas and the reflective areas to obtain a pixel change trend graph with better effect.
[0059] Based on this, in this application, through the optional implementation method of sub-steps S1022A to S1022B, a pixel change trend graph can be obtained conveniently, quickly and effectively to facilitate subsequent processing.
[0060] In some optional embodiments, the "other areas except the shadow area and the reflective area" in step S1022 can be determined in the following manner: other areas except the shadow area and the reflective area in the simulation image are determined by at least one of the following processes, namely: image binarization processing of the simulation image, field analysis processing of the simulation image, labeling processing of the shadow area of the simulation image, and labeling processing of the reflective area of the simulation image.
[0061] In the present application, by performing image binarization on the simulation image, the other areas in the simulation image except the shadow area and the reflective area are determined. An optional process can be: by performing image binarization on the simulation image, the simulation image after binarization is divided into two parts with pixel values of 1 and 0 (herein, the binarized pixel values), one of which is the part of the shadow area and the reflective area, and the other is the part of the other areas except the shadow area and the reflective area. Then, from the simulation image, the image area corresponding to the other areas of the simulation image after binarization is determined, and the other areas except the shadow area and the reflective area can be determined in the simulation image. For example, the binarization process can be implemented by setting a preset pixel value. For example, according to a preset pixel value threshold, the pixel value of the image area with a pixel value greater than or equal to the pixel value threshold is set to 1, and the pixel value of the image area with a pixel value less than the pixel value threshold is set to 0. The shadow area and the reflective area can be set to 1, and the other parts can be set to 0; or the shadow area and the reflective area can be set to 0, and the other parts can be set to 1, and the setting can be as needed. Taking a pure white image as an example (referring to the above description, an image with a pixel value greater than or equal to 250 can be considered a pure white image), then the areas other than the shadow area and the reflective area in the simulation image are generally pure white. Therefore, after the simulation image is binarized, the pure white areas with pixel values greater than 250 can be set to 1, and the areas with pixel values less than 250 can be set to 0. Then, in the simulation image after binarization, the areas with pixel values of 0 are the areas other than the shadow area and the reflective area. For another example, in an optional embodiment, the simulation image may include an object, the shadow area may be generated by the object, and the reflective area may be on the object. Then, the object and its corresponding shadow area in the simulation image may be detected by target detection, and then the shadow area and the reflective area may be identified separately from the detection results. Then, the pixel values of the shadow area and the reflective area are set to 1, and then the area other than the shadow area and the reflective area in the detection results is set to 0 (or may also be set to 1, 0 is taken as an example here), and then the pixel values of the part other than the object and its corresponding shadow area in the simulation image are all set to 0, and then a simulation image after binarization processing is obtained. In the simulation image after binarization processing, the area with a pixel value of 0 is the other area except the shadow area and the reflective area. After obtaining the simulation image after binarization processing, the image area corresponding to the other area of the simulation image after binarization processing is determined from the simulation image, and the other area except the shadow area and the reflective area can be determined in the simulation image. It should be understood that the above description is only an example and is not intended to limit the present application.
[0062] The neighborhood analysis processing of the simulation image can be implemented by any neighborhood analysis algorithm of the image, which is not limited in this application.
[0063] For the labeling of the shadow area of the simulation image, it can be done manually, for example, by labeling each pixel point of the shadow area in the simulation image or labeling each pixel point of the boundary of the shadow area according to manual identification, or it can be implemented through a preset labeling model. For the labeling of the reflective area of the simulation image, it can also be done manually, that is, by labeling each pixel point of the reflective area in the simulation image or labeling each pixel point of the boundary of the reflective area according to manual identification, or it can be implemented through a preset labeling model. After the shadow area and the reflective area in the simulation image are labeled, the area other than the shadow area and the reflective area in the simulation image is the area other than the shadow area and the reflective area.
[0064] According to at least one processing method of the above embodiments, other areas in the simulation image except the shadow area and the reflective area can be determined, and the pixel mean of the other areas can be calculated. The optional method of calculating the pixel mean has been introduced in the previous article and will not be repeated here.
[0065] Based on this, in this application, the above multiple methods are used to determine other areas in the simulation image except the shadow area and the reflective area, and the other areas can be determined conveniently and effectively to facilitate the calculation of the pixel mean of other areas, and then facilitate the calculation of the pixel change trend diagram for processing.
[0066] Alternatively, other processing methods may be selected in the present application to determine other areas in the simulation image except the shadow area and the reflective area, which is not limited here.
[0067] S104: Based on the pixel change trend graph, perform light and shadow processing on the original scanned document image to obtain a corresponding simulated scanned document image with shadow areas and / or reflective areas.
[0068] In this application, light and shadow processing can be performed on an original scanned document image, and light and shadow processing can be performed on the original scanned document image through a pixel change trend graph to add shadows and / or reflections to part of the image area, thereby generating a simulated scanned document image with shadow areas and / or reflection areas.
[0069] The original scanned document image can be an image without shadows or reflections. For example, the original scanned document image can be a document image obtained by any appropriate means (such as manual photography, image generation or synthesis software, or scanning software after initial training, etc.) (the document may include but is not limited to books, invoices, contracts, business licenses, work documents, certificate materials, identity cards, driver's licenses, etc.). For example, taking the document as an example of an invoice, the original scanned document image can be obtained as follows: the user places the invoice to be scanned on a support (such as a desktop, etc.) or holds it in the hand, and then uses the camera on the mobile terminal to shoot the invoice from the direction of the invoice, thereby obtaining the corresponding original scanned document image. Of course, this is just an example and does not constitute any limitation to this application. It should be understood that in this case, the original scanned document image can be obtained by shooting when needed and then directly obtained, or it can be shot in advance and then stored in a predetermined storage space (including but not limited to storage media such as disks, hard drives, memories, or databases, etc.), and directly obtained from the storage space when needed. This application does not impose any limitations on this.
[0070] Alternatively, the original scanned document image can also be obtained through a pre-generated electronic file, for example, it can be obtained by printing or scanning through software. The electronic file can be in PDF format, or it can also be in other file formats. The electronic file can include one or more pages, or the original scanned document image can be generated based on these pages. Similarly, the original scanned document image generated in this way can also be first stored in a predetermined storage space (including but not limited to storage media such as disks, hard disks, memories, or databases, etc.), and directly obtained from the storage space when needed, and there is no restriction on this in this application.
[0071] Optionally, as mentioned above, different simulation images can be obtained by performing different light and shadow processing on the solid color image. Here, it is assumed that a total of m simulation images are obtained, and then the corresponding m pixel change trend graphs can be obtained through the m simulation images; furthermore, if there are n original scanned document images, then in theory, m pixel change trend graphs can be selected as needed and randomly combined with n original scanned document images for light and shadow processing to obtain a rich number of different simulated scanned document images, which significantly improves the efficiency of generating simulated scanned document images and generating training sample pairs, and can also further expand the diversity of subsequent generated training sample pairs, so that when the training sample pairs are subsequently used to train the machine learning model used to implement software scanning, it is beneficial to improve the generalization of the machine learning model.
[0072] For example, an original scanned document image can be subjected to light and shadow processing through a pixel change trend graph, an original scanned document image can be subjected to light and shadow processing through multiple pixel change trend graphs, an original scanned document image can be subjected to light and shadow processing by changing the direction of the pixel change trend graph (for example, the pixel change trend graph can be subjected to light and shadow processing upright, reversed, or diagonally, etc.), an original scanned document image can be subjected to light and shadow processing after the size of the pixel change trend graph is changed, and so on. Of course, these are all examples and are not intended to limit the present application.
[0073] In some optional embodiments, step S104 includes: superimposing the pixel change trend graph onto the original scanned document image to add a shadow to an area in the original scanned document image corresponding to the shadow area of the pixel change trend graph, and / or adding a reflection to an area in the original scanned document image corresponding to the reflective area of the pixel change trend graph.
[0074] Based on this, by superimposing the pixel change trend graph onto the original scanned document image, a shadow is added to the area in the original scanned document image corresponding to the shadow area of the pixel change trend graph, and / or a reflection is added to the area in the original scanned document image corresponding to the reflective area of the pixel change trend graph, so as to conveniently and effectively obtain the corresponding simulated scanned document image with a shadow area and / or a reflective area. Moreover, since the pixel change trend graph is directly superimposed on the original scanned document image in this application, it is easier to avoid the problem of pixel misalignment, and there will be no problem of changing the background color of the original scanned document image.
[0075] The pixel change trend graph can be superimposed on the original document scan image by any suitable image superposition algorithm, which is not limited in this application.
[0076] In other optional embodiments, step S104 includes: subtracting the pixel change trend graph from the original scanned document image to add reflections in the area of the original scanned document image corresponding to the shadow area of the pixel change trend graph, and / or adding shadows in the area of the original scanned document image corresponding to the reflective area of the pixel change trend graph.
[0077] Based on this, by subtracting the pixel change trend graph from the original scanned document image, reflection is added to the area in the original scanned document image corresponding to the shadow area of the pixel change trend graph, and / or shadow is added to the area in the original scanned document image corresponding to the reflective area of the pixel change trend graph, so as to conveniently and effectively obtain the corresponding simulated scanned document image with shadow area and / or reflective area. Moreover, since the pixel change trend graph is directly superimposed on the original scanned document image in this application, it is easier to avoid the problem of pixel misalignment, and there will be no problem of changing the background color of the original scanned document image.
[0078] The pixel change trend graph can be subtracted from the original document scan image by any suitable image subtraction algorithm, which is not limited in this application.
[0079] It should be understood that in this application, other processing methods can also be selected to perform light and shadow processing on the original scanned document image to obtain a simulated scanned document image, and there is no limitation here.
[0080] S106: Generate training sample pairs based on the original scanned document image and the simulated scanned document image.
[0081] Optionally, training sample pairs are generated based on the original scanned document image and the simulated scanned document image with shadow areas and / or reflective areas obtained in step S104 to facilitate training of a machine learning model for implementing software scanning.
[0082] For example, the original scanned document image and the simulated scanned document image can be directly combined to form a training sample pair and then output. Alternatively, in other optional embodiments, the original scanned document image and the simulated scanned document image can be combined and processed, including but not limited to appropriate cropping of the same area, and the same enlargement or reduction, to generate a training sample pair.
[0083] Based on the optional implementation in steps S102 to S106 above, the pixel change trend graph can be used to determine the pixel change between the shadow area and the non-shadow area (i.e., other areas) under certain lighting conditions, and / or the pixel change between the reflective area and the non-reflective area (i.e., other areas). Based on this, a light and shadow processing can be performed on a certain original scanned document image. The original scanned document image can be an image without shadows and reflections. In this case, the original scanned document image is light and shadow processed using the pixel change trend graph to add shadows and / or reflections to some areas of the image, thereby generating a simulated scanned document image with shadow areas and / or reflective areas. Furthermore, by combining the original scanned document image and the simulated scanned document image, a training sample pair can be generated for training a machine learning model for implementing software scanning. Therefore, on the one hand, there is no need for manual long-term learning and complex parameter adjustment of rendering software, which not only reduces the burden of manual operation, but also improves the efficiency of light and shadow processing of images; on the other hand, the pixel change trend graph describes the difference and change trend between shadows and / or reflections and other parts, which can be applied to original scanned document images in various situations and has high flexibility. If the pixel change trend graph includes multiple, different light and shadow processing can be performed on the original scanned document image according to needs to obtain more simulated scanned document images and expand the number of training samples; on the other hand, the pixel change trend graph will not change the background color of the original scanned document image, and will not bring about pixel deviation and random color changes. Therefore, the generated training sample pair is of better quality, which can avoid the problem of difficult convergence of model training in traditional methods.
[0084] In some optional embodiments, the training sample generation method further includes: training a machine learning model for implementing software scanning based on the training sample pairs.
[0085] The training samples obtained through the technical solution of this application can be used to train the machine learning model used to implement software scanning, and the data distribution can be closer to the real sample distribution, thereby improving the training effect of the machine learning model and its subsequent performance on real data. In addition, since the technical solution in this application can efficiently generate different training sample pairs, it avoids the tedious sample labeling and manual operation process, which is also convenient for improving the generalization of the machine learning model and helping the machine learning model to iterate quickly and specifically.
[0086] From the above optional embodiments, it can be seen that the training sample generation scheme in the present application can effectively generate simulated scanned images with reflections and shadows, and then generate training samples with the original scanned document images, simplifying the production and collection of samples into a process of collecting pixel change trend graphs, reducing the difficulty of producing training data for removing shadows and reflections, and breaking through the difficulty of removing shadows and reflections in the machine learning model of software scanning by scanning software. The generated samples can fill the gaps in the data set in this field, and are simple and easy to use, which is conducive to helping the machine learning model to iterate quickly.
[0087] FIG5 is a schematic diagram of a scene example in an embodiment of the present application. Combined with FIG5 , the training sample generation method in the embodiment of the present application is understood as a whole. Referring to FIG5 , step S102 in the present application: "Obtain a pixel change trend diagram for indicating pixel changes between light and shadow areas and other areas other than light and shadow areas in the image", as shown in FIG5 , the pixel change trend diagram is obtained by processing a pure color image. Specifically, a simulation diagram can be generated for the pure color image, and then a pixel change trend diagram is obtained according to the simulation diagram (as shown in FIG5 , the pure color image and the pixel change trend diagram are composed of pixels. In the figure, a rectangular grid is used to represent pixels, and pixel values are marked in the rectangular grid. For example, the pixel value of each pixel in the pure color image in FIG5 is 255, which is a pure white image. The pixel values of the pixels in the pixel change trend diagram change sequentially. It should be understood that FIG5 is only a schematic diagram here, and it is more complicated in practice. It does not serve as any limitation to the present application). Through the pixel change trend diagram, the shadow area and the shadowless area (i.e., other areas) under a certain lighting condition can be determined. The pixel changes between the reflective areas and the non-reflective areas (i.e., other areas); then, step S104 is performed: "Based on the pixel change trend graph, the original scanned document image is subjected to light and shadow processing to obtain a corresponding simulated scanned document image with shadow areas and / or reflective areas". The original scanned document image may be an image without shadows and reflections. The original scanned document image is subjected to light and shadow processing through the pixel change trend graph to add shadows and / or reflections to part of the image to generate a simulated scanned document image with shadow areas and / or reflective areas. The shadow areas and / or reflective areas in the simulated scanned document image are represented by dotted lines in FIG5 ; then, step S106 is performed, and a training sample pair for training a machine learning model for implementing software scanning can be generated based on the original scanned document image and the simulated scanned document image. It should be understood that the scenario example shown in FIG5 is only used to facilitate the understanding of the embodiments of the present application and does not constitute any limitation to the present application.
[0088] In summary, the solution provided in the embodiments of the present application can determine the pixel changes between shadow areas and non-shadow areas (i.e., other areas) under certain lighting conditions, and / or the pixel changes between reflective areas and non-reflective areas (i.e., other areas) through a pixel change trend graph. Based on this, a light and shadow processing can be performed on a certain original scanned document image. The original scanned document image can be an image without shadows and without reflections. In this case, the light and shadow processing is performed on the original scanned document image through the pixel change trend graph to add shadows and / or reflections to some areas of the image, thereby generating a simulated scanned document image with shadow areas and / or reflective areas. Furthermore, by combining the original scanned document image and the simulated scanned document image, a training sample pair can be generated for training a machine learning model for implementing software scanning. Therefore, on the one hand, there is no need for manual long-term learning and complex parameter adjustment of rendering software, which not only reduces the burden of manual operation, but also improves the efficiency of light and shadow processing of images; on the other hand, the pixel change trend graph describes the difference and change trend between shadows and / or reflections and other parts, which can be applied to original scanned document images in various situations and has high flexibility. If the pixel change trend graph includes multiple, different light and shadow processing can be performed on the original scanned document image according to needs to obtain more simulated scanned document images and expand the number of training samples; on the other hand, the pixel change trend graph will not change the background color of the original scanned document image, and will not bring about pixel deviation and random color changes. Therefore, the generated training sample pair is of better quality, which can avoid the problem of difficult convergence of model training in traditional methods.
[0089] It is understandable that the above description of the training sample generation method is merely an exemplary description of the present application and does not constitute any limitation to the present application.
[0090] According to a second aspect of an embodiment of the present application, an electronic device is provided, comprising: a processor, a memory, a communication interface, and a communication bus, wherein the processor, the memory, and the communication interface communicate with each other via the communication bus; the memory is configured to store at least one executable instruction, wherein the executable instruction causes the processor to perform an operation corresponding to the method described in the first aspect. Referring to FIG6 , a schematic diagram of the structure of an electronic device according to an embodiment of the present application is shown, and the specific embodiments of the present application do not limit the specific implementation of the electronic device.
[0091] As shown in FIG6 , the electronic device 600 may include a processor 602 , a communications interface 604 , a memory 606 , and a communication bus 608 .
[0092] in:
[0093] The processor 602 , the communication interface 604 , and the memory 606 communicate with each other via a communication bus 608 .
[0094] The communication interface 604 is used to communicate with other electronic devices or servers.
[0095] The processor 602 is configured to execute the program 610 , and specifically to execute the relevant steps in the above-mentioned training sample generation method embodiment.
[0096] Specifically, the program 610 may include program codes, which include computer operation instructions.
[0097] Processor 602 may be a CPU, a Graphics Processing Unit (GPU), an Application Specific Integrated Circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present application. The one or more processors included in the smart device may be processors of the same type, such as one or more CPUs, or may be processors of different types, such as one or more CPUs and one or more ASICs.
[0098] The memory 606 is used to store the program 610. The memory 606 may include a high-speed RAM memory, and may also include a non-volatile memory (non-volatile memory), such as at least one disk memory.
[0099] The program 610 may include multiple computer instructions. Specifically, the program 610 may enable the processor 602 to execute operations corresponding to the training sample generation method described in any of the aforementioned method embodiments through the multiple computer instructions.
[0100] The specific implementation of each step in program 610 can refer to the corresponding description of the corresponding steps and units in the above-mentioned method embodiment, and has corresponding beneficial effects, which will not be repeated here. Those skilled in the art will clearly understand that for the convenience and brevity of description, the specific working process of the above-mentioned devices and modules can refer to the corresponding process description in the above-mentioned method embodiment, and will not be repeated here.
[0101] According to a third aspect of the embodiments of the present application, the embodiments of the present application further provide a computer storage medium having a computer program stored thereon, which, when executed by a processor, implements the method described in any of the aforementioned method embodiments. The computer storage medium includes, but is not limited to, a compact disc read-only memory (CD-ROM), random access memory (RAM), a floppy disk, a hard disk, or a magneto-optical disk.
[0102] An embodiment of the present application further provides a computer program product, including computer instructions, which instruct a computing device to execute operations corresponding to any training sample generation method in the above-mentioned multiple method embodiments.
[0103] The electronic device 600 / computer storage medium / computer program product embodiment in the embodiment of the present application has been described in detail in the aforementioned training sample generation method embodiment, so its relevant content and beneficial effects can be understood with reference to the above-mentioned method embodiment and will not be repeated here.
[0104] In addition, it should be noted that the user-related information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to sample data used to train the model, data used for analysis, stored data, displayed data, etc.) involved in the embodiments of this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data need to comply with the relevant laws, regulations and standards of relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse.
[0105] It should be pointed out that, according to the needs of implementation, the various components / steps described in the embodiments of the present application can be split into more components / steps, or two or more components / steps or partial operations of components / steps can be combined into new components / steps to achieve the purpose of the embodiments of the present application.
[0106] The above-mentioned method according to the embodiment of the present application can be implemented in hardware, firmware, or as software or computer code that can be stored in a recording medium (such as a CD-ROM, RAM, floppy disk, hard disk or magneto-optical disk), or as computer code that is originally stored in a remote recording medium or a non-temporary machine-readable medium downloaded via a network and will be stored in a local recording medium, so that the method described herein can be stored in such software processing on a recording medium using a general-purpose computer, a dedicated processor or programmable or dedicated hardware (such as an application-specific integrated circuit (ASIC) or a field programmable gate array (FPGA)). It can be understood that a computer, a processor, a microprocessor controller or programmable hardware includes a storage component (e.g., random access memory (RAM), read-only memory (ROM), flash memory, etc.) that can store or receive software or computer code. When the software or computer code is accessed and executed by a computer, a processor or hardware, the method described herein is implemented. In addition, when a general-purpose computer accesses the code for implementing the method shown here, the execution of the code converts the general-purpose computer into a dedicated computer for executing the method shown here.
[0107] Those skilled in the art will appreciate that the units and method steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for specific applications, but such implementation should not be considered to be beyond the scope of the embodiments of this application.
[0108] The above implementation methods are only used to illustrate the embodiments of the present application, and are not intended to limit the embodiments of the present application. Ordinary technicians in the relevant technical field can make various changes and modifications without departing from the spirit and scope of the embodiments of the present application. Therefore, all equivalent technical solutions also fall within the scope of the embodiments of the present application, and the scope of patent protection of the embodiments of the present application should be defined by the claims.
Claims
1. A training sample generation method, comprising: Obtaining a pixel change trend graph for indicating pixel changes between a light and shadow area and other areas except the light and shadow area in an image, wherein the light and shadow area includes a shadow area and / or a reflective area in the image; Based on the pixel change trend graph, performing light and shadow processing on the original scanned document image to obtain a corresponding simulated scanned document image having a shadow area and / or a reflective area; A training sample pair is generated according to the original scanned document image and the simulated scanned document image.
2. The method according to claim 1, wherein: The pixel change trend graph is generated in the following manner: Performing light and shadow processing on the pure color image to obtain a simulated image with shadow areas and / or reflective areas; A pixel change trend graph is obtained based on the pixel mean values of other regions in the simulation graph except the shadow region and the reflective region, and the pixel values of the shadow region and / or the reflective region.
3. The method according to claim 2, wherein: The step of obtaining a pixel change trend graph based on the pixel mean values of other regions in the simulation graph except the shadow region and the reflective region, and the pixel values of the shadow region and / or the reflective region, comprises: Calculate and obtain the pixel mean of other areas in the simulation image except the shadow area and the reflective area; The pixel values of the shadow area and / or the reflective area are used to subtract the pixel mean value to obtain a corresponding pixel change trend graph.
4. The method according to claim 2 or 3, wherein: The step of performing light and shadow processing on the pure color image to obtain a simulated image having a shadow area and / or a reflective area includes: Light and shadow processing including light source type setting and / or light source position setting is performed on the pure color image to obtain a simulated image with shadow areas and / or reflective areas.
5. The method according to any one of claims 2 to 4, wherein: The light and shadow processing also includes at least one of the following: background texture setting, background color setting, occluder position setting, and occluder type setting.
6. The method according to any one of claims 2 to 5, wherein the other regions are determined by: The other areas in the simulation image except the shadow area and the reflective area are determined by at least one of the following processing: performing image binarization processing on the simulation image, performing neighborhood analysis processing on the simulation image, performing shadow area labeling processing on the simulation image, and performing reflective area labeling processing on the simulation image.
7. The method according to any one of claims 2 to 6, wherein: The pure color image is a pure white image.
8. The method according to claim 1, wherein: The method of performing light and shadow processing on the original scanned document image based on the pixel change trend graph to obtain a corresponding simulated scanned document image having a shadow area and / or a reflective area includes: The pixel change trend graph is superimposed on the original scanned document image to add a shadow to an area in the original scanned document image that corresponds to the shadow area of the pixel change trend graph, and / or to add a reflection to an area in the original scanned document image that corresponds to the reflective area of the pixel change trend graph.
9. The method according to claim 1, wherein: The method of performing light and shadow processing on the original scanned document image based on the pixel change trend graph to obtain a corresponding simulated scanned document image having a shadow area and / or a reflective area includes: The pixel change trend graph is subtracted from the original scanned document image to add reflection to an area in the original scanned document image corresponding to a shadow area of the pixel change trend graph, and / or a shadow is added to an area in the original scanned document image corresponding to a reflective area of the pixel change trend graph.
10. An electronic device comprising: A processor, a memory, a communication interface and a communication bus, wherein the processor, the memory and the communication interface communicate with each other via the communication bus; The memory is used to store at least one executable instruction, and the executable instruction enables the processor to perform an operation corresponding to the method according to any one of claims 1 to 9.
11. A computer storage medium having a computer program stored thereon, wherein when the program is executed by a processor, the method according to any one of claims 1 to 9 is implemented.
12. A computer program product, comprising computer instructions, wherein the computer instructions instruct a computing device to execute operations corresponding to the method according to any one of claims 1 to 9.
Citation Information
Patent Citations
Document image processing method and device and training sample generation method and device
CN113744172A
Document image processing method and document image processing device
CN115861114A
Document image registration data synthesis method, system and device and medium
CN116452641A
Training sample generation method and electronic equipment
CN117745621A
Enhancing light text in scanned documents while preserving document fidelity
US20230260091A1