Training sample generation method, model training method, electronic device, and storage medium
By using pixel change trend charts in the generation of training samples for light and shadow processing, combined with random resolution adjustment, the problem that training sample processing in the prior art causes the model to be unable to effectively restore the scanned image, achieving a better image restoration effect.
Patent Information
- Application Number
- PCT/CN2024/130582
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-18
- Filing Date
- 2024-11-07
- Publication Date
- 2025-06-26
AI Technical Summary
In the prior art, when generating training samples for training machine learning models for scanning image processing, enhanced filter processing will remove shadows and reflections from the image, while also losing original information, such as image background color, texture, and style, resulting in the model being unable to effectively restore the true appearance of the scanned image.
By collecting pixel change trend charts from the original scan sample image, performing light and shadow processing, generating simulated scan sample images with shadowed areas and/or reflective areas, and combining the original scan sample images with random resolution adjustments, training sample pairs are generated to effectively train the machine learning model.
It realizes the generation of training samples that can effectively train machine learning models, so that the model can retain the original information of the image while removing shadows and reflections, thereby achieving a good image restoration effect.
Smart Images

Figure CN2024130582_26062025_PF_FP_ABST
Abstract
Description
Training sample generation method, model training method, electronic device and storage medium
[0001] This application claims priority to the Chinese patent application filed with the China Patent Office on December 18, 2023, with application number 202311750589.2, and application name “Training Sample Generation Method, Model Training Method, Electronic Device and Storage Medium”, all contents of which are incorporated by reference into this application. Technical Field
[0002] The embodiments of the present application relate to the field of computer technology, and in particular to a training sample generation method, a model training method, an electronic device, and a storage medium. Background Art
[0003] With the development of computer technology, people often need to electronically scan documents in their daily lives and work, such as electronic books, invoice reimbursement, work document scanning and printing, document material scanning, and other application scenarios.
[0004] Because professional scanners are expensive, scanning software has come into being. With scanning software, users can scan documents conveniently and at low cost anytime and anywhere. With the widespread use of scanning software, scanning scenarios are becoming more and more diverse. For example, in some scenarios, there is no need for the effect of an enhancement filter with high contrast for the scanned image, and it is only necessary to remove the shadows and reflections in the captured scanned image. To this end, training samples that are adapted to this requirement are needed to iteratively update the machine learning model that implements software scanning. In one existing method, the original scanned image with shadows and / or reflections is processed with an enhancement filter, and the obtained enhanced scanned image and the original scanned image are combined into an image pair to train the machine learning model that implements software scanning.
[0005] However, this enhanced filter processing method not only removes shadows and / or reflections, but also removes the original information in the original scanned image, such as image background color, texture, style, etc., so that the machine learning model trained using such training samples cannot effectively restore the true appearance of the scanned image, and the restoration effect is poor.
[0006] Therefore, how to generate training samples that can effectively train the above-mentioned machine learning models so that the trained models have better restoration effects has become an urgent problem to be solved.
[0007] Summary of the Invention
[0008] In view of this, an embodiment of the present application provides a training sample generation and model training solution to at least partially solve the above problems.
[0009] According to a first aspect of an embodiment of the present application, a training sample generation method is provided, comprising: collecting an original scanned sample image from an original sample set, and collecting a pixel change trend graph from a pixel change trend graph set, wherein the pixel change trend graph is used to indicate pixel changes between light and shadow areas in an image and areas other than the light and shadow areas, and the light and shadow areas include shadow areas and / or reflective areas in the image; performing random resolution adjustment on the original scanned sample image to obtain an adjusted sample image; performing light and shadow processing on the adjusted sample image based on the pixel change trend graph to obtain a corresponding simulated scanned sample image having a shadow area and / or a reflective area; generating a training sample pair for online training of a machine learning model based on the adjusted sample image and the simulated scanned sample image, wherein the machine learning model is used to process the scanned image.
[0010] According to the second aspect of the embodiments of the present application, a model training method is provided for online training of a machine learning model for processing scanned images, the method comprising: obtaining a training sample pair, wherein the training sample pair includes: an image after random resolution adjustment based on the original scanned sample image, and an image with shadow areas and / or reflective areas obtained after light and shadow processing of the adjusted image based on a pixel change trend graph; and online training of the machine learning model based on the training sample pair and a preset loss function.
[0011] According to a third aspect of an embodiment of the present application, an electronic device is provided, comprising: a processor, a memory, a communication interface and a communication bus, wherein the processor, the memory and the communication interface communicate with each other through the communication bus; the memory is used to store at least one executable instruction, and the executable instruction enables the processor to perform an operation corresponding to the method described in the first aspect or the method described in the second aspect.
[0012] According to a fourth aspect of the embodiments of the present application, a computer storage medium is provided, on which a computer program is stored. When the program is executed by a processor, the method described in the first aspect or the method described in the second aspect is implemented.
[0013] According to a fifth aspect of an embodiment of the present application, a computer program product is provided, comprising computer instructions, which instruct a computing device to perform operations corresponding to the method described in the first aspect or the method described in the second aspect.
[0014] According to the solution provided in the embodiment of the present application, for the training of a machine learning model for processing scanned images, when generating training samples for it, the original scanned sample image will be subjected to light and shadow processing based on the pixel change trend diagram, and a simulated scanned sample image will be generated based on the light and shadow processing. In addition, the original scanned sample image will be subjected to random resolution adjustment to generate an adjusted sample image. Based on the simulated scanned sample image and the adjusted sample image, a training sample pair for training the machine learning model is generated. Among them, through light and shadow processing, shadow areas and / or reflective areas can be added to the original scanned sample image, so that the generated simulated scanned sample image has shadow areas and / or reflective areas. Therefore, the machine learning model, through learning the training sample pair, enables the trained model to effectively remove shadows and / or reflections in the image, and avoid improper removal of other original information in the image, such as background color, texture, style, etc., thereby achieving the effect of effectively restoring the original appearance of the image. By randomly adjusting the resolution of the original scanned sample images, the sample images can have richer and more diverse resolutions and a wider sample coverage. After learning, the machine learning model can have better compatibility processing capabilities and better restoration effects for images of various resolutions after subsequent training is completed. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments recorded in the embodiments of the present application. For ordinary technicians in this field, other drawings can also be obtained based on these drawings.
[0016] FIG1 is a schematic diagram of an exemplary system applicable to the embodiment of the present application.
[0017] FIG2 is a flowchart of the steps of a method for generating training samples according to an embodiment of the present application.
[0018] FIG3 is a flowchart of an optional step for generating a set of pixel change trend graphs according to an embodiment of the present application.
[0019] FIG4 is a flowchart of an optional sub-step of step S301 in an embodiment of the present application.
[0020] FIG5 is a schematic diagram of a scenario example of a training sample generation solution according to an embodiment of the present application.
[0021] FIG6 is a flowchart of the steps of a model training method according to an embodiment of the present application.
[0022] Figure 7 is a schematic diagram of a scenario example of the model training solution of an embodiment of the present application.
[0023] FIG8 is a schematic structural diagram of an electronic device according to an embodiment of the present application. DETAILED DESCRIPTION
[0024] In order to enable those skilled in the art to better understand the technical solutions in the embodiments of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the embodiments of the present application, all other embodiments obtained by ordinary technicians in this field should fall within the scope of protection of the embodiments of the present application.
[0025] The specific implementation of the embodiment of the present application is further explained below in conjunction with the accompanying drawings of the embodiment of the present application.
[0026] Figure 1 illustrates an exemplary system applicable to the embodiments of the present application. As shown in Figure 1 , the system 100 may include a cloud service 102, a communication network 104, and / or one or more user devices 106, with Figure 1 illustrating multiple user devices. It should be noted that the embodiments of the present application can be independently implemented by the cloud service 102, or by a user device 106 with high-performance hardware and software.
[0027] The cloud server 102 can be any suitable device for storing information, data, programs, and / or any other suitable type of content, including but not limited to a distributed storage system, a server cluster, a computing cloud server cluster, etc. In some embodiments, the cloud server 102 can perform any suitable function. For example, when the cloud server 102 independently implements the solution of the embodiments of the present application, in some embodiments, the cloud server 102 can generate training sample pairs for training a machine learning model for processing scanned images. As an optional example, in some embodiments, the cloud service end 102 may first collect an original scan sample image from the original sample set, and collect a pixel change trend graph from the pixel change trend graph set, wherein the pixel change trend graph is used to indicate the pixel changes between the light and shadow area and other areas in the image except the light and shadow area, and the light and shadow area includes the shadow area and / or the reflective area in the image; then, the original scan sample image may be randomly adjusted in resolution to obtain an adjusted sample image; thereafter, the adjusted sample image may be subjected to light and shadow processing based on the pixel change trend graph to obtain a corresponding simulated scan sample image having a shadow area and / or a reflective area; and then, based on the adjusted sample image and the simulated scan sample image, a training sample pair for online training of the machine learning model is generated. In some embodiments, after generating the training sample pair, the cloud service end 102 may use the training sample pair and a preset loss function to perform online training on the machine learning model for processing the scanned image. After the training is completed, the cloud service end 102 can receive the scanned image with shadows and / or reflections sent by the user device 106, and process it through the machine learning model to generate a cleaner scanned image, and then return it to the user device 106.
[0028] In some embodiments, the communication network 104 can be any suitable combination of one or more wired and / or wireless networks. For example, the communication network 104 can include any one or more of the following: the Internet, an intranet, a wide area network (WAN), a local area network (LAN), a wireless network, a digital subscriber line (DSL) network, a frame relay network, an asynchronous transfer mode (ATM) network, a virtual private network (VPN), and / or any other suitable communication network. The user device 106 can be connected to the communication network 104 via one or more communication links (e.g., communication link 112), and the communication network 104 can be linked to the cloud service end 102 via one or more communication links (e.g., communication link 114). The communication link can be any communication link suitable for transmitting data between the user device 106 and the cloud service end 102, such as a network link, a dial-up link, a wireless link, a hard-wired link, any other suitable communication link, or any suitable combination of such links.
[0029] The user device 106 may include any one or more user devices suitable for presenting images, interacting with users, etc. When the solution of the embodiment of the present application is independently completed by the user device 106, the user device 106 can generate a training sample pair for training a machine learning model for processing scanned images. As an optional example, in some embodiments, the user device 106 can first collect an original scanned sample image from the original sample set, and collect a pixel change trend graph from the pixel change trend graph set, wherein the pixel change trend graph is used to indicate the pixel changes between the light and shadow area in the image and other areas other than the light and shadow area, and the light and shadow area includes the shadow area and / or the reflective area in the image; then, the original scanned sample image can be randomly adjusted in resolution to obtain an adjusted sample image; thereafter, the adjusted sample image can be subjected to light and shadow processing based on the pixel change trend graph to obtain a corresponding simulated scanned sample image with a shadow area and / or a reflective area; and then, based on the adjusted sample image and the simulated scanned sample image, a training sample pair for online training of the machine learning model is generated. In some optional embodiments, the user device 106 may send the generated training sample pairs to the cloud server 102 to perform online training on the machine learning model for processing the scanned image in the cloud server 102. In some embodiments, the user device 106 may include any suitable type of device. For example, in some embodiments, the user device 106 may include a mobile device, a tablet computer, a laptop computer, a desktop computer, and / or any other suitable type of user device.
[0030] Based on the above system, the embodiments of the present application provide a training sample generation scheme and a model training scheme, which are described below through multiple embodiments.
[0031] FIG2 is a flowchart of the steps of a training sample generation method according to an embodiment of the present application. According to the first aspect of the present application, a training sample generation method is provided. As shown in FIG2 , the method includes steps S202, S204, S206 and S208. Specifically:
[0032] S202: Collecting original scan sample images from the original sample set, and collecting pixel change trend graphs from the pixel change trend graph set.
[0033] In the embodiments of the present application, the original scanned sample images may be document sample images (documents may include, but are not limited to, books, invoices, contracts, business licenses, work documents, credentials, ID cards, driver's licenses, etc.) obtained by any appropriate means (such as manual photography, image generation or synthesis software, or scanning software after initial training). The original sample set may be a set consisting of multiple original scanned sample images. After obtaining the multiple original scanned sample images, the multiple original scanned sample images may be used to generate an original sample set, so as to facilitate subsequent collection and use from the original sample set when needed.
[0034] For example, taking the invoice as an example, the original scan sample image can be obtained in the following manner: the user places the invoice to be scanned on a support (such as a desktop, etc.) or holds it in his hand, and then uses the camera on the mobile terminal to shoot the invoice from the direction of the invoice, thereby obtaining the corresponding original scan sample image. Of course, this is just an example and does not serve as any limitation to this application. In this case, the original scan sample image can be obtained by shooting when needed and then directly obtained, or it can be shot in advance and then stored in a predetermined storage space (including but not limited to storage media such as disks, hard disks, memories, or databases, etc.), and directly obtained from the storage space when needed. This application does not impose any restrictions on this.
[0035] Alternatively, the original scanned sample image can also be obtained through a pre-generated electronic file, for example, it can be obtained by printing or scanning through software. The electronic file can be in PDF format, or it can also be in other file formats. The electronic file can include one or more pages, or the original scanned sample image can be generated based on these pages. Similarly, the original scanned sample image generated in this way can also be first stored in a predetermined storage space (including but not limited to storage media such as disks, hard disks, memories, or databases, etc.), and directly obtained from the storage space when needed. There is no restriction on this in this application.
[0036] In an embodiment of the present application, the pixel change trend diagram is used to indicate pixel changes in the light and shadow area in the image and other areas except the light and shadow area, and the light and shadow area includes the shadow area and / or reflective area in the image.
[0037] In the present application, a pixel change trend graph can be collected first. The pixel change trend graph can be generated based on an image that has both light and shadow areas and other areas (i.e., areas other than the light and shadow areas, which may include target content with actual information content, or background content, etc.). The pixel change trend graph can indicate the pixel changes between the light and shadow areas in the image and other areas other than the light and shadow areas. Through the pixel change trend graph, the pixel changes between the shadow area and the non-shadow area (i.e., other areas) under certain lighting conditions, and / or the pixel changes between the reflective area and the non-reflective area (i.e., other areas) can be determined.
[0038] In the present application, the pixel change trend graph set can be a set composed of multiple pixel change trend graphs. After obtaining multiple pixel change trend graphs, an original sample set can be generated with multiple original scanned sample images, so as to facilitate subsequent collection and use from the pixel change trend graph set when needed.
[0039] Optionally, each pixel change trend graph in the pixel change trend graph set in the present application may include an original pixel change trend graph and a new pixel change trend graph generated by processing the original pixel change trend graph, so that the images in the pixel change trend graph set are more diverse. The generation method of the pixel change trend graph set is not limited in the present application. Based on the above description, in some optional embodiments, with reference to the flowchart shown in Figure 3, the pixel change trend graph set is generated by the following steps S301, S302 and S303, specifically:
[0040] S301: Obtaining an original pixel change trend graph.
[0041] In the present application, the original pixel change trend graph may be pre-generated. The present application does not specifically limit the method for generating the original pixel change trend graph. In some optional embodiments, referring to the flowchart shown in FIG. 4 , the graph may be generated by the following sub-steps S3011 and S3012, specifically:
[0042] S3011: Perform light and shadow processing on the pure color image to obtain a simulated image with shadow areas and / or reflective areas.
[0043] Optionally, a pure color image may be obtained first, and its image size may be determined as required. The pure color image may be a pure color image of any color, and may be determined according to actual requirements.
[0044] In a preferred embodiment, the pure color image can be a pure white image. This has the following advantages: on the one hand, after light and shadow processing, the shadow area and / or reflective area of the pure white image is better; on the other hand, for software scanning of documents, many actual documents are on white paper. Therefore, using a pure white image as a pure color image to perform light and shadow processing to generate a simulation image, and then generating an original pixel change trend map based on the simulation image, can make the training sample pairs obtained by step-by-step processing based on the original pixel change trend map better fit the actual state of the actual document, and better meet the requirements for training machine learning models used to process scanned images.
[0045] It should be noted that a pure color image generally refers to an image in which each pixel has the same pixel value. However, in this application, if necessary, an image in which each pixel has a pixel value within a predetermined value ± a predetermined value may be considered a pure color image. It should be understood that the predetermined value may be set to 5 or any other value as needed.
[0046] For example, taking a pure white image as an example, a pure color image generally refers to a pure color image in which the pixel value of each pixel is 255 (that is, the predetermined pixel value is 255). In this application, when the needs are met, a pure white image may refer to an image in which the pixel value of each pixel is greater than or equal to 250 and less than or equal to 255 (that is, the predetermined value is 5).
[0047] Optionally, after obtaining a solid color image, light and shadow processing can be performed on the solid color image. For example, light simulation can be performed on the solid color image to achieve light and shadow processing, thereby obtaining a simulated image with shadow areas and / or reflective areas. The light simulation can be performed using image rendering software, which may include but is not limited to at least one of Unity3D, Blender, and Gmic.
[0048] The specific implementation method of step S3011 is not limited in this application. For example, in some optional methods, step S3011 includes: performing light and shadow processing on the pure color image including light source type setting and / or light source position setting to obtain a simulated image with shadow areas and / or reflective areas.
[0049] Based on this, in this application, by performing light and shadow processing on the pure color image including light source type setting and / or light source position setting, a simulation image can be obtained conveniently, quickly and effectively, so as to facilitate data processing through the simulation image in subsequent steps.
[0050] For example, the light source type setting can be to set the type of light source. For example, the light source type can be changed by adjusting at least one of the type of light source (such as sunlight, incandescent lamp, LED lamp, flashlight light, etc.), the color of the light source (for example, white light, red light, yellow light, green light, etc.), the light intensity of the light source, etc. Through the light source type setting, different shadow areas and / or reflective areas can be displayed on the solid color image. Light and shadow effects.
[0051] For another example, the light source position setting can be used to set the position of the light source relative to the image. This can be used to create different shadow and / or reflective lighting effects on a solid color image. For example, if the light source is set 1 meter above the image, the light source illuminates from directly above; if the light source is set 3 meters from the left front of the image, the light source illuminates from the left front, and so on.
[0052] It should be understood that the above examples of light source type settings and light source position settings are only examples and do not constitute any limitation to the embodiments of the present application.
[0053] Optionally, lighting simulation can be performed on a solid color image using image rendering software, and lighting simulation can be performed by setting the light source type and / or light source position to achieve light and shadow processing, so that under the influence of different types and / or different positions of light sources, different shadow areas and / or reflective areas can be generated on the solid color image, generating different simulation images to improve richness.
[0054] That is, in the present application, by changing the light source type and / or light source position, a plurality of different simulation images with different shadow areas and / or reflection areas can be obtained quickly and easily for the same pure color image.
[0055] In some optional embodiments, the light and shadow processing further includes at least one of the following: background texture setting, background color setting, occluder position setting, and occluder type setting.
[0056] In this application, these multiple optional light and shadow processing methods are used or combined to enrich the methods of obtaining simulation images in this application, and the simulation images can be obtained conveniently, quickly and effectively, so that data processing can be performed through the simulation images in subsequent steps.
[0057] For example, background texture setting may refer to setting a texture pattern of a shadow area and / or a reflective area on the background of a solid color image. These texture patterns may be pre-stored after being segmented from an actual shadowed image and / or reflective image, and may be directly obtained when needed, or they may be texture patterns preset in the image rendering software, as long as the requirements are met. Different background texture settings may be performed on the solid color image as needed (for example, different texture patterns may be set, including but not limited to moiré textures, etc.), thereby obtaining simulated images with different light and shadow effects. Alternatively, background texture setting may also be achieved through image rendering software.
[0058] For another example, background color setting may refer to setting the background color of a solid color image. Since the background of a solid color image is a solid color, this setting may be equivalent to modifying the color of at least a portion of the solid color image. For example, the background color of a pure white image may be set, and the background color may be at least partially adjusted to gray, so that the light and shadow effects of being blocked may be simulated to a certain extent (which may be regarded as a shadow effect, i.e., as a shadow area); for another example, the background color of a pure white image may be set, and the background color may be at least partially adjusted to red, so that the light and shadow effects of being illuminated by a red light may be simulated to a certain extent (which may be regarded as a reflective effect, i.e., as a reflective area). It should be understood that the above examples are merely examples and do not constitute any limitation on the embodiments of the present application. Alternatively, the background color setting may also be implemented by image rendering software.
[0059] For another example, the position setting of the obstruction can be to set the position of the obstacle relative to the image. The position of the obstruction will have a great impact on the shadow and / or reflection, so by setting the position of the obstruction, different shadow areas and / or light and shadow effects of the reflection area can be displayed on the solid color image. For example, when the light source is set directly above the image, the position of the obstruction is set to the upper left, lower right, or outside the image, etc., which may result in different shadow areas and / or light and shadow effects of the reflection area. It should be understood that the above examples are only examples and do not constitute any limitation to the embodiments of the present application. Optionally, the background color setting can also be achieved through image rendering software.
[0060] For another example, the occlusion type setting may be to set the type of occlusion, for example, by adjusting at least one of the occlusion type (e.g., a cup, a ball, a finger, etc.), the occlusion size, and the occlusion shape (e.g., a cube, a cylinder, a sphere, or a regular or irregular shape). By setting the occlusion type, different shadow and / or reflective lighting effects can be displayed on a solid color image.
[0061] It should be noted that the above light source type setting, light source position setting, background texture setting, background color setting, obstruction position setting, obstruction type setting and other light and shadow processing methods can be used in combination as needed, and this application does not limit this.
[0062] S3012: Obtain an original pixel change trend graph based on the pixel mean values of other areas except the shadow area and the reflective area in the simulation image, and the pixel values of the shadow area and / or the reflective area.
[0063] After obtaining the simulation image, the pixel mean of other areas in the simulation image except the shadow area and the reflective area can be calculated, and then combined with the pixel values of the shadow area and / or reflective area in the simulation image to obtain the original pixel change trend graph indicating the pixel changes between the light and shadow areas in the image and other areas except the light and shadow areas.
[0064] Based on this, in this application, through the optional implementation of sub-steps S3011 to S3022, the original pixel change trend graph can be obtained conveniently, quickly and effectively for subsequent processing.
[0065] In some optional embodiments, sub-step S3012 includes the following sub-steps S3012A and S3012B, specifically:
[0066] S3012A: Calculate and obtain the pixel mean of other areas in the simulation image except the shadow area and the reflective area.
[0067] Any pixel mean calculation method can be used to calculate the pixel mean of other areas in the simulation image except the shadow area and the reflective area. For example, the pixel mean can be calculated by summing the pixel values of each pixel in the other area and then averaging the sum to obtain the pixel mean of the other area.
[0068] S3012B: Use the pixel values of the shadow area and / or the reflective area, subtract the pixel mean, and obtain the corresponding original pixel change trend graph.
[0069] In the present application, after obtaining the pixel mean, for each pixel in the shadow area and / or the reflective area, the pixel mean is subtracted from the pixel value of each pixel to obtain the corresponding pixel change value.
[0070] It should be understood that when the simulation image only includes shadow areas or reflective areas, it is only necessary to subtract the pixel mean from the pixel values of the shadow areas or reflective areas; if the simulation image includes both shadow areas and reflective areas, the pixel mean is subtracted from the pixel values of both the shadow areas and the reflective areas to obtain a better original pixel change trend graph.
[0071] Based on this, in this application, through the optional implementation method of sub-steps S3012A to S3012B, the original pixel change trend diagram can be obtained conveniently, quickly and effectively for subsequent processing.
[0072] In some optional embodiments, the "other areas except the shadow area and the reflective area" in step S3012 can be determined in the following manner: other areas except the shadow area and the reflective area in the simulation image are determined by at least one of the following processes, namely: image binarization processing of the simulation image, field analysis processing of the simulation image, labeling processing of the shadow area of the simulation image, and labeling processing of the reflective area of the simulation image.
[0073] In the present application, by performing image binarization on the simulation image, the other areas in the simulation image except the shadow area and the reflective area are determined. An optional process can be: by performing image binarization on the simulation image, the simulation image after binarization is divided into two parts with pixel values of 1 and 0 (herein, the binarized pixel values), one of which is the part of the shadow area and the reflective area, and the other is the part of the other areas except the shadow area and the reflective area. Then, from the simulation image, the image area corresponding to the other areas of the simulation image after binarization is determined, and the other areas except the shadow area and the reflective area can be determined in the simulation image. For example, the binarization process can be implemented by setting a preset pixel value. For example, according to a preset pixel value threshold, the pixel value of the image area with a pixel value greater than or equal to the pixel value threshold is set to 1, and the pixel value of the image area with a pixel value less than the pixel value threshold is set to 0. The shadow area and the reflective area can be set to 1, and the other parts can be set to 0; or the shadow area and the reflective area can be set to 0, and the other parts can be set to 1, and the setting can be as needed. Taking a pure white image as an example (referring to the above description, an image with a pixel value greater than or equal to 250 can be considered a pure white image), then the areas other than the shadow area and the reflective area in the simulation image are generally pure white. Therefore, after the simulation image is binarized, the pure white areas with pixel values greater than 250 can be set to 1, and the areas with pixel values less than 250 can be set to 0. Then, in the simulation image after binarization, the areas with pixel values of 0 are the areas other than the shadow area and the reflective area. For another example, in an optional embodiment, the simulation image may include an object, the shadow area may be generated by the object, and the reflective area may be on the object. Then, the object and its corresponding shadow area in the simulation image may be detected by target detection, and then the shadow area and the reflective area may be identified separately from the detection results. Then, the pixel values of the shadow area and the reflective area are set to 1, and then the area other than the shadow area and the reflective area in the detection results is set to 0 (or may also be set to 1, 0 is taken as an example here), and then the pixel values of the part other than the object and its corresponding shadow area in the simulation image are all set to 0, and then a simulation image after binarization processing is obtained. In the simulation image after binarization processing, the area with a pixel value of 0 is the other area except the shadow area and the reflective area. After obtaining the simulation image after binarization processing, the image area corresponding to the other area of the simulation image after binarization processing is determined from the simulation image, and the other area except the shadow area and the reflective area can be determined in the simulation image. It should be understood that the above description is only an example and is not intended to limit the present application.
[0074] The neighborhood analysis processing of the simulation image can be implemented by any neighborhood analysis algorithm of the image, which is not limited in this application.
[0075] For the labeling of the shadow area of the simulation image, it can be done manually, for example, by labeling each pixel point of the shadow area in the simulation image or labeling each pixel point of the boundary of the shadow area according to manual identification, or it can be implemented through a preset labeling model. For the labeling of the reflective area of the simulation image, it can also be done manually, that is, by labeling each pixel point of the reflective area in the simulation image or labeling each pixel point of the boundary of the reflective area according to manual identification, or it can be implemented through a preset labeling model. After the shadow area and the reflective area in the simulation image are labeled, the area other than the shadow area and the reflective area in the simulation image is the area other than the shadow area and the reflective area.
[0076] According to at least one processing method of the above embodiments, other areas in the simulation image except the shadow area and the reflective area can be determined, and the pixel mean of the other areas can be calculated. The optional method of calculating the pixel mean has been introduced in the previous article and will not be repeated here.
[0077] Based on this, in this application, the above multiple methods are used to determine other areas in the simulation image except the shadow area and the reflective area, and the other areas can be determined conveniently and effectively to facilitate the calculation of the pixel mean of other areas, and then facilitate the calculation of the original pixel change trend map for processing.
[0078] Alternatively, other processing methods may be selected in the present application to determine other areas in the simulation image except the shadow area and the reflective area, which is not limited here.
[0079] It should be understood that in addition to the above method of generating the original pixel change trend graph, the embodiments of the present application can also be implemented in other reasonable ways, and the relevant methods should be considered within the scope of protection of the embodiments of the present application.
[0080] S302: Based on a preset range of change factors, randomly select a change factor, and adjust the change trend of the original pixel change trend graph based on the selected change factor to generate a new pixel change trend graph.
[0081] After obtaining the original pixel change trend graph, a change factor can be randomly selected from a preset change factor range (for example, the change factor can be recorded as alpha) to adjust the change trend of the original change trend graph according to the selected change factor alpha, thereby generating a new pixel change trend graph.
[0082] Through the above method, more new pixel change trend graphs can be generated based on the original pixel change trend graph, thereby improving the richness and diversity of the pixel change trend graphs in the pixel change trend graph set; and, in such a way, new pixel change trend graphs and pixel change trend sets can be generated in real time when the machine learning model is trained online, which can effectively adapt to the needs of online model training and effectively reduce the storage burden and storage consumption in the online training scenario.
[0083] Optionally, in step S302, the change trend of the original pixel change trend graph is adjusted based on the selected change factor alpha to generate a new pixel change trend graph. The change trend adjustment can be achieved by multiplying the pixel value of each pixel point in the original pixel change trend graph by the selected change factor alpha. For example, if the pixel values of the pixel points of the original pixel trend graph are divided into three areas, the pixel values of the pixel points in the first area are all A, the pixel values of the pixel points in the second area are all 1.2A, and the pixel values of the pixel points in the third area are all 1.5A. If the randomly selected change factor is b, the pixel values of the three areas of the pixel change trend graph after the change trend adjustment can be A×alpha, 1.2A×alpha, and 1.5A×alpha respectively. If the calculated pixel value of some pixel points exceeds 255, the pixel value of the pixel point is directly set to 255. It should be understood that this example is only a simple example and does not serve as any limitation to the present application.
[0084] Optionally, in step S302, a change factor may be randomly selected from the range of change factors to adjust the change trend of the multiple original pixel change trend graphs obtained to obtain multiple new pixel change trend graphs. Alternatively, multiple change factors may be randomly selected from the range of change factors to adjust the change trend of the one original pixel change trend graph obtained to obtain multiple new pixel change trend graphs. This is not specifically limited in the present application.
[0085] The range of the variation factor can be preset as needed and is not specifically limited in this application. For example, the variation factor range can be set to 0.5 to 1.5, and any value thereof can be randomly selected when randomly selecting the variation factor, for example, 0.5, 0.6, 0.7, 0.8, 0.9, 1.0, 1.1, 1.2, 1.3, 1.4, 1.5, etc. Preferably, the variation factor range can be set to 0.7 to 1.3.
[0086] In some optional embodiments, the training sample generation method in the present application further includes: dynamically adjusting the range of the variation factor according to a preset adjustment rule.
[0087] Based on this, in the embodiment of the present application, by dynamically adjusting the range of the change factor, it is possible to more flexibly and randomly select the change factor to generate a new pixel change trend graph, thereby improving the richness and diversity of the generated new pixel change trend graph, and can effectively adapt to the needs of online training of the machine learning model, effectively reducing the storage burden and storage consumption in the online training scenario.
[0088] The preset adjustment rules can be set as needed. For example, in some embodiments, the preset adjustment rules may be to dynamically adjust the range of the change factor after the user issues an adjustment instruction. For another example, in other embodiments, the preset adjustment rules may be to dynamically adjust the range of the change factor after every predetermined time. For another example, in some other embodiments, the preset adjustment rules may be to dynamically adjust the range of the change factor after each predetermined number of new pixel change trend graphs are generated. For another example, in some other embodiments, the range of the change factor may be dynamically adjusted according to the image effect of the new pixel change trend graph generated. Or it may be other adjustment rules, which are not particularly limited in this application.
[0089] Taking the above-mentioned "dynamically adjusting the range of the variation factor after every predetermined period" as an example, for example, the initial variation factor range may be 0.5 to 1.5. Then, after a time period of 10 epochs (training rounds), it can be dynamically adjusted to 0.7 to 1.3. After another time period of 10 epochs (training rounds), it can be adjusted back to 0.5 to 1.5, and this cycle will continue.
[0090] S303: Generate a pixel change trend graph set based on the new pixel change trend graph and the original pixel change trend graph.
[0091] Optionally, after obtaining several new pixel change trend graphs and several original pixel change trend graphs, these new pixel change trend graphs and original pixel change trend graphs can be combined into a set to generate a pixel change trend graph set, so as to collect pixel change trend graphs from the pixel change trend graph set when needed. It can be understood that the pixel change trend graph collected at this time may be either a new pixel change trend graph or an original pixel change trend graph.
[0092] Based on this, the present application includes the optional implementation method of the above-mentioned steps S301 to S303. On the one hand, it can effectively generate a set of pixel change trend graphs and effectively improve the richness and diversity of the pixel change trend graphs in the set of pixel change trend graphs; on the other hand, it can be effectively used to generate a pixel change trend set in real time when the machine learning model is trained online, which can adapt to the needs of online model training and effectively reduce the storage burden and storage consumption in the online training scenario.
[0093] Through the description of the above optional embodiments, an original sample set including multiple original scanned sample images and a pixel change trend graph set including multiple pixel change trend graphs (which may be the aforementioned original pixel change trend graphs and new pixel change trend graphs) can be obtained.
[0094] In some optional embodiments, in step S202, the original scanned sample image is collected from the original sample set, which may be a random collection of the original scanned sample image from the original sample set; the pixel change trend graph is collected from the pixel change trend graph set, which may be a random collection of the pixel change trend graph from the pixel change trend graph set.
[0095] Based on this, in this application, by randomly collecting original scanned sample images from the original sample set and randomly collecting pixel change trend graphs from the pixel change trend graph set, the richness and diversity of subsequently generated training sample pairs can be effectively improved to meet the needs of online training of machine learning models.
[0096] S204: Randomly adjust the resolution of the original scanned sample image to obtain an adjusted sample image.
[0097] In this application, after the original scanned sample images are collected from the original sample set, the original scanned sample images can be randomly adjusted in resolution, and the original scanned sample images can be randomly adjusted to images of other resolutions to achieve scaling of different scales and form multi-scale images to obtain adjusted sample images. This can effectively improve the richness and diversity of the subsequent generated training sample pairs, meeting the needs of online training of machine learning models. In addition, by performing random resolution adjustment processing on the original scanned sample images, the sample images can have richer and more diverse resolutions and a wider sample coverage. After learning, the machine learning model can have better compatibility and processing capabilities for images of various resolutions and better restoration effects after subsequent training is completed.
[0098] It is understood that random resolution adjustment, i.e., if the resolution of the original scanned sample image is a×b, can be adjusted to c×d through random resolution adjustment, where at least one of a and b is different from both c and d. Random resolution adjustment can be performed in any manner. For example, a and b can be randomly added or subtracted. For example, if the original scanned sample image has a=300 and b=500, i.e., a resolution of 300×500, then a can be randomly added by 300 and b by 200, resulting in a random resolution adjustment of c=600 and d=700, i.e., 600×700. Alternatively, a and b can be randomly multiplied by non-zero multiples. For example, if the original scanned sample image has a=300 and b=500, i.e., a resolution of 300×500, then a can be randomly multiplied by 2 and b by 2, resulting in a random resolution adjustment of c=600 and d=1000, i.e., 600×1000. Alternatively, a and b can be directly set to fixed values c and d. For example, if the original scanned sample image has a=300 and b=500, i.e., a resolution of 300×500, a can be randomly set to 2000 and b to 2000, i.e., the random resolution is adjusted to c=2000 and d=2000, i.e., 2000×2000. Alternatively, a can be randomly set to 1000 and b to 1000, i.e., the random resolution is adjusted to c=1000 and d=1000, i.e., 1000×1000. The above examples are merely for ease of understanding and do not constitute any limitation on this application. Alternatively, the original scanned sample image can be randomly adjusted in other ways, which is not specifically limited in this application.
[0099] In some optional embodiments, “randomly adjusting the resolution of the original scanned sample images” in step S204 may include: randomly adjusting the resolution of the original scanned image samples randomly collected from the original sample set according to a preset random adjustment ratio.
[0100] In this application, a random adjustment ratio is used to perform random resolution adjustment on original scanned image samples randomly collected from an original sample set. This can be achieved by multiplying the resolution of the original scanned image sample by the random adjustment ratio to achieve random resolution adjustment and obtain the resolution of the adjusted sample image. That is, if the resolution of the original scanned sample image is a×b and the random adjustment ratio is n, then the random resolution adjustment can be used to adjust it to c×d, where c=a×n and d=b×n. This means that the resolution of the adjusted sample image is (a×n)×(b×n).
[0101] It should be noted that, in such an optional embodiment, multiple different random adjustment ratios may be preset, and then a random adjustment ratio may be randomly selected from the multiple different random adjustment ratios to perform random resolution adjustment on the original scanned image samples randomly collected from the original sample set.
[0102] For example, taking the original scanned sample image with a=300 and b=500, i.e., a resolution of 300×500, as an example, if the preset random adjustment ratios include 9, namely n1, n2, ..., n9, and if n2 is randomly selected and the original scanned image sample randomly collected from the original sample set is subjected to random resolution adjustment, then the resolution of the adjusted image obtained after random resolution adjustment can be (a×n2)×(b×n2)=300n2×500n2. Other similar situations can be deduced by analogy and will not be elaborated here.
[0103] Based on this, in this application, the original scanned image samples randomly collected from the original sample set are randomly adjusted in resolution through a preset random adjustment ratio, which can effectively randomly adjust the original scanned sample images to images of other resolutions, realize scaling of different scales, and form multi-scale images to obtain adjusted sample images, which can effectively improve the richness and diversity of subsequent generated training sample pairs, and meet the needs of online training of machine learning models; and, in this way, it is more suitable to make the sample images have richer and more diverse resolutions and wider sample coverage. Through learning, the machine learning model can have better compatible processing capabilities and better restoration effects for images of various resolutions after subsequent training is completed.
[0104] S206: Based on the pixel change trend graph, perform light and shadow processing on the adjusted sample image to obtain a corresponding simulated scan sample image having a shadow area and / or a reflective area.
[0105] In the present application, after obtaining the adjusted sample image, the adjusted sample image can be subjected to light and shadow processing. The adjusted sample image is subjected to light and shadow processing by using the collected pixel change trend graph to add shadows and / or reflections to some areas of the adjusted sample image, thereby generating a simulated scanned sample image with shadow areas and / or reflection areas. Thus, the machine learning model learns the training sample pairs, so that the trained model can not only effectively remove shadows and / or reflections in the image, but also avoid improper removal of other original information in the image, such as background color, texture, style, etc., thereby achieving the effect of effectively restoring the original appearance of the image.
[0106] It is understandable that, in theory, multiple pixel change trend graphs can be selected and randomly combined with multiple adjusted sample images for light and shadow processing as needed to obtain a rich number of different simulated scanning sample images, which significantly improves the efficiency of generating simulated scanning sample images and generating training sample pairs, and can also further expand the diversity of subsequent generated training sample pairs, thereby helping to improve the generalization of the machine learning model when the training sample pairs are subsequently used to train the machine learning model used to process the scanned images.
[0107] For example, an adjusted sample image can be subjected to light and shadow processing through a pixel change trend graph, an adjusted sample image can be subjected to light and shadow processing through multiple pixel change trend graphs, an adjusted sample image can be subjected to light and shadow processing by changing the direction of the pixel change trend graph (for example, the pixel change trend graph can be subjected to light and shadow processing upright, reversed, or diagonally, etc.), an adjusted sample image can be subjected to light and shadow processing after changing the size of the pixel change trend graph, and so on. Of course, these are all examples and are not intended to limit the present application.
[0108] In some optional embodiments, step S206 includes: superimposing the pixel change trend graph onto the adjusted sample image to add a shadow to an area in the adjusted sample image corresponding to the shadow area of the pixel change trend graph, and / or adding a reflection to an area in the adjusted sample image corresponding to the reflective area of the pixel change trend graph.
[0109] Based on this, by superimposing the pixel change trend graph onto the adjusted sample image, a shadow is added to the area in the adjusted sample image corresponding to the shadow area of the pixel change trend graph, and / or a reflection is added to the area in the adjusted sample image corresponding to the reflective area of the pixel change trend graph, so as to conveniently and effectively obtain the corresponding simulated scanned sample image with a shadow area and / or a reflective area. Moreover, since the pixel change trend graph is directly superimposed on the adjusted sample image in this application, it is easier to avoid the problem of pixel misalignment, and there will be no problem of changing the background color of the adjusted sample image.
[0110] The pixel change trend graph can be superimposed on the adjusted sample image by any suitable image superposition algorithm, which is not limited in this application.
[0111] In other optional embodiments, step S206 includes: subtracting the pixel change trend graph from the adjusted sample image to add reflections in an area in the adjusted sample image corresponding to the shadow area of the pixel change trend graph, and / or adding shadows in an area in the adjusted sample image corresponding to the reflective area of the pixel change trend graph.
[0112] Based on this, by subtracting the pixel change trend graph from the adjusted sample image, reflection is added to the area in the adjusted sample image corresponding to the shadow area of the pixel change trend graph, and / or shadow is added to the area in the adjusted sample image corresponding to the reflective area of the pixel change trend graph, so as to conveniently and effectively obtain the corresponding simulated scanned sample image with shadow area and / or reflective area. Moreover, since the pixel change trend graph is directly superimposed on the adjusted sample image in this application, it is easier to avoid the problem of pixel misalignment, and there will be no problem of changing the background color of the adjusted sample image.
[0113] The pixel change trend graph can be subtracted from the adjusted sample image by any suitable image subtraction algorithm, which is not limited in this application.
[0114] It should be understood that in the present application, other processing methods may be selected to perform light and shadow processing on the adjusted sample image to obtain a simulated scanned sample image, and this is not limited here.
[0115] S208: Generate training sample pairs for online training of a machine learning model based on the adjusted sample images and the simulated scanned sample images, wherein the machine learning model is used to process the scanned images.
[0116] Optionally, training sample pairs are generated based on the adjusted sample image obtained in step S204 and the simulated scanned sample image with shadow areas and / or reflective areas obtained in step S206 to facilitate training of a machine learning model for processing scanned images.
[0117] For example, the adjusted sample image and the simulated scanned sample image can be directly combined to form a training sample pair and then output. Alternatively, in other optional embodiments, the adjusted sample image can be combined with the simulated scanned sample image and subjected to processing operations including but not limited to appropriate cropping of the same area, and the same enlargement or reduction, to generate a training sample pair.
[0118] Based on the optional implementation in the above steps S202 to S208, for the training of the machine learning model for processing scanned images, when generating training samples for it, the original scanned sample image will be subjected to light and shadow processing based on the pixel change trend diagram, and a simulated scanned sample image will be generated based on the light and shadow processing. In addition, the original scanned sample image will be subjected to random resolution adjustment to generate an adjusted sample image. Based on the simulated scanned sample image and the adjusted sample image, a training sample pair for training the machine learning model is generated. Among them, through light and shadow processing, shadow areas and / or reflective areas can be added to the original scanned sample image, so that the generated simulated scanned sample image has shadow areas and / or reflective areas. Therefore, the machine learning model, through learning the training sample pair, enables the trained model to effectively remove shadows and / or reflections in the image, and avoid improper removal of other original information in the image, such as background color, texture, style, etc., thereby achieving the effect of effectively restoring the original appearance of the image. By randomly adjusting the resolution of the original scanned sample images, the sample images can have richer and more diverse resolutions and a wider sample coverage. After learning, the machine learning model can have better compatibility processing capabilities and better restoration effects for images of various resolutions after subsequent training is completed.
[0119] In some optional embodiments, the training sample generation method further includes: online training of a machine learning model for processing scanned images based on the training sample pairs.
[0120] The training sample pairs obtained through the technical solution of the present application are used to perform online training on the machine learning model used to process the scanned images. On the one hand, the machine learning model learns the training sample pairs so that the trained model can not only effectively remove shadows and / or reflections in the image, but also avoid improper removal of other original information in the image, such as background color, texture, style, etc., thereby achieving the effect of effectively restoring the original appearance of the image; on the other hand, the machine learning model can have better compatible processing capabilities and better restoration effects for images of various resolutions after subsequent training is completed through learning; on the other hand, in the present application, when the machine learning model is trained online, the training sample pairs can be generated online in real time, which can effectively meet the needs of online model training and effectively reduce the storage burden and storage consumption in the online training scenario.
[0121] Optionally, in each optional embodiment of the above-mentioned training sample generation scheme of the present application, at least some steps can be integrated into the dataloader (data loading module in the training framework) and performed in real time during the online training of the machine learning model, for example, it may include: step S202 "collecting the original scanned sample image from the original sample set, and collecting the pixel change trend graph from the pixel change trend graph set", step S204 "randomly adjusting the resolution of the original scanned sample image to obtain the adjusted sample image", step S206 "based on the pixel change trend graph, performing light and shadow processing on the adjusted sample image to obtain the corresponding simulated scanned sample image with shadow area and / or reflective area", step S208 "generating a training sample pair for online training of the machine learning model based on the adjusted sample image and the simulated scanned sample image", and may also include the steps of generating the original pixel change trend graph and the generation of the pixel change trend graph set. In addition, from the above optional embodiments, the training sample generation scheme in this application can effectively improve the generation efficiency of training sample pairs, reduce the complexity of training data preparation, and ensure sample diversity and richness. It can also effectively improve the stability of online model training, and further improve the use effect of the machine learning model trained by the training samples. Moreover, through the above optional embodiments, the training sample generation scheme in this application can effectively reduce the difficulty of producing shadow removal and reflection removal training data, breaking through the difficulty of removing shadows and reflections in the machine learning model used to process scanned images by scanning software. The generated samples can fill the gaps in the data set in this field, and are simple and easy to use, which is conducive to helping the machine learning model to iterate quickly.
[0122] FIG5 is a schematic diagram of a scenario example of the training sample generation scheme of the embodiment of the present application. Combined with FIG5, the training sample generation scheme in the embodiment of the present application can be understood as a whole. It should be noted that in FIG5, for the sake of easy distinction, the pure color image, the original pixel change trend graph, the new pixel change trend graph and the original scanned sample image are all illustrated with "double dotted lines", while the adjusted sample image, the simulated scanned sample image and the training sample pair are all illustrated with "solid lines". These illustrations do not represent other meanings. As shown in the upper content of Figure 5, the pixel change trend graph set includes the original pixel change trend graph and the new pixel change trend graph generated after the change trend of the original pixel change trend graph is adjusted. The original pixel change trend graph in Figure 5 is obtained by processing the pure color image. Specifically, a simulation graph can be generated for the pure color image, and then the original pixel change trend graph is obtained according to the simulation graph (the pure color image, the original pixel change trend graph, and the new pixel change trend graph are composed of pixels. In Figure 5, pixels are represented by rectangular grids, and pixel values are marked in the rectangular grids. For example, the pixel value of each pixel in the pure color image in Figure 5 is 255, which is a pure white image; the pixel values of the pixels in the original pixel change trend graph change sequentially; in the example of Figure 5, the change factor alpha = 2 is randomly selected for adjusting the change trend of the original pixel change trend graph, and the pixel values of the pixels in the obtained new pixel change trend graph change sequentially; it should be understood that Figure 5 here is only a schematic diagram, and it is more complicated in practice, and it does not serve as any limitation to the present application). 5 , step S202 in the present application: “collecting the original scanned sample image from the original sample set, and collecting the pixel change trend graph from the pixel change trend graph set”, can both be randomly collected. The pixel change trend graph randomly collected in step S202 is the new pixel change trend graph of the content above FIG5 ; then step S204: “randomly adjust the resolution of the original scanned sample image to obtain an adjusted sample image”, as shown in FIG5 , the resolution of the original scanned sample image is a×b, and the resolution of the adjusted sample image obtained after step S204 is c×d; then step S206: “perform light and shadow processing on the adjusted sample image based on the pixel change trend graph to obtain the corresponding sample image with shadow”. The adjusted sample image is obtained by randomly adjusting the resolution of the original scanned sample image. For example, the original scanned sample image in the example of FIG5 is an image without shadows and reflections. The original scanned sample image is subjected to light and shadow processing through the pixel change trend graph to add shadows and / or reflections to part of the image, thereby generating a simulated scanned sample image with shadow areas and / or reflection areas. The shadow areas and / or reflection areas in the simulated scanned sample image are represented by “dash-dotted lines” in FIG5 . After step S208, a training sample pair for online training of a machine learning model for processing scanned images can be generated based on the adjusted sample image and the simulated scanned sample image.It should be understood that the scenario example shown in Figure 5 is only used to facilitate understanding of the embodiments of the present application and does not constitute any limitation to the present application.
[0123] It is understandable that the above description of the training sample generation method is merely an exemplary description of the present application and does not constitute any limitation to the present application.
[0124] FIG6 is a flowchart of the steps of a model training method according to an embodiment of the present application. According to the first aspect of the present application, a training sample generation method is provided for online training of a machine learning model for processing scanned images. As shown in FIG6 , the method includes steps S602 and S604, specifically:
[0125] S602: Acquire a training sample pair, wherein the training sample pair includes: an image obtained by randomly adjusting the resolution based on the original scanned sample image, and an image having shadow areas and / or reflective areas obtained by performing light and shadow processing on the adjusted image based on a pixel change trend graph.
[0126] Optionally, the training sample pairs can be generated by any of the training sample generation methods provided in the first aspect. The relevant content of the training sample pairs in step S602 can be understood with reference to the embodiment of the training sample generation method in the first aspect above, and will not be repeated here.
[0127] By obtaining the training sample pairs generated by the training sample generation method provided in the first aspect, the machine learning model is trained online. On the one hand, the machine learning model learns the training sample pairs so that the trained model can not only effectively remove shadows and / or reflections in the image, but also avoid improper removal of other original information in the image, such as background color, texture, style, etc., thereby effectively restoring the original appearance of the image. On the other hand, the machine learning model can have better compatibility processing capabilities and better restoration effects for images of various resolutions after subsequent training is completed. On the other hand, in this application, when the machine learning model is trained online, the training sample pairs can be generated online in real time, which can effectively meet the needs of online model training and effectively reduce the storage burden and storage consumption in the online training scenario. Therefore, the effect of online model training can be more effectively improved.
[0128] S604: Based on the training sample pairs and the preset loss function, the machine learning model is trained online.
[0129] In this application, the loss function preset by the model learning model can be used through training samples to perform online training on the machine learning model until the training end conditions of the model online training are met, thereby ending the online training.
[0130] For example, the training end condition can be that when the training rounds reach a predetermined number of rounds, the training end condition is considered to be met, thereby terminating the model training process. Alternatively, the training end condition can also be that when the loss value calculated by a preset loss function is within a preset loss value range, the training end condition is considered to be met, thereby terminating the model training process.
[0131] Any suitable loss function may be used in this application to implement online model training, and is not specifically limited herein. For example, in some optional embodiments, the preset loss function includes at least one of the following: a loss function for evaluating the degree of pixel restoration, or a loss function for evaluating the degree of image style restoration.
[0132] Based on this, in this application, the machine learning model is trained online through training sample pairs and loss functions for the degree of pixel restoration, so that when the machine learning model after online training processes the scanned image, the pixel restoration is better and more stable restoration is achieved, thereby improving the effect of the model's online training; in this application, the machine learning model is trained online through training sample pairs and loss functions for evaluating the degree of image style restoration, which can enable the machine learning model to better learn the image style, so that when the machine learning model after online training processes the scanned image, it can better restore the image style and more realistically restore the original appearance of the scanned image, thereby improving the effect of the model's online training.
[0133] For example, the loss function used to evaluate the degree of pixel restoration may be an L1 loss function, or other feasible loss functions may be used.
[0134] For example, the loss function used to evaluate the degree of image style restoration can be a perceptual loss function. Alternatively, other feasible loss functions can be used.
[0135] In some optional embodiments, the preset loss function also includes a loss function for evaluating the overall image restoration degree. Based on this, in this application, the machine learning model is trained online by using training sample pairs and the loss function for evaluating the overall image restoration degree. This allows the trained machine learning model to achieve better overall image restoration and more stable restoration when processing scanned images, thereby improving the effectiveness of the model's online training.
[0136] For example, the loss function used to evaluate the overall restoration degree of the image may be a Peak Signal-to-Noise Ratio (PSNR) loss function. Alternatively, other feasible loss functions may be used.
[0137] Optionally, in order to further improve the model training effect in this application, the loss function used to evaluate the overall image restoration degree can be combined with at least one of the loss function used to evaluate the pixel restoration degree and the loss function used to evaluate the image style restoration degree. It is better to use all three in combination. When the model is trained online, when the loss function used to evaluate the overall image restoration degree (such as the PSNR loss function) converges stably, the convergence of the loss function used to evaluate the pixel restoration degree (such as the L1 loss function) and the convergence of the loss function used to evaluate the image style restoration degree (such as the Perceptual Loss loss function) are superimposed to achieve better training effect.
[0138] Based on the model training method of the above steps S602 to S604, on the one hand, the machine learning model learns the training sample pairs so that the trained model can not only effectively remove the shadows and / or reflections in the image, but also avoid improper removal of other original information in the image, such as background color, texture, style, etc., thereby achieving the effect of effectively restoring the original appearance of the image. The learning model is trained online; on the other hand, the machine learning model can have better compatible processing capabilities and better restoration effects for images of various resolutions after subsequent training is completed. Therefore, the model training method in this application can effectively improve the effect of model online training.
[0139] FIG7 is a schematic diagram of a scenario example of the training sample generation scheme of the embodiment of the present application. Combined with FIG7, the training sample generation scheme in the embodiment of the present application is understood as a whole. Referring to FIG7, the training sample generation method of the first aspect is used to obtain a training sample pair and a preset loss function, and the model training method (steps S602 to S604) provided by the second aspect is used to train the machine learning model for processing the scanned image. A trained machine learning model can be obtained, and the scanned image to be processed (as shown in FIG7, the shadow area and the reflective area in the scanned image are also schematically shown by "dotted lines") is input into the trained machine learning model. The trained machine learning model can output a processed output image, and the shadow area and the reflective area in the output image are removed and the useful content and details in the original scanned image are retained. The model has a good restoration effect. It should be understood that the scenario example shown in FIG7 is only used to facilitate the understanding of the embodiment of the present application and does not serve as any limitation to the present application.
[0140] It should be understood that the above description of the model training method is merely an exemplary description of the present application and does not constitute any limitation to the present application.
[0141] According to a third aspect of an embodiment of the present application, an electronic device is provided, comprising: a processor, a memory, a communication interface, and a communication bus, wherein the processor, the memory, and the communication interface communicate with each other via the communication bus; the memory is used to store at least one executable instruction, and the executable instruction causes the processor to perform operations corresponding to the training sample generation method described in the first aspect or the model training method described in the second aspect. Referring to Figure 8, a schematic diagram of the structure of an electronic device according to an embodiment of the present application is shown, and the embodiment of the present application does not limit the specific implementation of the electronic device.
[0142] As shown in FIG8 , the electronic device 800 may include a processor 802 , a communications interface 804 , a memory 806 , and a communication bus 808 .
[0143] in:
[0144] The processor 802 , the communication interface 804 , and the memory 806 communicate with each other via a communication bus 808 .
[0145] The communication interface 804 is used to communicate with other electronic devices or servers.
[0146] The processor 802 is used to execute the program 810, and specifically can execute the relevant steps in the above-mentioned training sample generation method or model training method embodiment.
[0147] Specifically, the program 810 may include program codes, which include computer operation instructions.
[0148] Processor 802 may be a CPU, a Graphics Processing Unit (GPU), an Application Specific Integrated Circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present application. The one or more processors included in the smart device may be processors of the same type, such as one or more CPUs, or may be processors of different types, such as one or more CPUs and one or more ASICs.
[0149] The memory 806 is used to store the program 810. The memory 806 may include a high-speed RAM memory, and may also include a non-volatile memory (non-volatile memory), such as at least one disk memory.
[0150] The program 810 may include multiple computer instructions. Specifically, the program 810 may enable the processor 802 to execute operations corresponding to the training sample generation method or model training method described in any of the aforementioned method embodiments through multiple computer instructions.
[0151] The specific implementation of each step in program 810 can refer to the corresponding description of the corresponding steps and units in the above-mentioned method embodiment, and has corresponding beneficial effects, which will not be repeated here. Those skilled in the art will clearly understand that for the convenience and brevity of description, the specific working process of the above-mentioned devices and modules can refer to the corresponding process description in the above-mentioned method embodiment, and will not be repeated here.
[0152] According to a fourth aspect of the embodiments of the present application, the embodiments of the present application further provide a computer storage medium having a computer program stored thereon, which, when executed by a processor, implements the training sample generation method or model training method described in any of the aforementioned method embodiments. The computer storage medium includes, but is not limited to, a compact disc read-only memory (CD-ROM), random access memory (RAM), a floppy disk, a hard disk, or a magneto-optical disk.
[0153] An embodiment of the present application also provides a computer program product, including computer instructions, which instruct a computing device to perform operations corresponding to any training sample generation method or model training method in the above-mentioned multiple method embodiments.
[0154] The electronic device 800 / computer storage medium / computer program product embodiment in the embodiment of the present application has been described in detail in the aforementioned training sample generation method embodiment, so its relevant content and beneficial effects can be understood with reference to the above-mentioned method embodiment and will not be repeated here.
[0155] In addition, it should be noted that the user-related information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to sample data used to train the model, data used for analysis, stored data, displayed data, etc.) involved in the embodiments of this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data need to comply with the relevant laws, regulations and standards of relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse.
[0156] It should be pointed out that, according to the needs of implementation, the various components / steps described in the embodiments of the present application can be split into more components / steps, or two or more components / steps or partial operations of components / steps can be combined into new components / steps to achieve the purpose of the embodiments of the present application.
[0157] The above-mentioned method according to the embodiment of the present application can be implemented in hardware, firmware, or as software or computer code that can be stored in a recording medium (such as a CD-ROM, RAM, floppy disk, hard disk or magneto-optical disk), or as computer code that is originally stored in a remote recording medium or a non-temporary machine-readable medium downloaded via a network and will be stored in a local recording medium, so that the method described herein can be stored in such software processing on a recording medium using a general-purpose computer, a dedicated processor or programmable or dedicated hardware (such as an application-specific integrated circuit (ASIC) or a field programmable gate array (FPGA)). It can be understood that a computer, a processor, a microprocessor controller or programmable hardware includes a storage component (e.g., random access memory (RAM), read-only memory (ROM), flash memory, etc.) that can store or receive software or computer code. When the software or computer code is accessed and executed by a computer, a processor or hardware, the method described herein is implemented. In addition, when a general-purpose computer accesses the code for implementing the method shown here, the execution of the code converts the general-purpose computer into a dedicated computer for executing the method shown here.
[0158] Those skilled in the art will appreciate that the units and method steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for specific applications, but such implementation should not be considered to be beyond the scope of the embodiments of this application.
[0159] The above implementation methods are only used to illustrate the embodiments of the present application, and are not intended to limit the embodiments of the present application. Ordinary technicians in the relevant technical field can make various changes and modifications without departing from the spirit and scope of the embodiments of the present application. Therefore, all equivalent technical solutions also fall within the scope of the embodiments of the present application, and the scope of patent protection of the embodiments of the present application should be defined by the claims.
Claims
1. A training sample generation method, comprising: Collecting an original scanned sample image from an original sample set, and collecting a pixel change trend graph from a pixel change trend graph set, wherein the pixel change trend graph is used to indicate pixel changes between a light and shadow area in an image and other areas except the light and shadow area, and the light and shadow area includes a shadow area and / or a reflective area in the image; Randomly adjusting the resolution of the original scanned sample image to obtain an adjusted sample image; Based on the pixel change trend graph, performing light and shadow processing on the adjusted sample image to obtain a corresponding simulated scanned sample image having a shadow area and / or a reflective area; Based on the adjusted sample image and the simulated scanned sample image, a training sample pair is generated for online training of a machine learning model, wherein the machine learning model is used to process the scanned image.
2. The method according to claim 1, wherein: The collecting of original scanned sample images from the original sample set includes: randomly collecting original scanned sample images from the original sample set; The collecting of the pixel change trend graph from the pixel change trend graph set includes: randomly collecting the pixel change trend graph from the pixel change trend graph set.
3. The method according to claim 2, wherein: The randomly adjusting the resolution of the original scanned sample image comprises: According to a preset random adjustment ratio, the original scanned image samples randomly collected from the original sample set are randomly adjusted in resolution.
4. The method according to any one of claims 1 to 3, wherein: The pixel change trend graph set is generated in the following manner: Get the original pixel change trend graph; Based on a preset range of variation factors, randomly selecting variation factors, and adjusting the variation trend of the original pixel variation trend graph based on the selected variation factors to generate a new pixel variation trend graph; A pixel change trend graph set is generated based on the new pixel change trend graph and the original pixel change trend graph.
5. The method according to claim 4, wherein: The method further comprises: The range of the variation factor is dynamically adjusted according to a preset adjustment rule.
6. The method according to claim 4 or 5, wherein: The original pixel change trend graph is generated in the following way: Performing light and shadow processing on the pure color image to obtain a simulated image with shadow areas and / or reflective areas; Based on the pixel averages of other regions in the simulation image except the shadow region and the reflective region, and, The pixel values of the shadow area and / or the reflective area are used to obtain an original pixel change trend graph.
7. The method according to claim 6, wherein: The step of performing light and shadow processing on the pure color image to obtain a simulated image having a shadow area and / or a reflective area includes: Light and shadow processing including light source type setting and / or light source position setting is performed on the pure color image to obtain a simulated image with shadow areas and / or reflective areas.
8. The method according to claim 6 or 7, wherein: The light and shadow processing also includes at least one of the following: background texture setting, background color setting, occluder position setting, and occluder type setting.
9. A model training method for online training of a machine learning model for processing scanned images, the method comprising: Acquire a training sample pair, wherein the training sample pair includes: an image after random resolution adjustment based on the original scanned sample image, and an image with a shadow area and / or a reflective area obtained by performing light and shadow processing on the adjusted image based on a pixel change trend graph; Based on the training sample pairs and a preset loss function, the machine learning model is trained online.
10. The method according to claim 9, wherein: The preset loss function includes at least one of the following: a loss function for evaluating the degree of pixel restoration, and a loss function for evaluating the degree of image style restoration.
11. The method according to claim 10, wherein: The preset loss function also includes: a loss function for evaluating the overall restoration degree of the image.
12. The method according to any one of claims 9 to 11, wherein: The training sample pairs are generated by the method according to any one of claims 1-8.
13. An electronic device comprising: A processor, a memory, a communication interface and a communication bus, wherein the processor, the memory and the communication interface communicate with each other via the communication bus; The memory is used to store at least one executable instruction, and the executable instruction enables the processor to perform an operation corresponding to the method according to any one of claims 1 to 12.
14. A computer storage medium having a computer program stored thereon, wherein when the program is executed by a processor, the method according to any one of claims 1 to 12 is implemented.
15. A computer program product, comprising computer instructions, wherein the computer instructions instruct a computing device to execute operations corresponding to the method according to any one of claims 1 to 12.
Citation Information
Patent Citations
Correction method and device for scanned image
CN106791268A
Method for homogenizing background light and shadow in photo and text shooting and scanning system
CN116723412A
Training sample generation method, model training method, electronic equipment and storage medium
CN117745787A
Image processing apparatus, image forming apparatus and image processing method
US20100085611A1