Image enhancement method and device
By separating the foreground and background of the picture, keeping the foreground features unchanged, and randomly replacing the background features, the problem of destroying the main features by existing data augmentation algorithms is solved, and the recognition ability and generalization of the model are improved.
Patent Information
- Application Number
- CN202110630620.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-06-07
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2041-06-07
AI Technical Summary
The existing data augmentation algorithm randomly adds noise to the overall picture, resulting in the damage to the main feature structure and affecting the generalization of the model.
By separating the foreground and background of the picture, keeping the foreground features unchanged, randomly substituting the pixels of the background features to generate enhanced pictures for image classification.
It effectively improves the model's ability to identify and extract main features, and improves the accuracy and generalization of the model.
Smart Images

Figure CN113506207B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image processing, and in particular to an image enhancement method and device. Background Art
[0002] Before the image to be recognized or classified, noise is randomly added to the training image to enrich the sample so that the model can better learn the main features and increase the generalization of the model. In the process of implementing the present invention, the applicant found that there are at least the following problems in the prior art: However, the existing data enhancement algorithm randomly adds noise to the entire image, and the noise will have a destructive effect on the main feature structure. Summary of the invention
[0003] The embodiment of the present invention provides a method and device for image enhancement, which separates the foreground and background of an image, keeps important features unchanged, performs random pixel replacement on pixels of background features to change secondary features, and enhances the model to identify and extract main features.
[0004] To achieve the above object, on the one hand, an embodiment of the present invention provides a method for enhancing an image, comprising:
[0005] Determine the foreground features and background features of the original image, separate the foreground features and background features in the original image, and form corresponding foreground images and background images respectively; wherein the foreground features are features corresponding to the target image of the original image, and the background features are features corresponding to the non-target image of the original image;
[0006] The pixels of the background image are randomly transformed to generate a background transformation image, and a plurality of corresponding background transformation images are generated by randomly transforming the pixels of the background image multiple times;
[0007] Each background transformation image is fused with the foreground image to regenerate a corresponding enhanced image, which is used for image classification.
[0008] On the other hand, an embodiment of the present invention provides a picture enhancement device, including:
[0009] An important feature separation module is used to determine the foreground features and background features of the original image, separate the foreground features and background features in the original image, and form corresponding foreground images and background images respectively; wherein the foreground features are features corresponding to the target image of the original image, and the background features are features corresponding to the non-target image of the original image;
[0010] A background feature transformation module, used to randomly transform the pixels of the background image to generate a background transformation image, and to generate multiple corresponding background transformation images by randomly transforming the pixels of the background image multiple times;
[0011] The fusion module is used to fuse each background transformation image with the foreground image respectively to regenerate a corresponding enhanced image, and the enhanced image is used for image classification.
[0012] The above technical solution has the following beneficial effects: for a picture in a given sample, the foreground picture and the background picture of the picture are obtained through the important feature separation module, the pixels of the background picture are randomly replaced and merged with the foreground picture to obtain an enhanced picture, and then the enhanced picture is used as a new sample of the deep learning model. By keeping the important features unchanged and changing the sample enhancement of the secondary features, the model can be well improved for the recognition and extraction of the main features. When using enhanced pictures for training, the training model has higher accuracy and generalization, which has a very high practical improvement effect. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0014] Figure 1 is a flow chart of a method for picture enhancement according to an embodiment of the present invention;
[0015] Figure 2 is a structural diagram of a picture enhancement device according to an embodiment of the present invention;
[0016] Figure 3 It is a principle diagram of the image enhancement method according to an embodiment of the present invention. DETAILED DESCRIPTION
[0017] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0018] like Figure 1 As shown, in combination with an embodiment of the present invention, a method for enhancing an image is provided, comprising:
[0019] S101: Determine foreground features and background features of the original image, separate the foreground features and background features in the original image, and form corresponding foreground images and background images respectively; wherein the foreground features are features corresponding to the target image of the original image, and the background features are features corresponding to the non-target image of the original image;
[0020] S102: randomly transforming pixels of the background image to generate a background transformation image, and generating a plurality of corresponding background transformation images by randomly transforming pixels of the background image multiple times;
[0021] S103: Merge each background transformation picture with the foreground picture respectively to regenerate a corresponding enhanced picture, which is used for image classification.
[0022] Preferably, step 101 specifically includes:
[0023] S1011: Segmenting the original image using an image segmentation model to obtain foreground features of the original image, saving the foreground features into an image to form a foreground image, wherein the foreground image includes blank areas except for the area where the foreground features are located, and the blank areas of the foreground image overlap with the areas where the background features of the original image are located;
[0024] S1012: setting an empty matrix mask, wherein the empty matrix mask has the same width and height as the foreground image;
[0025] S1013: each point in the empty matrix mask is represented by a coordinate (u, v), and the empty matrix mask is adjusted by assigning a coordinate (u, v) value of a specific point in the empty matrix mask to 1, to obtain a mask adjustment matrix; the specific point is a point corresponding to a blank area of the foreground image;
[0026] S1014: A specific point with a value of 1 in the mask adjustment matrix represents that the pixel at the corresponding position of the original image is a background pixel, and the mask adjustment matrix is used as a mask matrix of the background image;
[0027] S1015: The original image is processed by the background image mask matrix to obtain background features, pixels at corresponding positions in the original image and whose values in the background image mask matrix are 1 are retained, and pixels at corresponding positions in the original image and whose values in the background image mask matrix are 0 are assigned a value of 0, to obtain a background image of the background features of the original image; wherein the positions in the background image where the pixels are assigned a value of 0 are blank areas of the background image.
[0028] Preferably, the step 102 specifically includes:
[0029] The pixels of the background features in the background image are determined by the background image mask matrix, and the background image is transformed by randomly transforming the pixels of the background features in the background image to obtain a background transformed image.
[0030] Preferably, the step 103 specifically includes:
[0031] S1031: aligning the foreground image with the background transformation image so that the foreground features in the foreground image completely match the blank area of the background transformation image;
[0032] S1032: Add and fuse the pixels of each background transformation image with the pixels at the corresponding position of the foreground image to generate an enhanced image; wherein the pixels of the blank area of the foreground image are set to 0.
[0033] Preferably, the image segmentation model refers to a model formed by training the rembg background removal tool using a UNet network.
[0034] like Figure 2 As shown, in combination with an embodiment of the present invention, a picture enhancement device is provided, comprising:
[0035] The important feature separation module 21 is used to determine the foreground features and background features of the original image, separate the foreground features and background features in the original image, and form corresponding foreground images and background images respectively; wherein the foreground features are features corresponding to the target image of the original image, and the background features are features corresponding to the non-target image of the original image;
[0036] A background feature transformation module 22, used for randomly transforming pixels of a background image to generate a background transformation image, and generating a plurality of corresponding background transformation images by randomly transforming pixels of the background image multiple times;
[0037] The fusion module 23 is used to fuse each background transformation picture with the foreground picture to regenerate a corresponding enhanced picture, and the enhanced picture is used for image classification.
[0038] Preferably, the important feature separation module 21 includes:
[0039] The foreground separation submodule 211 is used to segment the original image through the image segmentation model to obtain the foreground features of the original image, and save the foreground features into an image to form a foreground image, wherein the other areas in the foreground image except the area where the foreground features are located are blank areas, and the blank areas of the foreground image overlap with the areas where the background features of the original image are located;
[0040] The empty matrix construction submodule 212 is used to set an empty matrix mask, wherein the empty matrix mask has the same width and height as the foreground image; each point in the empty matrix mask is represented by a coordinate (u, v), and the empty matrix mask is adjusted by assigning a coordinate (u, v) value of a specific point in the empty matrix mask to 1, thereby obtaining a mask adjustment matrix; the specific point refers to a point corresponding to a blank area of the foreground image;
[0041] Background construction submodule 213, the specific point with a value of 1 in the mask adjustment matrix represents that the pixel at the corresponding position of the original image is a background pixel, and the mask adjustment matrix is used as the background image mask matrix; the original image is processed by the background image mask matrix to obtain background features, the pixels at the corresponding positions in the original image and the background image mask matrix with a value of 1 are retained, and the pixels at the corresponding positions in the original image and the background image mask matrix with a value of 0 are assigned a value of 0, so as to obtain a background image of the background features of the original image; wherein the positions in the background image where the pixels are assigned a value of 0 are blank areas of the background image.
[0042] Preferably, the background feature transformation module 22 is specifically used for:
[0043] The pixels of the background features in the background image are determined by the background image mask matrix, and the background image is transformed by randomly transforming the pixels of the background features in the background image to obtain a background transformed image.
[0044] Preferably, the fusion module 23 includes:
[0045] An alignment submodule 231 is used to align the foreground image with the background transformation image so that the foreground features in the foreground image completely match the blank areas of the background transformation image;
[0046] The fusion submodule 232 is used to add and fuse the pixels of each background transformation image with the pixels at the corresponding position of the foreground image to generate an enhanced image; wherein the pixels of the blank area of the foreground image are set to 0.
[0047] Preferably, the image segmentation model refers to a model formed by training the rembg background removal tool using a UNet network.
[0048] The beneficial effects brought by the present invention are:
[0049] For a picture X in a given sample, the foreground and background pictures of picture X are obtained through the important feature separation module, and the pixels of the background picture are randomly replaced and merged with the foreground picture to obtain a new sample (enhanced picture) picture X1. Then X1 is used as a new sample for the deep learning model. By keeping the important features unchanged and changing the sample enhancement of the secondary features, the model can be well improved for the recognition and extraction of the main features. The use of enhanced picture training makes the training model have higher accuracy and generalization. The technology does not seem to be complicated but has a very high practical improvement effect.
[0050] The above technical solution of the embodiment of the present invention is described in detail below in conjunction with specific application examples. For technical details not introduced during the implementation process, please refer to the relevant description in the previous text.
[0051] The present invention is a deep learning training sample image enhancement method, which uses a new image enhancement method to enhance image samples, and uses enhanced image training to make the training model have higher accuracy and generalization. The technology seems not too complicated but has a very high practical improvement effect. It belongs to the field of machine learning algorithms and can be used for marketing image classification and recognition.
[0052] The principle diagram of the present invention is as follows Figure 3 As shown, first of all, the picture is composed of pixels. For any pixel, we can use RGB to represent it. Each value range of R, G, and B is 0-255, that is, (0-255, 0-255, 0-255). Secondly, in image classification, the image foreground is often the target object, which is an important feature of image classification, and the image background is a secondary feature. The present invention extracts the foreground of the picture as the main feature. Since this feature (foreground feature) is the main part of model learning, the pixels of this feature remain unchanged, and the pixel values of the secondary features (image background) are randomly transformed. Finally, the unchanged main features and the randomly transformed secondary features are fused to generate an enhanced picture, thereby achieving the effect of enhancing the picture sample with the main features unchanged and the background constantly changing. The specific operations are as follows:
[0053] 1. Separation of important features
[0054] This is achieved through the important feature separation module. This process separates the foreground and background of the image. The foreground is the important feature of the image, and the background is the secondary feature of the image. For example, Figure 3 As shown in the picture separation module.
[0055] 1. Use the rembg background removal tool, which uses the UNet network (UNet neural network) to train the segmentation model. The segmentation model trained by the tool obtains the foreground of the image and saves it as an image (such as Figure 3 The important feature images in the image are formed into a foreground image.
[0056] 2. By analyzing the foreground image, we know that the blank area (pixel points less than 10) of the foreground image overlaps with the background of the original image. Here, we define an empty matrix mask with the same size as the width and height of the foreground image. In the mask, the value at the position corresponding to the blank area of the foreground image is assigned to 1 to form a mask adjustment matrix; at this time, the value at the (u, v) position in the mask matrix is 1, which means that the pixel at the (u, v) position of the original image is the background pixel, and the mask adjustment matrix is used as the background image mask matrix.
[0057] 3. Process the background image mask matrix and the original image. Note that the position with a value of 1 in the background image mask matrix is (u, v), and the pixels at (u, v) in the original image are retained. The other pixels are assigned a value of 0. This can generate Figure 3 The background image is backgound. The logic code snippet is as follows:
[0058] rows,cols,channels=image.shape#image is the original image
[0059] foreground=rembg(image)#Use rembg tool to extract foreground image
[0060] mask = np.zeros(shape = (rows, cols)) #Record the position of the background image
[0061] background = np.zeros(shape = image.shape) #define background image
[0062] for row in range(rows):
[0063] for col in range(cols):
[0064] if foreground[row,col]<=10:mask[row,col]=1
[0065] background = np.zeros(shape = image.shape) #define background image
[0066] for row in range(rows):
[0067] for col in range(cols):
[0068] if mask[row,col]==1:
[0069] #Extract the background image from the original image
[0070] [background.itemset((row,col,i),image.item(row,col,i))for iin range(3)]
[0071] Second, random transformation of secondary features
[0072] This is achieved through the background feature transformation module. This process combines the mask matrix (background image mask matrix) and backgound (background image) according to a certain logic to generate a randomly transformed background image recorded as
[0073] backgoundRandom (background transformation picture), the result is as follows Figure 3 As shown in the picture of the secondary feature random transformation module, the logical code snippet is as follows:
[0074] rows,cols,channels=image.shape
[0075] randRGB = [random.randint(0,256),random.randint(0,256),random.randint(0,256)] # Randomly generate pixel values for each channel
[0076] backgroundRandom = np.zeros(shape = image.shape) #define an empty background image
[0077] for row in range(rows):
[0078] for col in range(cols):
[0079] if mask[row,col]==1:
[0080] #The random pixel value of each channel is combined with the pixel value of the background image
[0081] randRGBList=[(background.item(row,col,i)+randRGB[i]) / / 2for iin range(3)]
[0082] #Generate a random background image
[0083] [backgroundRandom.itemset((row,col,i),abs(randRGBList[i]))for iinrange(3)]
[0084] 3. Enhanced Image Generation
[0085] After the important feature separation module and the secondary feature random transformation module, a foreground image (containing important features) and a randomly transformed background image (backgoundRandom) are generated. It is easy to see from the analysis of the process of the two modules that the foreground image and the background image backgoundRandom are aligned in size and position, and the object (feature structure) in the foreground image coincides with the black area (or blank area, pixel is 0, and the pixel is 0 when it is displayed as white or black) in backgoundRandom. The new image can be generated by adding and merging the pixels of the two. Among them, the pixels in the blank area of the foreground image are set to 0 during the addition process, and the others remain unchanged.
[0086] The present invention aims to alleviate the problem that important features may be lost in existing data enhancement algorithms by combining the characteristics of pictures (foreground features are more important) by separating the foreground and background of the picture, keeping the foreground unchanged, and randomly transforming the pixels of the background, thereby alleviating the problem of losing important features in data enhancement. The unchanged important features can improve the model's ability to recognize the main features of the picture, and enhance the sample data, thereby improving the model's generalization ability.
[0087] The beneficial effects brought by the present invention are:
[0088] For a picture X in a given sample, the important feature separation module is used to obtain the foreground and background pictures of picture X, and the pixels of the background picture are randomly replaced and merged with the foreground picture to obtain a new sample picture X1. Then X1 is used as a new sample for the deep learning model. This can improve the model's recognition and extraction of main features by keeping the important features unchanged and changing the sample enhancement of the secondary features. This avoids the occurrence of "randomly adding noise to the foreground and background of the picture without distinguishing between them, which results in adding noise to the structure of the main features and destroying the main features".
[0089] It should be understood that the specific order or hierarchy of steps in the disclosed process is an example of an exemplary method. Based on design preferences, it should be understood that the specific order or hierarchy of steps in the process can be rearranged without departing from the scope of protection of the present disclosure. The attached method claims present the elements of the various steps in an exemplary order and are not intended to be limited to the specific order or hierarchy described.
[0090] In the above detailed description, various features are grouped together in a single embodiment to simplify the disclosure. This method of disclosure should not be interpreted as reflecting an intention that the embodiments of the claimed subject matter require more features than are clearly stated in each claim. On the contrary, as reflected in the appended claims, the invention is in a state of having less than all the features of the disclosed individual embodiments. Therefore, the appended claims are hereby expressly incorporated into the detailed description, with each claim standing on its own as a separate preferred embodiment of the invention.
[0091] The disclosed embodiments are described above to enable any person skilled in the art to implement or use the present invention. Various modifications of these embodiments are obvious to those skilled in the art, and the general principles defined herein may also be applied to other embodiments without departing from the spirit and scope of the present disclosure. Therefore, the present disclosure is not limited to the embodiments given herein, but is consistent with the broadest scope of the principles and novel features disclosed in this application.
[0092] The above description includes examples of one or more embodiments. Of course, it is impossible to describe all possible combinations of components or methods for the purpose of describing the above embodiments, but it should be recognized by those skilled in the art that the various embodiments may be further combined and arranged. Therefore, the embodiments described herein are intended to cover all such changes, modifications and variations that fall within the scope of protection of the appended claims. In addition, with respect to the term "comprising" used in the specification or claims, the word is covered in a manner similar to the term "including", just as "including," is explained as a transitional word in the claims. In addition, any term "or" used in the specification of the claims is intended to mean "non-exclusive or".
[0093] Those skilled in the art may also understand that the various illustrative logical blocks, units, and steps listed in the embodiments of the present invention may be implemented by electronic hardware, computer software, or a combination of the two. In order to clearly demonstrate the interchangeability of hardware and software, the various illustrative components, units, and steps described above have generally described their functions. Whether such functions are implemented by hardware or software depends on the specific application and the design requirements of the entire system. Those skilled in the art may use various methods to implement the described functions for each specific application, but such implementation should not be understood as exceeding the scope of protection of the embodiments of the present invention.
[0094] The various illustrative logic blocks or units described in the embodiments of the present invention can be implemented or operated by a general-purpose processor, a digital signal processor, an application-specific integrated circuit (ASIC), a field programmable gate array or other programmable logic device, a discrete gate or transistor logic, a discrete hardware component, or any combination of the above. The general-purpose processor can be a microprocessor, and optionally, the general-purpose processor can also be any conventional processor, controller, microcontroller or state machine. The processor can also be implemented by a combination of computing devices, such as a digital signal processor and a microprocessor, a plurality of microprocessors, one or more microprocessors combined with a digital signal processor core, or any other similar configuration.
[0095] The steps of the method or algorithm described in the embodiments of the present invention can be directly embedded in hardware, a software module executed by a processor, or a combination of the two. The software module can be stored in a RAM memory, a flash memory, a ROM memory, an EPROM memory, an EEPROM memory, a register, a hard disk, a removable disk, a CD-ROM, or other storage media of any form in the art. Exemplarily, the storage medium can be connected to the processor so that the processor can read information from the storage medium and can write information to the storage medium. Optionally, the storage medium can also be integrated into the processor. The processor and the storage medium can be arranged in an ASIC, and the ASIC can be arranged in a user terminal. Optionally, the processor and the storage medium can also be arranged in different components in the user terminal.
[0096] In one or more exemplary designs, the above functions described in the embodiments of the present invention can be implemented in hardware, software, firmware or any combination of the three. If implemented in software, these functions can be stored on a computer-readable medium, or transmitted in the form of one or more instructions or codes on a computer-readable medium. Computer-readable media include computer storage media and communication media that facilitate the transfer of computer programs from one place to another. The storage medium can be any available medium that can be accessed by any general or special computer. For example, such computer-readable media can include but are not limited to RAM, ROM, EEPROM, CD-ROM or other optical disk storage, disk storage or other magnetic storage devices, or any other medium that can be used to carry or store program codes in the form of instructions or data structures and other forms that can be read by general or special computers, or general or special processors. In addition, any connection can be appropriately defined as a computer-readable medium, for example, if the software is transmitted from a website site, server or other remote resource through a coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL) or wirelessly, such as infrared, wireless and microwave, it is also included in the defined computer-readable medium. The disk and disc include compact disk, laser disk, optical disk, DVD, floppy disk and blue-ray disk. Disks usually copy data magnetically, while discs usually copy data optically with lasers. The above combination can also be included in computer readable media.
[0097] The specific implementation methods described above further illustrate the objectives, technical solutions and beneficial effects of the present invention in detail. It should be understood that the above description is only a specific implementation method of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. A method for image enhancement, characterized in that: include: Determine the foreground features and background features of the original image, separate the foreground features and background features in the original image, and form corresponding foreground images and background images respectively; wherein the foreground features are features corresponding to the target image of the original image, and the background features are features corresponding to the non-target image of the original image; The pixels of the background image are randomly transformed to generate a background transformation image, and a plurality of corresponding background transformation images are generated by randomly transforming the pixels of the background image multiple times; Merge each background transformation image with the foreground image to regenerate a corresponding enhanced image, wherein the enhanced image is used for image classification; The determining of the foreground features and the background features of the original image, separating the foreground features and the background features in the original image, and forming corresponding foreground images and background images respectively, specifically includes: Segmenting the original image through the image segmentation model to obtain the foreground features of the original image, saving the foreground features into an image to form a foreground image, wherein the foreground image includes blank areas except for the area where the foreground features are located, and the blank areas of the foreground image overlap with the areas where the background features of the original image are located; Set an empty matrix mask, where the empty matrix mask has the same width and height as the foreground image; Each point in the empty matrix mask is represented by the coordinates (u, v), and the empty matrix mask is adjusted by assigning the coordinate (u, v) value of a specific point in the empty matrix mask to 1, thereby obtaining a mask adjustment matrix; The specific point with a value of 1 in the mask adjustment matrix represents that the pixel at the corresponding position of the original image is a background pixel, and the mask adjustment matrix is used as the mask matrix of the background image; The original image is processed by the background image mask matrix to obtain background features, the pixels at the corresponding positions in the original image and the background image mask matrix with values of 1 are retained, and the pixels at the corresponding positions in the original image and the background image mask matrix with values of 0 are assigned values of 0 to obtain a background image of the background features of the original image; wherein the positions in the background image where the pixels are assigned values of 0 are blank areas of the background image.
2. The image enhancement method according to claim 1, characterized in that: The randomly transforming pixels of the background image to generate a background transformation image, and generating a plurality of corresponding background transformation images by randomly transforming pixels of the background image multiple times, specifically includes: The pixels of the background features in the background image are determined by the background image mask matrix, and the background image is transformed by randomly transforming the pixels of the background features in the background image to obtain a background transformed image.
3. The image enhancement method according to claim 2, characterized in that: The step of fusing each background transformation picture with the foreground picture to regenerate a corresponding enhanced picture specifically includes: Align the foreground image with the background transformation image so that the foreground features in the foreground image completely match the blank areas of the background transformation image; The pixels of each background transformation image are added and fused with the pixels at the corresponding position of the foreground image to generate an enhanced image; wherein the pixels of the blank area of the foreground image are set to 0.
4. The image enhancement method according to claim 1, characterized in that: The image segmentation model refers to a model formed by training the rembg background removal tool using the UNet network.
5. A picture enhancement device, characterized in that: include: An important feature separation module is used to determine the foreground features and background features of the original image, separate the foreground features and background features in the original image, and form corresponding foreground images and background images respectively; wherein the foreground features are features corresponding to the target image of the original image, and the background features are features corresponding to the non-target image of the original image; A background feature transformation module, used to randomly transform the pixels of the background image to generate a background transformation image, and to generate multiple corresponding background transformation images by randomly transforming the pixels of the background image multiple times; A fusion module, used to fuse each background transformation image with the foreground image to regenerate a corresponding enhanced image, wherein the enhanced image is used for image classification; The important feature separation module comprises: A foreground separation submodule is used to segment the original image through the image segmentation model to obtain the foreground features of the original image, save the foreground features into an image to form a foreground image, and the other areas in the foreground image except the area where the foreground features are located are blank areas, and the blank areas of the foreground image overlap with the areas where the background features of the original image are located; The empty matrix construction submodule is used to set an empty matrix mask, where the empty matrix mask has the same width and height as the foreground image; each point in the empty matrix mask is represented by a coordinate (u, v), and the empty matrix mask is adjusted by assigning a coordinate (u, v) value of a specific point in the empty matrix mask to 1, thereby obtaining a mask adjustment matrix; the specific point refers to a point corresponding to a blank area of the foreground image; A background construction submodule, wherein a specific point with a value of 1 in the mask adjustment matrix represents that the pixel at the corresponding position of the original image is a background pixel, and the mask adjustment matrix is used as a background image mask matrix; the original image is processed by the background image mask matrix to obtain background features, pixels at corresponding positions in the original image with a value of 1 in the background image mask matrix are retained, and pixels at corresponding positions in the original image with a value of 0 in the background image mask matrix are assigned a value of 0, to obtain a background image of the background features of the original image; wherein the positions in the background image where the pixels are assigned a value of 0 are blank areas of the background image.
6. The image enhancement device according to claim 5, characterized in that: The background feature transformation module is specifically used for: The pixels of the background features in the background image are determined by the background image mask matrix, and the background image is transformed by randomly transforming the pixels of the background features in the background image to obtain a background transformed image.
7. The image enhancement device according to claim 6, characterized in that: The fusion module includes: An alignment submodule, used to align the foreground image with the background transformation image so that the foreground features in the foreground image completely match the blank areas of the background transformation image; The fusion submodule is used to add and fuse the pixels of each background transformation image with the pixels at the corresponding position of the foreground image to generate an enhanced image; wherein the pixels of the blank area of the foreground image are set to 0.
8. The image enhancement device according to claim 5, characterized in that: The image segmentation model refers to a model formed by training the rembg background removal tool using the UNet network.
Citation Information
Patent Citations
Image processing method and related device
WO2019114571A1