Model training method and electronic equipment

Through image synthesis, the training data containing the embedding is generated and the neural network model is trained, which solves the problems of low precision of embedding image removal and difficult to obtain training data in the prior art, and achieves an efficient embedding image removal effect.

CN114492643BActive Publication Date: 2025-05-13VIVO MOBILE COMM CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210103440.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-01-27
Publication Date
2025-05-13
Estimated Expiration
2042-01-27

AI Technical Summary

Technical Problem

The existing neural network model has low accuracy in removing banding images, and requires a large number of pairs of banding-free images and banding-free images as training data, which is difficult to obtain.

Method used

By obtaining the first image and the preset embedding image without the embedding, the second image containing the embedding is synthesized, and inputting it into the original model for training, a model for extracting and removing the embedding image is generated.

Benefits of technology

Without a large amount of paired training data, sufficient training data is generated through image synthesis, which improves the accuracy and efficiency of removing the strap image of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114492643B_ABST
    Figure CN114492643B_ABST
Patent Text Reader

Abstract

The present application discloses a model training method and electronic device, belonging to the field of image processing technology. The model training method comprises: obtaining a first image without a fillet and a preset fillet image; synthesizing the first image and the preset fillet image, and outputting a second image containing the fillet; and inputting the second image into the original model for training to obtain a trained model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application belongs to the field of artificial intelligence technology, and specifically relates to a model training method and electronic equipment. Background Art

[0002] At present, when users use mobile terminals, taking photos is one of the most commonly used functions of the terminals. When shooting moving objects under indoor lighting, if the object moves too fast, a shorter shutter time must be used to ensure that the photographed object is not blurred. Figure 1 As shown in the figure, when the shutter speed is slow, the hand in the picture is more seriously blurred. When the shutter time is too short, although the hand motion blur disappears, black or colored stripes appear in the captured image, also known as the banding phenomenon, which affects the image quality.

[0003] There is a method for removing the banding of the image by using a neural network model. However, the accuracy of removing the banding of the image by using the existing neural network model is low. Summary of the invention

[0004] The purpose of the embodiments of the present application is to provide a model training method and electronic device that can solve the problem of low accuracy in removing banding of image embeddings using existing neural network models.

[0005] In a first aspect, an embodiment of the present application provides a model training method, comprising:

[0006] Acquire a first image without a fillet and a preset fillet image;

[0007] Synthesize the first image and the preset fillet image, and output a second image containing the fillet;

[0008] The second image is input into the original model for training to obtain a trained model.

[0009] In a second aspect, an embodiment of the present application provides a method for removing a striped image by using a model, wherein the model is trained by the method of the first aspect, and the method comprises:

[0010] Obtaining a to-be-processed image of the photographed object containing a fillet and a reference image of the photographed object not containing a fillet;

[0011] Inputting a to-be-processed image of the photographed object containing a fillet and a reference image of the photographed object not containing a fillet into the model, and outputting a fillet image;

[0012] Divide the pixel value of the image to be processed containing the fillet by the pixel value of the fillet image, and output a fifth image;

[0013] The fifth image is adjusted to a target image format, and an image to be processed with the stripe image removed is output.

[0014] In a third aspect, an embodiment of the present application provides a model training device, comprising:

[0015] An acquisition module, used for acquiring a first image without a fillet and a preset fillet image;

[0016] A synthesis module, used for synthesizing the first image and the preset fillet image, and outputting a second image containing the fillet;

[0017] The training module is used to input the second image into the original model for training to obtain a trained model.

[0018] In a fourth aspect, an embodiment of the present application provides a device for removing a stripe image using a model, wherein the model is obtained by training the device of the third aspect, and the device includes:

[0019] An acquisition module, used for acquiring a to-be-processed image of the photographed object containing a stripe and a reference image of the photographed object not containing a stripe;

[0020] An output module, used for inputting a to-be-processed image of the photographed object containing a fillet and a reference image of the photographed object not containing a fillet into the model, and outputting a fillet image;

[0021] An adjustment module, used for dividing the pixel value of the image to be processed containing the fillet by the pixel value of the fillet image, and outputting a fifth image;

[0022] The adjustment module is further used to adjust the fifth image to a target image format and output an image to be processed with the stripe image removed.

[0023] In a fifth aspect, an embodiment of the present application provides an electronic device, which includes a processor and a memory, wherein the memory stores programs or instructions that can be run on the processor, and when the programs or instructions are executed by the processor, the steps of the methods described in the first and second aspects are implemented.

[0024] In a sixth aspect, an embodiment of the present application provides a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, the steps of the method described in the first aspect and the second aspect are implemented.

[0025] In the seventh aspect, an embodiment of the present application provides a chip, which includes a processor and a communication interface, the communication interface is coupled to the processor, and the processor is used to run programs or instructions to implement the methods described in the first and second aspects.

[0026] In an eighth aspect, an embodiment of the present application provides a computer program product, which is stored in a storage medium and is executed by at least one processor to implement the method described in the first aspect and the second aspect.

[0027] In the embodiment of the present application, a large number of pairs of images without banding and images with banding are not required as training data. Instead, a first image without fillets and a preset fillet image are synthesized to output a second image containing fillets, and the second image is input into an original model for training. The original model is used to extract the fillet image of the input image and output the input image after the fillet image is removed. There is no need to actually collect pairs of images without banding and images with banding of the same object. Instead, a first image without fillets and a preset fillet image are obtained, and then a second image is obtained by synthesizing the first image and the preset fillet image. The second image is used as a banding image of the same object as the first image. In this way, a banding image of the same object as the image without banding is generated by image synthesis, and the training data requirements for model training can be met by only obtaining images without banding and images with fillets. In practical applications, the first image without the fillet and the image with the fillet can be easily obtained, and the number of images containing the fillet that are the same as the subject without the fillet is very small. The present application does not need to obtain the image containing the fillet, but instead generates the second image containing the fillet through synthesis, and uses the second image containing the fillet and the first image without the fillet as training data, which can ensure sufficient training data and improve model accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] Figure 1 It is a flowchart of a model training method provided in an embodiment of the present application;

[0029] Figure 2 It is a schematic diagram of a raw image and a four-channel image array in a Bayer format provided in an embodiment of the present application;

[0030] Figure 3 It is a schematic diagram of the improved principle of a Unet model provided in an embodiment of the present application;

[0031] Figure 4 It is a schematic diagram of an existing Unet model structure provided in an embodiment of the present application;

[0032] Figure 5 It is a schematic diagram of an improved Unet model structure provided in an embodiment of the present application;

[0033] Figure 6 It is a flowchart of a method for removing a fillet image by applying a model provided in an embodiment of the present application;

[0034] Figure 7 It is a structural schematic diagram of a model training device provided in an embodiment of the present application;

[0035] Figure 8 It is a structural schematic diagram of a device for removing a fillet image using a model provided in an embodiment of the present application;

[0036] Fig. 9 It is a structural schematic diagram of an electronic device provided in an embodiment of the present application;

[0037] Fig.10 It is a schematic diagram of the structure of another electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0038] The following will be combined with the drawings in the embodiments of the present application to clearly describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments in the present application belong to the scope of protection of this application.

[0039] The terms "first", "second", etc. in the specification and claims of the present application are used to distinguish similar objects, and are not used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable under appropriate circumstances, so that the embodiments of the present application can be implemented in an order other than those illustrated or described here, and the objects distinguished by "first", "second", etc. are generally of one type, and the number of objects is not limited. For example, the first object can be one or more. In addition, "and / or" in the specification and claims represents at least one of the connected objects, and the character " / " generally indicates that the objects associated with each other are in an "or" relationship.

[0040] Existing methods for removing banding from image inserts using neural network models require a large number of paired images with and without banding as training data. However, in practice, training data is difficult to obtain, resulting in low accuracy in removing banding from image inserts using existing neural network models.

[0041] The following is a detailed description of the model training method provided by the embodiment of the present application through specific embodiments and their application scenarios in conjunction with the accompanying drawings.

[0042] Figure 1 FIG. 1 is a flow chart of a method for training a model provided by an embodiment of the present application. Figure 1 As shown, the method may include the following steps:

[0043] S110, acquiring a first image without a fillet and a preset fillet image.

[0044] The first image without the fillet can be obtained from an electronic device storing the first image, and the first image can be a raw image file with internal colors in a wide color gamut, which can be precisely adjusted, and some simple modifications can be made before conversion to improve the accuracy of subsequent model training. The preset fillet image can be obtained by pre-cutting the fillet image from an image containing the fillet image.

[0045] S120, synthesizing the first image and the preset fillet image, and outputting a second image including the fillet.

[0046] Among them, the image synthesis algorithm can synthesize images. When generating the second image, the image synthesis algorithm can be applied to synthesize the first image and the preset fillet image. By using the second image as the training data of the model, there is no need to actually collect the non-banding image and the banding image of the same photographed object, but to obtain the first image without fillet and the preset fillet image, and then obtain the second image by synthesizing the first image and the preset fillet image, and use the second image as the same banding image of the first image photographed object. In this way, the banding image with the same photographed object as the non-banding image is generated by image synthesis, and only the non-banding image and the fillet image need to be obtained to meet the training data requirements of the model training. In practical applications, the first image without fillet and the fillet image can be easily obtained, and the number of images containing fillets that are the same as the photographed object without fillet is very small. In this application, there is no need to obtain the image containing fillet, but to generate the second image containing fillet by synthesis, and use the second image containing fillet and the first image without fillet as training data, which can ensure sufficient training data and improve the accuracy of the model.

[0047] In one embodiment, S120: synthesizing the first image and the preset fillet image to output a second image containing the fillet may include:

[0048] S1201: Convert a first image and a preset fillet image into a preset image format.

[0049] Among them, considering that the two images to be synthesized need to be in the same format during image synthesis, when synthesizing the second image, the embodiment of the present application first converts the first image and the preset fillet image into a preset image format, and the preset image format includes image size and image channel array, etc. For example, the first image is a 14-bit original image file raw image in grbg format stored in bayer format. First, convert it to uint16 format, then subtract the black level and normalize the data to [0, 1]. At this time, assuming that the raw image size is [4080, 3060, 1], convert it to a grbg four-channel image array, the bayer format raw image and the four-channel image array are as follows: Figure 2 As shown, the image size is [2040, 1530, 4]. After resizing it to [300, 300, 4], the image size is reduced to improve the efficiency of subsequent model training, and the conversion of the first image to the preset image format is completed. Similarly, assuming that the preset fillet image is a grayscale single-channel mask image stored in png format, the single channel of the mask image is copied to four channels, which is consistent with the grbg image of the raw image, and then the pixel values ​​of the mask image are normalized to between [0, 1]. Then, the four channels of the mask image are multiplied by four random numbers respectively, and the grayscale mask is converted to a color mask. Finally, the image is adjusted to [300, 300, 4], which completes the conversion of the preset fillet image to the preset image format.

[0050] S1202: randomly intercept at least two first image blocks from a first image converted into a preset image format, and randomly intercept at least one second image block from a preset fill image converted into a preset image format.

[0051] Among them, it is considered that the effect of banding elimination for a single banding image will be relatively poor. Based on the principle of banding, in actual scenes, banding is likely to occur when the shutter speed of the camera exceeds 1 / 100s, and it is not easy to occur when the shutter speed is lower than 1 / 100. Therefore, using a low shutter motion blurred image and a high shutter banded image, and inputting them into the deep learning model at the same time will bring better results. However, paired long exposure + short exposure images are difficult to obtain. Two raw images of the same size can be randomly cropped in a single raw image to simulate the difference between long exposure and short exposure images (due to hand movement when taking images, there is often a certain offset between the long exposure and short exposure images). In the embodiment of the present application, the two first image blocks are respectively used as long exposure images and short exposure images. For example, two first image blocks M1 and M2 of size [256, 256, 4] are randomly obtained from the first image converted to a preset image format. Wherein M1 simulates a long exposure image, M2 simulates a short exposure image, the two first image blocks are involved in the model training, and at least one second image block is randomly intercepted from the preset fillet image converted into the preset image format to cooperate with the first image block to participate in the model training, simulating the banding principle and improving the model training accuracy. For example, an image block M3 with a size of [256, 256, 4] is intercepted as the second image block.

[0052] S1203: synthesize at least two first image blocks and at least one second image block, and output a second image.

[0053] In one embodiment, S1203: synthesizing at least two first image blocks and at least one second image block, and outputting a second image, may include:

[0054] Multiplying the pixel values ​​of one of the first image blocks and the second image to output a third image block;

[0055] The image channels of the third image block and another first image block are added together to output a second image.

[0056] Among them, the pixel values ​​of one of the first image blocks and the second image are multiplied to output a third image block, and then the image channels of the third image block and another first image block are added. The output second image can simulate an image containing banding, thereby realizing the synthesis of an image containing a fillet.

[0057] S130, inputting the second image into the original model for training to obtain a trained model.

[0058] The original model is used to extract the fillet image of the input image and output the input image after removing the fillet image; the original model can be an image processing model, which can extract the target image in the image and remove the target image in the image, and the model can be obtained by inputting the second image into the original model for training. Considering that the sub-model for extracting the fillet image and the sub-model for removing the fillet image are trained simultaneously in the model training stage in order to improve the model accuracy, after the model meets the training stop condition in the training stage, that is, when the first loss function value, the second loss function value, and the third loss function value all meet the preset conditions, the model training is considered to be completed, and the sub-model for outputting the fillet image in the original model can be applied to the fillet image removal method, and the fillet in the input image can be accurately removed by the model.

[0059] In one embodiment, S130: inputting the second image into the original model for training may include:

[0060] The second image is input into the original model to extract the fillet image and remove the fillet image, and the fillet extraction image and the third image are output.

[0061] Among them, the original model can be selected as the Unet model, but considering that only using the Unet model for stripe image extraction and stripe image removal, it is easy to have problems such as uneven color blocks of the extracted stripe image and unclean stripe image removal. Figure 3 As shown, in the embodiment of the present application, an output branch Output2 is added to the original output branch Output1 of the Unet model. One of the two branches is used for stripe image extraction, and the other is used for stripe image removal. The two branches perform corresponding work independently. Compared with the existing Unet model that only uses one output branch to simultaneously extract and remove stripe images, it avoids mixing of parameters, improves the model accuracy, and makes the extracted stripe image color blocks uniform and the stripe image is removed cleanly.

[0062] Specifically, Figure 4 As shown in , the existing Unet model consists of 4 downsampling units, 4 upsampling units, and two double conv layers. Each downsampling unit is a down conv layer structure, which consists of a downsampling layer and a double conv unit. Each upsampling unit consists of a double conv unit and a concatenation layer. The upsampling feature map and the downsampling feature map of the same size are concatenated as the input of the next convolution layer. Figure 5 As shown, compared with the Unet model, the improved Unet model of the present application adds an upsampling branch that is the same as the original upsampling branch on the basis of the Unet model.

[0063] A first loss function value is calculated based on the fillet extraction image and the second image block, a second loss function value is calculated based on the third image and one of the first image blocks, and a third loss function value is calculated based on the third image block and a fourth image, wherein the fourth image is an image output after the pixel value of the third image and the fillet extraction image are multiplied. When the first loss function value, the second loss function value, and the third loss function value all meet the preset conditions, the training is completed.

[0064] Among them, considering that the construction of the loss function directly affects the accuracy of the model, the present application redesigns the method for calculating the loss function value, respectively calculating the first loss function value according to the fillet extraction image and the second image block, calculating the second loss function value according to the third image and one of the first image blocks, and calculating the third loss function value according to the third image block and the fourth image, wherein the fourth image is the image output after the pixel value of the third image and the fillet extraction image is multiplied. Then, when the first loss function value, the second loss function value, and the third loss function value all meet the preset conditions, the training is completed. For example, assuming that the first loss function value, the second loss function value, and the third loss function value obtained by directly settling the model input and output are A, B, and C, respectively. During the training process, the orders of magnitude of these three losses may differ in magnitude. For example, assuming A=400, B=4000, and C=40000, then the three losses need to be weighted and adjusted to the same order of magnitude. For example, the weights of A, B, and C are 1, 0.1, and 0.01, respectively. The adjusted A, B, and C are in the same order of magnitude, and then A, B, and C are summed. The preset condition can be that the sum of A, B, and C fluctuates within a range of less than 0.01 in five consecutive rounds of training.

[0065] The embodiment of the present application redesigns the loss function and uses the fillet extraction image, the second image block, the third image and the first image block to participate in the determination of model stop, thereby improving the model training accuracy.

[0066] In the embodiment of the present application, a large number of paired images without banding and images with banding are not required as training data. Instead, a first image without fillets and a preset fillet image are synthesized to output a second image containing fillets. The second image is input into an original model for training. The original model is used to extract a fillet image of the input image and output the input image after the fillet image is removed. There is no need to actually collect images without banding and images with banding of the same object. Instead, a first image without fillets and a preset fillet image are obtained, and then a second image is obtained by synthesizing the first image and the preset fillet image. The second image is used as a banding image of the same object as the first image. In this way, a banding image of the same object as the image without banding is generated by image synthesis. Only images without banding and fillet images need to be obtained to meet the training data requirements for model training. In practical applications, the first image without the fillet and the fillet image can be easily obtained, and the number of images containing fillets that are the same as the subject without the fillet is very small. The present application does not need to obtain the image containing the fillet, but generates the second image containing the fillet by synthesis, and uses the second image containing the fillet and the first image without the fillet as training data, which can ensure sufficient training data and improve the model accuracy. In addition, in the model training, the embodiment of the present application uses smaller first image blocks and second image blocks for model training, which is much smaller than the existing model in terms of data processing volume, and the model training efficiency is higher.

[0067] The above describes the training method of the model provided in the embodiment of the present application. Based on the model obtained by training the training method of the application model, the embodiment of the present application also provides a method for applying the model to remove the stripe image, such as Figure 6 As shown, the method includes:

[0068] S610: Acquire a to-be-processed image of the photographed object that contains a stripe and a reference image of the photographed object that does not contain a stripe.

[0069] When the model is applied to remove the mosaic image from the image, a long exposure image of the object can be actually obtained, that is, a reference image without mosaics, and a short exposure image of the object can be obtained, that is, an image to be processed containing mosaics.

[0070] S620: Input the to-be-processed image of the photographed object containing the fillet and the reference image of the photographed object not containing the fillet into the model, and output the fillet image.

[0071] Before the image to be processed containing the fillet of the photographed object and the reference image of the photographed object without the fillet are input into the model, the input image needs to be preprocessed first. The embodiment of the present application preprocesses the image to be processed and the reference image, including image size adjustment and image channel array adjustment. For example, the image to be processed and the reference image are converted as follows: converted to uint16 format. Then the black level is subtracted and the data is normalized to between [0,1]. At this time, it is assumed that the image width and height are [4080, 3060], and then it is converted to the grbg four-channel image array [2040, 1530, 4], and then resized to [256, 256, 4] to obtain the preset image format. Then, the image channels of the image to be processed and the reference image converted to the preset image format are added, and the output image is used as the model input image. Among them, adding the image channels can facilitate the matching of the image input format of the model, improve the recognition efficiency of the fillet image, and thus improve the efficiency of removing the fillet image.

[0072] S630: Divide the pixel value of the image to be processed containing the fillet by the pixel value of the fillet image, and output a fifth image.

[0073] After the fillet image is identified by the model, the fillet image in the to-be-processed image containing the fillet can be removed by a pixel division method.

[0074] S640: Adjust the fifth image to a target image format, and output the image to be processed with the stripe image removed.

[0075] In one embodiment, S640 may include:

[0076] The fifth image is adjusted to a target image size, a target pixel value, and a target black level value, and an image to be processed with the stripe image removed is output.

[0077] Among them, considering that the fifth image is an image format that is convenient for model output and is not suitable for user viewing, after the fifth image is output, the fifth image can also be adjusted to the target image format for user viewing, and the fifth image can be adjusted from dimensions such as image size, pixel value and black level value to match the light sensitivity range of the human eye. For example, the image size of the fifth image is adjusted to [4080, 3060, 1], and then the pixel value range of [0, 1] is proportionally enlarged to the pixel value range of [0, 2^14], and then the black level value is increased, and finally the pixel values ​​below 0 are set to 0, and the pixel values ​​greater than or equal to 2^14 are set to 2^13, which ensures that the image can have a better viewing effect for the human eye and optimizes the user experience.

[0078] In an embodiment of the present application, a model trained with sufficient training data is used to remove the inlaid image from the target image. The removal accuracy is high, and the inlaid image of the target image can be cleanly removed. The removal effect is good, and it is ensured that the image to be processed after removing the inlaid image can have a good viewing effect for the human eye, thereby optimizing the user experience.

[0079] The model training method provided in the embodiment of the present application can be executed by a model training device; the method for removing the fillet image by applying the model provided in the embodiment of the present application can be executed by a device for removing the fillet image by applying the model. In the embodiment of the present application, each device performs its own corresponding method to illustrate each device provided in the embodiment of the present application.

[0080] Figure 7 A schematic diagram of the structure of a training device for a model provided by an embodiment of the present application is shown. Figure 7 Each module in the device shown has the function of realizing Figure 1 The functions of each step in the process can achieve the corresponding technical effects. Figure 7 As shown, the device may include:

[0081] The acquisition module 710 is used to acquire a first image without a fillet and a preset fillet image.

[0082] The synthesis module 720 is used to synthesize the first image and the preset fillet image, and output a second image containing the fillet.

[0083] The training module 730 is used to input the second image into the original model for training to obtain a trained model.

[0084] In one embodiment, the synthesis module 720 includes a conversion unit 7201 , a capture unit 7202 and a synthesis unit 7203 .

[0085] The conversion unit 7201 is used to convert the first image and the preset fillet image into a preset image format.

[0086] The interception unit 7202 is configured to randomly intercept at least two first image blocks from the first image converted into a preset image format, and randomly intercept at least one second image block from the preset fill image converted into the preset image format.

[0087] The synthesis unit 7203 is used to synthesize at least two first image blocks and at least one second image block, and output a second image.

[0088] In one embodiment, the synthesis unit 7203 is specifically used to:

[0089] Multiply the pixel values ​​of one of the first image blocks and the second image to output a third image block.

[0090] The image channels of the third image block and another first image block are added together to output a second image.

[0091] In one embodiment, the training module is specifically used to:

[0092] The second image is input into the original model to extract the fillet image and remove the fillet image, and the fillet extraction image and the third image are output.

[0093] A first loss function value is calculated based on the fillet extraction image and the second image block, a second loss function value is calculated based on the third image and one of the first image blocks, and a third loss function value is calculated based on the third image block and a fourth image. The fourth image is an image output after the pixel values ​​of the third image and the fillet extraction image are multiplied.

[0094] When the first loss function value, the second loss function value, and the third loss function value all meet the preset conditions, the training is completed.

[0095] In the embodiment of the present application, a large number of paired images without banding and images with banding are not required as training data. Instead, a first image without fillets and a preset fillet image are synthesized to output a second image containing fillets. The second image is input into an original model for training. The original model is used to extract a fillet image of the input image and output the input image after the fillet image is removed. There is no need to actually collect images without banding and images with banding of the same object. Instead, a first image without fillets and a preset fillet image are obtained, and then a second image is obtained by synthesizing the first image and the preset fillet image. The second image is used as a banding image of the same object as the first image. In this way, a banding image of the same object as the image without banding is generated by image synthesis. Only images without banding and fillet images need to be obtained to meet the training data requirements for model training. In practical applications, the first image without the fillet and the fillet image can be easily obtained, and the number of images containing fillets that are the same as the subject without the fillet is very small. The present application does not need to obtain the image containing the fillet, but generates the second image containing the fillet by synthesis, and uses the second image containing the fillet and the first image without the fillet as training data, which can ensure sufficient training data and improve the model accuracy. In addition, in the model training, the embodiment of the present application uses smaller first image blocks and second image blocks for model training, which is much smaller than the existing model in terms of data processing volume, and the model training efficiency is higher.

[0096] Figure 8 A schematic diagram of the structure of an apparatus for removing a fillet image by applying a model provided by an embodiment of the present application is shown. Figure 8 Each module in the device shown has the function of realizing Figure 6 The functions of each step in the process can achieve the corresponding technical effects. Figure 8 As shown, the device may include:

[0097] An acquisition module 810 is used to acquire a to-be-processed image of the photographed object containing a stripe and a reference image of the photographed object not containing a stripe;

[0098] An output module 820 is used to input the to-be-processed image of the photographed object containing the fillet and the reference image of the photographed object not containing the fillet into the model, and output the fillet image;

[0099] An adjustment module 830, configured to divide the pixel value of the image to be processed containing the fillet by the pixel value of the fillet image, and output a fifth image;

[0100] The adjustment module 830 is further configured to adjust the fifth image to a target image format, and output an image to be processed with the stripe image removed.

[0101] In one embodiment, the adjustment module 830 may be specifically configured to:

[0102] The fifth image is adjusted to a target image size, a target pixel value, and a target black level value, and an image to be processed with the stripe image removed is output.

[0103] In the embodiment of the present application, a model trained with sufficient training data is used to remove the stripe image from the target image. The removal accuracy is high, the stripe image of the target image can be cleanly removed, and the removal effect is good.

[0104] The training device of the model in the embodiment of the present application can be a device, or a component, integrated circuit, or chip in a terminal. The device can be a mobile electronic device or a non-mobile electronic device. Exemplarily, the mobile electronic device can be a mobile phone, a tablet computer, a laptop computer, a PDA, a car electronic device, a mobile Internet device (Mobile Internet Device, MID), an augmented reality (augmented reality, AR) / virtual reality (virtual reality, VR) device, a robot, a wearable device, an ultra-mobile personal computer (ultra-mobile personal computer, UMPC), a netbook or a personal digital assistant (personal digital assistant, PDA), etc., and can also be a server, a network attached storage (Network Attached Storage, NAS), a personal computer (personal computer, PC), a television (television, TV), a teller machine or a self-service machine, etc., which is not specifically limited in the embodiment of the present application.

[0105] The training device of the model in the embodiment of the present application may be a device having an operating system. The operating system may be an Android operating system, an iOS operating system, or other possible operating systems, which are not specifically limited in the embodiment of the present application.

[0106] The device provided in the embodiment of the present application can achieve Figures 1 to 6 To avoid repetition, the various processes implemented by the method embodiment are not described here.

[0107] Alternatively, if Fig. 9 As shown, an embodiment of the present application also provides an electronic device 900, including a processor 901, a memory 902, and a program or instruction stored in the memory 902 and executable on the processor 901. When the program or instruction is executed by the processor 901, each process of the training method embodiment of the above-mentioned model is implemented, and the same technical effect can be achieved. To avoid repetition, it will not be repeated here.

[0108] It should be noted that the electronic devices in the embodiments of the present application include the mobile electronic devices and non-mobile electronic devices mentioned above.

[0109] Fig.10 A schematic diagram of the hardware structure of an electronic device to implement an embodiment of the present application.

[0110] The electronic device 100 includes but is not limited to components such as a radio frequency unit 101, a network module 102, an audio output unit 103, an input unit 104, a sensor 105, a display unit 106, a user input unit 107, an interface unit 108, a memory 109, and a processor 110.

[0111] Those skilled in the art will appreciate that the electronic device 100 may also include a power source (such as a battery) for supplying power to various components, and the power source may be logically connected to the processor 110 through a power management system, thereby implementing functions such as managing charging, discharging, and power consumption management through the power management system. Fig.10 The electronic device structure shown in the figure does not constitute a limitation on the electronic device. The electronic device may include more or fewer components than shown in the figure, or combine certain components, or arrange the components differently, which will not be described in detail here.

[0112] The input unit 104 is used to obtain a first image without a fillet and a preset fillet image; synthesize the first image and the preset fillet image to output a second image containing the fillet; input the second image into an original model for training, the original model is used to extract the fillet image of the input image and output the input image after the fillet image is removed; and determine the sub-model in the original model for outputting the fillet image as the model.

[0113] In the embodiment of the present application, a large number of paired images without banding and images with banding are not required as training data. Instead, a first image without fillets and a preset fillet image are synthesized to output a second image containing fillets. The second image is input into an original model for training. The original model is used to extract a fillet image of the input image and output the input image after the fillet image is removed. There is no need to actually collect images without banding and images with banding of the same object. Instead, a first image without fillets and a preset fillet image are obtained, and then a second image is obtained by synthesizing the first image and the preset fillet image. The second image is used as a banding image of the same object as the first image. In this way, a banding image of the same object as the image without banding is generated by image synthesis. Only images without banding and fillet images need to be obtained to meet the training data requirements for model training. In practical applications, the first image without the fillet and the image with the fillet can be easily obtained, and the number of images containing the fillet that are the same as the subject without the fillet is very small. The present application does not need to obtain the image containing the fillet, but instead generates the second image containing the fillet through synthesis, and uses the second image containing the fillet and the first image without the fillet as training data, which can ensure sufficient training data and improve model accuracy.

[0114] Optionally, the input unit 104 is further used to convert the first image and the preset fillet image into a preset image format; randomly intercept at least two first image blocks from the first image converted into the preset image format, and randomly intercept at least one second image block from the preset fillet image converted into the preset image format; synthesize the at least two first image blocks and the at least one second image block, and output the second image.

[0115] Optionally, the input unit 104 is further configured to multiply pixel values ​​of one of the first image blocks and the second image to output a third image block; and add image channels of the third image block and another first image block to output a second image.

[0116] Among them, the pixel values ​​of one of the first image blocks and the second image are multiplied to output a third image block, and then the image channels of the third image block and another first image block are added. The output second image can simulate an image containing banding, thereby realizing the synthesis of an image containing a fillet.

[0117] Optionally, the input unit 104 is further used to input the second image into the original model to extract the fillet image and remove the fillet image, and output the fillet extracted image and the third image; calculate the first loss function value according to the fillet extracted image and the second image block, calculate the second loss function value according to the third image and one of the first image blocks, and calculate the third loss function value according to the third image block and the fourth image, wherein the fourth image is an image output after the pixel value of the third image and the fillet extracted image is multiplied. When the first loss function value, the second loss function value, and the third loss function value all meet the preset conditions, the training is completed.

[0118] The embodiment of the present application redesigns the loss function and uses the fillet extraction image, the second image block, the third image and the first image block to participate in the determination of model stop, thereby improving the model training accuracy.

[0119] In the embodiment of the present application, there is no need to actually collect the image without banding and the image with banding of the same photographic object. Instead, a first image without fillet and a preset fillet image are obtained, and then the second image is obtained by synthesizing the first image and the preset fillet image, and the second image is used as the image with banding of the same photographic object as the first image. In this way, the image with banding of the same photographic object as the image without banding is generated by image synthesis, and only the image without banding and the fillet image need to be obtained to meet the training data requirements of model training. In practical applications, the first image without fillet and the fillet image can be easily obtained, and the number of images containing fillets that are the same as the photographic object without fillet is very small. In the present application, there is no need to obtain the image containing fillet, but the second image containing fillet is generated by synthesis, and the second image containing fillet and the first image without fillet are used as training data, which can ensure sufficient training data and improve model accuracy. In addition, in the model training, the embodiment of the present application uses smaller first image blocks and second image blocks for model training, which is much smaller than the existing model in terms of data processing volume, and the model training efficiency is high.

[0120] In one embodiment, the input unit 104 is further used to obtain a to-be-processed image of the photographed object containing a fillet and a reference image of the photographed object not containing a fillet; convert the to-be-processed image and the reference image into a preset image format; synthesize the to-be-processed image and the reference image after being converted into the preset image format, and output a target image; input the target image into the model to remove the fillet image, and obtain the to-be-processed image after the fillet image is removed.

[0121] Optionally, the input unit 104 is further configured to add image channels of the image to be processed and the reference image after being converted into a preset image format, and output a target image.

[0122] In the embodiment of the present application, a model trained with sufficient training data is used to remove the stripe image from the target image. The removal accuracy is high, the stripe image of the target image can be cleanly removed, and the removal effect is good.

[0123] It should be understood that in the embodiment of the present application, the input unit 104 may include a graphics processor (Graphics Processing Unit, GPU) 1041 and a microphone 1042, and the graphics processor 1041 processes the image data of a static picture or video obtained by an image capture device (such as a camera) in a video capture mode or an image capture mode. The display unit 106 may include a display panel 1061, and the display panel 1061 may be configured in the form of a liquid crystal display, an organic light emitting diode, etc. The user input unit 107 includes a touch panel 1071 and at least one of other input devices 1072. The touch panel 1071 is also called a touch screen. The touch panel 1071 may include two parts: a touch detection device and a touch controller. Other input devices 1072 may include, but are not limited to, a physical keyboard, function keys (such as volume control keys, switch keys, etc.), a trackball, a mouse, and a joystick, which will not be repeated here.

[0124] The memory 109 can be used to store software programs and various data. The memory 109 may mainly include a first storage area for storing programs or instructions and a second storage area for storing data, wherein the first storage area may store an operating system, an application program or instructions required for at least one function (such as a sound playback function, an image playback function, etc.), etc. In addition, the memory 109 may include a volatile memory or a non-volatile memory, or the memory 109 may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), a static random access memory (SRAM), a dynamic random access memory (DRAM), a synchronous dynamic random access memory (SDRAM), a double data rate synchronous dynamic random access memory (DDRSDRAM), an enhanced synchronous dynamic random access memory (ESDRAM), a synchronous link dynamic random access memory (SLDRAM) and a direct memory bus random access memory (DRRAM). The memory 109 in the embodiment of the present application includes but is not limited to these and any other suitable types of memory.

[0125] The processor 110 may include one or more processing units; optionally, the processor 110 integrates an application processor and a modem processor, wherein the application processor mainly processes operations related to an operating system, a user interface, and application programs, and the modem processor mainly processes wireless communication signals, such as a baseband processor. It is understandable that the modem processor may not be integrated into the processor 110.

[0126] An embodiment of the present application also provides a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, the various processes of the above-mentioned method embodiment are implemented and the same technical effect can be achieved. To avoid repetition, it will not be repeated here.

[0127] The processor is the processor in the electronic device described in the above embodiment. The readable storage medium includes a computer readable storage medium, such as a computer read-only memory ROM, a random access memory RAM, a magnetic disk or an optical disk.

[0128] An embodiment of the present application further provides a chip, which includes a processor and a communication interface, wherein the communication interface is coupled to the processor, and the processor is used to run programs or instructions to implement the various processes of the above-mentioned method embodiment, and can achieve the same technical effect. To avoid repetition, it will not be repeated here.

[0129] It should be understood that the chip mentioned in the embodiments of the present application can also be called a system-level chip, a system chip, a chip system or a system-on-chip chip, etc.

[0130] An embodiment of the present application provides a computer program product, which is stored in a storage medium. The program product is executed by at least one processor to implement the various processes of the above-mentioned method embodiment and can achieve the same technical effect. To avoid repetition, it will not be repeated here.

[0131] It should be noted that, in this article, the terms "comprise", "include" or any other variant thereof are intended to cover non-exclusive inclusion, so that the process, method, article or device including a series of elements includes not only those elements, but also includes other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise one..." do not exclude the presence of other identical elements in the process, method, article or device including the element. In addition, it should be pointed out that the scope of the methods and devices in the embodiments of the present application is not limited to performing functions in the order shown or discussed, and may also include performing functions in a substantially simultaneous manner or in reverse order according to the functions involved, for example, the described method may be performed in an order different from that described, and various steps may also be added, omitted, or combined. In addition, the features described with reference to certain examples may be combined in other examples.

[0132] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus a necessary general hardware platform, and of course by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a computer software product, which is stored in a storage medium (such as ROM / RAM, a disk, or an optical disk), and includes a number of instructions for a terminal (which can be a mobile phone, a computer, a server, or a network device, etc.) to execute the methods described in each embodiment of the present application.

[0133] The embodiments of the present application are described above in conjunction with the accompanying drawings, but the present application is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Under the guidance of the present application, ordinary technicians in this field can also make many forms without departing from the purpose of the present application and the scope of protection of the claims, all of which are within the protection of the present application.

Claims

1. A model training method, characterized in that: include: Acquire a first image without a fillet and a preset fillet image; Synthesize the first image and the preset fillet image, and output a second image containing the fillet; Inputting the second image into the original model for training to obtain a trained model; The step of synthesizing the first image and the preset fillet image to output a second image containing a fillet includes: An image synthesis algorithm is applied to synthesize the first image and the preset fillet image to obtain the second image.

2. The training method of the model as claimed in claim 1, characterized in that: The step of synthesizing the first image and the preset fillet image to output a second image containing a fillet includes: Converting the first image and the preset fillet image into a preset image format; Randomly intercepting at least two first image blocks from a first image converted into a preset image format, and randomly intercepting at least one second image block from a preset fill image converted into a preset image format; The at least two first image blocks and the at least one second image block are synthesized, and the second image is output.

3. The training method of the model as claimed in claim 2, characterized in that: The synthesizing the at least two first image blocks and the at least one second image block, and outputting the second image, comprises: multiplying pixel values ​​of one of the first image blocks and the second image to output a third image block; The third image block and another image channel of the first image block are added together to output the second image.

4. The training method of the model as claimed in claim 3, characterized in that: The step of inputting the second image into the original model for training to obtain a trained model comprises: Input the second image into the original model, extract the fillet image and remove the fillet image, and output the fillet extracted image and the third image; Calculating a first loss function value according to the fillet extraction image and the second image block, calculating a second loss function value according to the third image and one of the first image blocks, and calculating a third loss function value according to the third image block and a fourth image, wherein the fourth image is an image output after the third image is multiplied by a pixel value of the fillet extraction image; When the first loss function value, the second loss function value, and the third loss function value all meet preset conditions, the training is completed.

5. A method for removing a stripe image by using a model, characterized in that: The model is trained by the method according to claim 1, wherein the method comprises: Obtaining a to-be-processed image of the photographed object containing a fillet and a reference image of the photographed object not containing a fillet; Inputting a to-be-processed image of the photographed object containing a fillet and a reference image of the photographed object not containing a fillet into the model, and outputting a fillet image; Dividing the pixel value of the image to be processed containing the fillet by the pixel value of the fillet image, and outputting a fifth image; The fifth image is adjusted to a target image format, and an image to be processed with the stripe image removed is output.

6. A model training device, characterized in that: include: An acquisition module, used for acquiring a first image without a fillet and a preset fillet image; A synthesis module, used for synthesizing the first image and the preset fillet image, and outputting a second image containing the fillet; A training module, used for inputting the second image into an original model for training to obtain a trained model; The synthesis module is specifically used for: An image synthesis algorithm is applied to synthesize the first image and the preset fillet image to obtain the second image.

7. The training device of the model as claimed in claim 6, characterized in that: The synthesis module includes a conversion unit, an interception unit and a synthesis unit; The conversion unit is used to convert the first image and the preset fillet image into a preset image format; The interception unit is used to randomly intercept at least two first image blocks from the first image converted into the preset image format, and randomly intercept at least one second image block from the preset fill image converted into the preset image format; The synthesis unit is used to synthesize the at least two first image blocks and the at least one second image block, and output the second image.

8. The training device of the model as claimed in claim 7, characterized in that: The synthesis unit is specifically used for: multiplying pixel values ​​of one of the first image blocks and the second image to output a third image block; The third image block and another image channel of the first image block are added together to output the second image.

9. The training device of the model as claimed in claim 8, characterized in that: The training module is specifically used for: Input the second image into the original model, extract the fillet image and remove the fillet image, and output the fillet extracted image and the third image; Calculating a first loss function value according to the fillet extraction image and the second image block, calculating a second loss function value according to the third image and one of the first image blocks, and calculating a third loss function value according to the third image block and a fourth image, wherein the fourth image is an image output after the third image is multiplied by a pixel value of the fillet extraction image; When the first loss function value, the second loss function value, and the third loss function value all meet preset conditions, the training is completed.

10. A device for removing a stripe image using a model, characterized in that: The model is trained by the device according to claim 6, wherein the device comprises: An acquisition module, used for acquiring a to-be-processed image of the photographed object containing a stripe and a reference image of the photographed object not containing a stripe; An output module, used for inputting a to-be-processed image of the photographed object containing a fillet and a reference image of the photographed object not containing a fillet into the model, and outputting a fillet image; An adjustment module, configured to divide the pixel value of the image to be processed containing the fillet by the pixel value of the fillet image, and output a fifth image; The adjustment module is further used to adjust the fifth image to a target image format, and output an image to be processed with the stripe image removed.

11. An electronic device, characterized in that: The method comprises a processor and a memory, wherein the memory stores a program or instruction that can be run on the processor, and when the program or instruction is executed by the processor, the steps of the method according to any one of claims 1 to 5 are implemented.

12. A readable storage medium, characterized in that: The readable storage medium stores a program or instruction, and when the program or instruction is executed by a processor, the steps of the method according to any one of claims 1 to 5 are implemented.

Citation Information

Patent Citations

  • Debanding Using A Novel Banding Metric

    US20230131228A1