Image processing method, dataset augmentation method, storage medium, and electronic device

By acquiring images from different data sets and performing white balance processing and Laplace pyramid fusion, the problems of lack and low quality of image data sets are solved, and the image fusion quality and model performance are improved.

CN114663320BActive Publication Date: 2025-07-11ALIBABA GROUP HOLDING LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202011527321.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-12-22
Publication Date
2025-07-11
Estimated Expiration
2040-12-22

AI Technical Summary

Technical Problem

In the prior art, due to the lack of image data sets or low image quality, the performance of neural network models is poor, especially in image processing, and the fusion quality is poor.

Method used

By acquiring images from different data sets and performing white balance processing, the target fusion image is generated using Laplace pyramid fusion technology to enhance the image data set and improve the image fusion quality.

Benefits of technology

It improves the efficiency and quality of image fusion processing, enriches the training data of the image segmentation model, and improves the segmentation performance of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114663320B_ABST
    Figure CN114663320B_ABST
Patent Text Reader

Abstract

The present application discloses an image processing method and apparatus, a method and apparatus for augmenting an image data set, a method for training an image segmentation model, an image processing method for a video conference, a method for processing video data in a video live broadcast, a computer storage medium, and an electronic device. The image processing method includes: obtaining a first image including target image elements from a first data set; obtaining a second image from a second data set; performing white balance processing on the first image and the second image to obtain a first white balance image corresponding to the first image and a second white balance image corresponding to the second image; performing image fusion processing on the first white balance image and the second white balance image to generate a target fusion image in which the target image elements are fused into the second image; which can improve the efficiency of image fusion processing and the image quality of the target fusion image, and avoid the defects of boundary formation of the fused image and color differentiation of the fused image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer image processing applications, and particularly relates to an image processing method and apparatus. The present application also relates to a method and apparatus for expanding an image data set, a method for training an image segmentation model, a method for processing background images in a video conference, and a method for processing video data in a video live broadcast, a computer storage medium, and an electronic device. Background Art

[0002] With the continuous development of artificial intelligence technology, deep learning has been widely applied in various fields. Deep learning can be understood as a branch of machine learning, referring to an artificial intelligence algorithm based on deep neural networks and a large amount of data. However, for deep learning to achieve better learning effects, it needs to be obtained through training with a large amount of sample data. Therefore, the amount of sample data also becomes one of the reasons affecting the performance of the neural network model.

[0003] In the prior art, in the face of a neural network model applied to image processing, due to the lack of an image data set or low-quality synthetic images, the model performance is poor. Summary of the Invention

[0004] The present application provides an image processing method to solve the problem of poor image fusion quality in the prior art.

[0005] The present application provides an image processing method, including:

[0006] Obtaining a first image including target image elements from a first data set;

[0007] Obtaining a second image from a second data set;

[0008] Performing white balance processing on the first image and the second image to obtain a first white balance image corresponding to the first image and a second white balance image corresponding to the second image;

[0009] Performing image fusion processing on the first white balance image and the second white balance image to generate a target fusion image in which the target image elements are fused into the second image.

[0010] In some embodiments, the performing white balance processing on the first image and the second image to obtain a first white balance image corresponding to the first image and a second white balance image corresponding to the second image includes:

[0011] Determining the gray mean value of the first image according to the RGB three-channel mean value of the first image;

[0012] Determine the grayscale mean value of the second image according to the RGB three-channel mean values of the second image;

[0013] Determine the RGB three-channel gain coefficients of the first image according to the grayscale mean value of the first image;

[0014] Determine the RGB three-channel gain system of the second image according to the grayscale mean value of the second image;

[0015] Adjust the pixels of the first image according to the original pixel RGB three-channel values of the first image and the RGB three-channel gain coefficients of the first image to obtain the first white balance image;

[0016] Adjust the pixels of the second image according to the original pixel RGB three-channel values of the second image and the RGB three-channel gain coefficients of the second image to obtain the second white balance image.

[0017] In some embodiments, the adjusting the pixels of the first image according to the original pixel RGB three-channel values of the first image and the RGB three-channel gain coefficients of the first image to obtain the first white balance image includes:

[0018] Obtain the RGB three-channel pixel values according to the product of the original pixel RGB three-channel values of the first image and the RGB three-channel gain coefficients of the first image;

[0019] Adjust the original pixel RGB three-channel values of the first image according to the RGB three-channel pixel values;

[0020] Determine the adjusted image as the first white balance image.

[0021] In some embodiments, the adjusting the pixels of the second image according to the original pixel RGB three-channel values of the second image and the RGB three-channel gain coefficients of the second image to obtain the second white balance image includes:

[0022] Obtain the RGB three-channel pixel values according to the product of the original pixel RGB three-channel values of the second image and the RGB three-channel gain coefficients of the second image;

[0023] Adjust the original pixel RGB three-channel values of the second image according to the RGB three-channel pixel values;

[0024] Determine the adjusted image as the second white balance image.

[0025] In some embodiments, the image fusion processing of the first white balance image and the second white balance image to generate a target fusion image in which the target image elements are fused into the second image includes:

[0026] Obtain a mask image of the first image;

[0027] Perform image fusion processing on the first white balance image, the second white balance image, and the mask image to generate a target fusion image that fuses the target image elements into the second image.

[0028] In some embodiments, the performing image fusion processing on the first white balance image, the second white balance image, and the mask image to generate a target fusion image that fuses the target image elements into the second image includes:

[0029] Perform image fusion processing on the first white balance image, the second white balance image, and the mask image by using Laplacian pyramid fusion to generate a target fusion image that fuses the target image elements into the second image.

[0030] In some embodiments, the performing image fusion processing on the first white balance image, the second white balance image, and the mask image by using Laplacian pyramid fusion to generate a target fusion image that fuses the target image elements into the second image includes:

[0031] Construct Gaussian pyramids for the first white balance image, the second white balance image, and the mask image respectively to obtain a first Gaussian pyramid of the first white balance image, a second Gaussian pyramid of the second white balance image, and a third Gaussian pyramid of the mask image;

[0032] Construct a first Laplacian pyramid corresponding to the first Gaussian pyramid, a second Laplacian pyramid corresponding to the second Gaussian pyramid, and a third Laplacian pyramid corresponding to the third Gaussian pyramid respectively according to the first Gaussian pyramid, the second Gaussian pyramid, and the third Gaussian pyramid;

[0033] Fuse the corresponding image layers in the first Laplacian pyramid, the second Laplacian pyramid, and the third Laplacian pyramid respectively to generate a new Laplacian image pyramid;

[0034] Reconstruct the new Laplacian image pyramid to generate the target fusion image.

[0035] In some embodiments, the reconstructing the new Laplacian image pyramid to generate the target fusion image includes:

[0036] Start from the top-layer image of the new Laplacian image pyramid and perform upsampling in sequence to generate a sampled image;

[0037] Determine the sampled image as the target fusion image.

[0038] In some embodiments, it further includes:

[0039] Perform data enhancement processing on the first white balance image and the second white balance image to obtain a first enhanced image and a second enhanced image;

[0040] The step of performing image fusion processing on the first white balance image and the second white balance image to generate a target fusion image in which the target image elements are fused into the second image includes:

[0041] Perform image fusion processing on the first enhanced image of the first white balance image and the second enhanced image of the second white balance image to generate a target fusion image in which the target image elements are fused into the second image.

[0042] This application also provides an image processing apparatus, including:

[0043] A first acquisition unit, configured to acquire a first image including target image elements from a first data set;

[0044] A second acquisition unit, configured to acquire a second image from a second data set;

[0045] A first processing unit, configured to perform white balance processing on the first image and the second image to obtain a first white balance image corresponding to the first image and a second white balance image corresponding to the second image;

[0046] A second processing unit, configured to perform image fusion processing on the first white balance image and the second white balance image to generate a target fusion image in which the target image elements are fused into the second image.

[0047] This application also provides a method for expanding an image data set, including:

[0048] Acquire a target fusion image generated according to the above image processing method;

[0049] Expand the image data set according to the target fusion image.

[0050] In some embodiments, the step of expanding the image data set according to the target fusion image includes:

[0051] Use the target fusion image as training data for an image segmentation model and store it in the image data set.

[0052] This application also provides an apparatus for expanding an image data set, including:

[0053] An acquisition unit, configured to acquire a target fusion image generated according to the above image processing method;

[0054] An expansion unit, configured to expand an image data set according to the target fusion image.

[0055] This application also provides a method for training an image segmentation model, including:

[0056] Acquire image data according to the image data set expanded by the above image data set expansion method;

[0057] Use the image data as training parameters of the image segmentation model and input them into the image segmentation model for training to obtain a trained image segmentation model; wherein, the trained image segmentation model can segment a foreground image or a background image from the image data input into the image segmentation model.

[0058] This application also provides a method for processing images in a video conference, including:

[0059] Acquire a video conference image in the video conference;

[0060] Input the video conference image into the image segmentation model for learning to identify a human body image in the video conference image; wherein, the image segmentation model is a model trained using the image data in the image data set obtained by the above image data set expansion method as training data;

[0061] Blur or replace the image outside the range of the human body image area.

[0062] This application also provides a method for processing video data in a video live broadcast, including:

[0063] Acquire a live broadcast screen image in the video live broadcast;

[0064] Input the live broadcast screen image into the image segmentation model for learning to identify a human body image in the live broadcast screen image; wherein, the image segmentation model is a model trained using the image data in the image data set obtained by the above image data set expansion method as training data;

[0065] Add preset information to the image outside the range of the human body image area.

[0066] This application also provides a computer storage medium for storing data generated by a network platform and a program for processing the data generated by the network platform;

[0067] When the program is read and executed, it performs the steps of the image processing method described above; or, it performs the steps of the method for augmenting the image data set described above; or, it performs the steps of the method for training the image segmentation model described above; or, it performs the steps of the image processing method for video conferencing described above; or, it performs the steps of the method for processing video data in video live streaming described above.

[0068] This application also provides an electronic device, including:

[0069] A processor;

[0070] A memory for storing a program for processing data generated by a network platform. When the program is read and executed by the processor, it performs the steps of the image processing method described above; it performs the steps of the method for augmenting the image data set described above; or, it performs the steps of the method for training the image segmentation model described above; or, it performs the steps of the image processing method for video conferencing described above; or, it performs the steps of the method for processing video data in video live streaming described above.

[0071] Compared with the prior art, this application has the following advantages:

[0072] An image processing method provided by this application can obtain a first image, a second image, and a mask image of the first image from different data sets. The first image can be an image including target image elements. Then, through white balance processing of the first image and the second image, the difference in image boundaries and hues during later fusion caused by different shooting environments can be reduced. The first white balance image and the second white balance image are subjected to image fusion processing to generate a target fusion image in which the target image elements are fused into the second image, thereby improving the efficiency of image fusion processing and the image quality of the target fusion image, and avoiding the defects of boundary formation and color differentiation of the fusion image.

[0073] A method for augmenting an image dataset provided by the present application can obtain a first image, a second image, and a mask image of the first image from different datasets. The first image can be an image including target image elements. Then, by performing white balance processing on the first image and the second image, the differences in image boundaries and hues during subsequent fusion due to different shooting environments can be reduced. The first white balance image and the second white balance image are subjected to image fusion processing to generate a target fusion image in which the target image elements are fused into the second image, thereby improving the efficiency of image fusion processing and the image quality of the target fusion image. Since the target fusion image is generated from images of different datasets, it can provide rich, high-quality, and realistic training data for an image segmentation model to augment the image dataset to meet the requirements of image segmentation model training and ensure the performance requirements of the model for image segmentation after training. Description of the Drawings

[0074] Figure 1 is a flowchart of an embodiment of an image processing method provided by the present application;

[0075] Figure 2 is a schematic structural diagram of an embodiment of an image processing apparatus provided by the present application;

[0076] Figure 3 is a flowchart of an embodiment of a method for augmenting an image dataset provided by the present application;

[0077] Figure 4 is a schematic principle structural diagram of an embodiment of a method for augmenting an image dataset provided by the present application;

[0078] Figure 5 is a schematic structural diagram of an embodiment of an apparatus for augmenting an image dataset provided by the present application;

[0079] Figure 6 is a flowchart of an embodiment of a method for training an image segmentation model provided by the present application;

[0080] Figure 7 is a flowchart of an embodiment of a method for processing images in a video conference provided by the present application;

[0081] Figure 8 is a flowchart of an embodiment of a method for processing video data in a video live broadcast provided by the present application;

[0082] Figure 9 is a schematic structural diagram of an embodiment of an electronic device provided by the present application. Detailed Embodiments

[0083] In the following description, numerous specific details are set forth in order to provide a thorough understanding of the present application. However, the present application can be implemented in many other ways different from those described herein, and those skilled in the art can make similar generalizations without departing from the spirit of the present application. Therefore, the present application is not limited by the specific implementations disclosed below.

[0084] The terms used in this application are for the purpose of describing specific embodiments only and are not intended to limit the present application. The descriptive methods used in this application and the appended claims, such as "a", "first", and "second", etc., are not intended to limit the quantity or the order, but are used to distinguish information of the same type from each other.

[0085] Combined with the above background art, the technical concept of the method for expanding an image data set provided by this application originates from the fact that in the prior art, when performing image segmentation, due to the lack of image data during the training of the image segmentation model, the application and the technical development of the image segmentation model are both hindered. To solve this problem, the technical concept provided by this application is that the image data in different data sets can be synthesized, and the synthesized image data is used to expand the data set of the image segmentation model, thereby enriching the data volume of the data set of the image segmentation model. For an image segmentation model or a neural network model, the quality of the image is also an important indicator affecting the performance of the model. Therefore, when fusing the image data of two data sets respectively, the problem of the image quality after fusion also needs to be considered. Therefore, this application provides a method for expanding an image data set, which can ensure the image quality of the expanded image data set while expanding the image data set, and avoid the performance degradation caused by the interference of the image quality on model training.

[0086] Based on the above content, first, an embodiment of an image processing method provided by this application will be described below. Please refer to Figure 1 and Figure 2 as shown, Figure 1 is a flowchart of an embodiment of an image processing method provided by this application. The embodiment of the image processing method includes:

[0087] Step S101: Obtain a first image including target image elements from a first data set;

[0088] The purpose of step S101 is to obtain a first image from the first data set.

[0089] The first dataset in step S101 can be an existing dataset in a deep learning task; the first image can be an image including target image elements, and the target image elements can be determined according to the image requirements of the image segmentation model. For example, for a portrait segmentation model, an image including portrait elements is required, so the target image elements can be a portrait (which can be the whole body or a partial image of the human body), that is to say, the target image elements can be understood as a human body image; if it is a model for segmenting animals or buildings, etc., the target image elements can be an animal image or a building image, etc. Of course, the target image elements can be partial information or all information of the image elements.

[0090] In this embodiment, a portrait image is taken as an example for illustration.

[0091] From the perspective of images, the target image elements can also be understood as the target foreground image. For a target image of a portrait, the foreground image should be portrait information. Therefore, the first image can be understood as a foreground image including portrait information.

[0092] Step S102: Obtain a second image from a second dataset;

[0093] The purpose of step S102 is to obtain a second image.

[0094] In this embodiment, the second image can be obtained from the second dataset. The second dataset and the first dataset are different image datasets. If the first image is a foreground image including a portrait, the second image can be understood as a background image.

[0095] The first dataset in step S101 and the second dataset in step S102 can be classification and recognition datasets, such as: COCO, ImageNet, PASCAL VOC, Label me, SUN, Caltech, Corel5k, etc.; face datasets, such as LFW (Labeled Faces in the Wild), VGG Face dataset, etc.; pedestrian detection datasets, etc.

[0096] In this embodiment, the first image can be obtained by using the above COCO dataset to obtain an image with portrait as the target image element; the second image can also be an image obtained from the above datasets or other datasets. From the perspective of image processing, the target image elements can be understood as a portrait, and the first image can be understood as a first image with the portrait as the foreground image. Correspondingly, the second image can be a background image.

[0097] Step S103: Perform white balance processing on the first image and the second image to obtain a first white balance image corresponding to the first image and a second white balance image corresponding to the second image;

[0098] The purpose of step S103 is to perform preprocessing on the first image and the second image. Because, whether the image data comes from the same dataset or different datasets, there are situations where the hues are inconsistent. Therefore, it is necessary to perform color correction on the first image and the second image to avoid performance problems of the neural network model due to hue issues.

[0099] In this embodiment, image color correction is mainly achieved through white balance processing. The specific implementation process may include:

[0100] Step S103-1: Determine the gray mean value of the first image according to the RGB three-channel mean values of the first image;

[0101] Step S103-2: Determine the gray mean value of the second image according to the RGB three-channel mean values of the second image;

[0102] Step S103-3: Determine the RGB three-channel gain coefficients of the first image according to the gray mean value of the first image;

[0103] Step S103-4: Determine the RGB three-channel gain system of the second image according to the gray mean value of the second image;

[0104] Step S103-5: Adjust the pixels of the first image according to the original pixel RGB three-channel values of the first image and the RGB three-channel gain coefficients of the first image to obtain the first white balance image;

[0105] Step S103-6: Adjust the pixels of the second image according to the original pixel RGB three-channel values of the second image and the RGB three-channel gain coefficients of the second image to obtain the second white balance image.

[0106] Among them, the RGB three-channel mean value of the first image in step S103-1 is The gray mean value of the first image and the gray mean value of the second image in step S103-2 can be calculated using the following formula:

[0107]

[0108] Among them, is the gray mean value of the RGB three channels.

[0109] The RGB three-channel gain coefficients in the steps S103-3 and S103-4 can be calculated using the following formula:

[0110]

[0111]

[0112]

[0113] where K r is the gain coefficient of the R channel, K g is the gain coefficient of the G channel, and K b is the gain coefficient of the B channel. Different images can be identified with different subscripts.

[0114] The specific implementation process of the step S103-5 can include:

[0115] Step S103-51: Obtain the RGB three-channel pixel values according to the product of the RGB three-channel pixel values of the first image and the RGB three-channel gain coefficients of the first image;

[0116] Step S103-52: Adjust the original RGB three-channel pixel values of the first image according to the RGB three-channel pixel values;

[0117] Step S103-53: Determine the adjusted image as the first white balance image.

[0118] The specific implementation process of the step S103-6 can include:

[0119] Step S103-61: Obtain the RGB three-channel pixel values according to the product of the RGB three-channel pixel values of the second image and the RGB three-channel gain coefficients of the second image;

[0120] Step S103-62: Adjust the original RGB three-channel pixel values of the second image according to the RGB three-channel pixel values;

[0121] Step S103-63: Determine the adjusted image as the second white balance image.

[0122] That is, the original RGB three-channel pixel values of the first image and the second image can be adjusted through the Von Kries diagonal model, for example, using the following formula:

[0123] P(R′) = P(R) × K r ;

[0124] P(G′) = P(G) × K g ;

[0125] P(B′)=P(B)×K b ;

[0126] Perform white balance processing on the first image and the second image through the above content to obtain a first white balance image and a second white balance image.

[0127] Step S104: Perform image fusion processing on the first white balance image and the second white balance image to generate a target fusion image that fuses the target image elements into the second image;

[0128] The purpose of step S104 is to generate a target fusion image that can fuse the target image elements into the second image. That is, fuse the foreground portrait image in the first image into the background image provided by the second image.

[0129] For generating the target fusion image in step S104, image fusion technology can be used. To achieve a better image fusion effect, while saving the operation time cost, it can avoid the problem of low quality of the fusion image caused by boundary differences in the fusion image.

[0130] In this embodiment, the specific implementation of step S104 may include:

[0131] Step S104-1: Obtain a mask image of the first image;

[0132] The mask image (Mask) is an image formed by blocking the target image elements in the first image, that is: the mask image is a binary image composed of 0 and 1. When the mask is applied in a certain function, the 1-value area is processed, and the masked 0-value area is not included in the calculation. In this embodiment, the mask image can have the same size as the first image. In the mask image, the pixel value of the target area is 1, and the pixel value of the non-target area is 0, so as to extract the pixel values of the target image elements from the first image according to the mask image. Therefore, after obtaining the first image, the mask image can be obtained by blocking the target image elements in the first image. There is no specific limitation on the timing of obtaining the mask image. For example: the first image can be masked when the first image is obtained to obtain the mask image, or the mask image can be obtained when generating the target fusion image.

[0133] Step S104-2: Perform image fusion processing on the first white balance image, the second white balance image, and the mask image by using the Laplacian pyramid fusion method to generate a target fusion image that fuses the target image elements into the second image.

[0134] Combined with Figure 4As shown, the specific implementation process of step S104-2 may include:

[0135] Step S104-21: Construct Gaussian pyramids for the first white balance image, the second white balance image, and the mask image respectively, to obtain the first Gaussian pyramid of the first white balance image, the second Gaussian pyramid of the second white balance image, and the third Gaussian pyramid of the mask image. The Gaussian pyramid is the most basic image pyramid.

[0136] The specific implementation process of step S104-21 may be to perform downsampling convolution operations on the first white balance image, the second white balance image, and the mask image respectively, so as to obtain the Gaussian pyramid images of the first white balance image, the second white balance image, and the mask image respectively. The following takes the first white balance image as an example to illustrate the process of obtaining the first Gaussian pyramid:

[0137] Take the first white balance image as the bottom layer image G0 of the Gaussian pyramid image, that is, the 0th layer image. Convolve the bottom layer image with a Gaussian kernel (n×n), and then downsample the convolved image (remove even rows and columns) to obtain the upper layer image G1 adjacent to and above the bottom layer image. Take image G1 as the input, and repeat the convolution and downsampling operations in G0 to obtain the upper layer image G2 above G1. Iterate multiple times, and the formed pyramid-shaped image data structure is the first Gaussian pyramid. It is expressed by the following formula:

[0138] G i =Down(G i-1 );

[0139] Where Down is the downsampling function, and G i represents the Gaussian image of the i-th layer. Based on the above content, it can be understood that downsampling can be achieved by discarding even rows and even columns in the image, so that the length and width of the image are each reduced by half, and the area is reduced by a quarter.

[0140] Similarly, the second white balance image can also be downsampled to obtain the second Gaussian pyramid, and the mask image can be downsampled to obtain the third Gaussian pyramid.

[0141] Step S104-22: Construct the first Laplacian pyramid corresponding to the first Gaussian pyramid, the second Laplacian pyramid corresponding to the second Gaussian pyramid, and the third Laplacian pyramid corresponding to the third Gaussian pyramid respectively according to the first Gaussian pyramid, the second Gaussian pyramid, and the third Gaussian pyramid;

[0142] The Laplacian pyramid in the step S104-22 can be understood as a pyramid of residual image structures. That is, in the process of operating the Gaussian pyramid, some high-frequency detail information will be lost during operations such as convolution and downsampling of the image, and the Laplacian pyramid describes the high-frequency detail information.

[0143] The specific implementation process of the step S104-22 can be to subtract the predicted image obtained by upsampling and Gaussian convolution of the image in the previous adjacent layer from each layer image of the Gaussian pyramid, and a series of difference images are obtained as the image structure of the Laplacian pyramid. In other words, first perform upsampling (enlarging the image, which can also be called image interpolation) on the image layer 3 in the Gaussian pyramid to obtain an image A with the same size as the image layer 2 after sampling, then perform downsampling (shrinking the image, that is, subsampling) on the image A to obtain a blurred A' with the same size as the image layer 3, and the image formed by the difference between A' and A is called the difference image (residual image) which is the Laplacian image. That is to say, the Laplacian pyramid is to record the difference between the upsampling and then downsampling and the pre-sampling at each level of the Gaussian pyramid. That is, the following formula:

[0144] L i =G i -U p (Down(G i ));

[0145] Wherein, L i represents the Laplacian pyramid image, U p represents upsampling, G i represents the Gaussian image of the i-th layer, and Down(G i ) can be understood as G i+1 , that is, the Gaussian image of the (i + 1)-th layer.

[0146] The formula can also be expressed as: L i =G i -U p (G i+1 );

[0147] Generally, the image at the top layer of the Gaussian pyramid is the same as the image at the top layer of the Laplacian.

[0148] Step S104-23: Fuse the corresponding image layers in the first Laplacian pyramid, the second Laplacian pyramid, and the third Laplacian pyramid respectively to generate a new Laplacian image pyramid;

[0149] The specific implementation process of step S104-23 can be to add the image of the first Laplacian pyramid and the image of the second Laplacian pyramid according to the third Laplacian pyramid, that is, the Laplacian pyramid of the mask image. The mask image is used to determine the fusion part, and in this embodiment, the portrait part in the first image is the mask part. The result after addition is a new Laplacian image pyramid.

[0150] Step S104-24: Reconstruct the new Laplacian image pyramid to generate the target fusion image.

[0151] The purpose of the reconstruction in step S104-24 is to construct the final target fusion image according to the new Laplacian pyramid. The construction process can include:

[0152] Step S104-241: Start from the top-layer image of the new Laplacian image pyramid and perform upsampling in sequence to generate a sampled image;

[0153] Step S104-242: Determine the sampled image as the target fusion image.

[0154] The above is the description of the implementation process of step S104. Since the Gaussian pyramid and the Laplacian pyramid belong to the prior art, the above description process is relatively brief.

[0155] In order to provide more image data to expand the image dataset, in this embodiment, data enhancement processing can also be performed on the first white balance image and the second white balance image. The purpose of performing data enhancement processing on the image data is to generate more image data, so as to improve the generalization ability during image fusion. Usually, the common methods of image data enhancement include: elastic deformation, image blurring, image rotation, adding noise, etc. Therefore, the specific implementation of step S104 can include:

[0156] Step S104-31: Perform image fusion processing on the first enhanced image of the first white balance image and the second enhanced image of the second white balance image to generate a target fusion image that fuses the target image elements into the second image.

[0157] Based on step S104-31, during image fusion processing, fusion processing can also be performed according to the image after data enhancement processing, so as to obtain more target fusion images.

[0158] The above is a description of an embodiment of an image processing method provided by the present application. Through the embodiment of the image processing method provided by the present application, two images with different backgrounds can be fused, that is, the target image elements in the first image are fused into the second image, so as to reduce the boundary difference in the fusion area and the color difference between the fused image and the background image during image fusion, and improve the quality of the fused image and the efficiency of image fusion processing.

[0159] The above is a specific description of an embodiment of an image processing method provided by the present application. In this embodiment, different image data are obtained from different data sets, and the two different image data are subjected to white balance processing to reduce the environmental difference between them. Then, image fusion is performed through an image pyramid to generate a target fused image, so that the target fused image can avoid problems such as low quality, obvious boundaries, and large color differences when the target image elements are fused into the background image due to the environmental difference of the original images.

[0160] Combined with the above content, corresponding to the embodiment of an image processing method provided above, the present application also provides an embodiment of an image processing apparatus. Please refer to Figure 2 , since the apparatus embodiment is basically similar to the method embodiment, the description is relatively simple. For related parts, please refer to the partial description of the method embodiment. The apparatus embodiment described below is only illustrative.

[0161] Please refer to Figure 2 as shown in Figure 2 is a schematic structural diagram of an embodiment of an apparatus for expanding an image data set provided by the present application. This apparatus embodiment includes:

[0162] A first acquisition unit 201, configured to acquire a first image including target image elements from a first data set;

[0163] The specific implementation process of the first acquisition unit 201 can refer to the specific content of step S101 above, and will not be repeated here.

[0164] A second acquisition unit 202, configured to acquire a second image from a second data set;

[0165] The specific implementation process of the second acquisition unit 202 can refer to step S102 above, and will not be repeated here.

[0166] A first processing unit 203, configured to perform white balance processing on the first image and the second image to obtain a first white balance image corresponding to the first image and a second white balance image corresponding to the second image;

[0167] The first processing unit includes: a first determination subunit, a second determination subunit, a third determination subunit, a fourth determination subunit, a first acquisition subunit, and a second acquisition subunit.

[0168] The first determination subunit is configured to determine the grayscale mean value of the first image according to the RGB three-channel mean value of the first image.

[0169] The second determination subunit is configured to determine the grayscale mean value of the second image according to the RGB three-channel mean value of the second image.

[0170] The third determination subunit is configured to determine the RGB three-channel gain coefficient of the first image according to the grayscale mean value of the first image in the first determination subunit.

[0171] The fourth determination subunit is configured to determine the RGB three-channel gain system of the second image according to the grayscale mean value of the second image in the first determination subunit.

[0172] The first acquisition subunit is configured to adjust the pixels of the first image according to the original pixel RGB three-channel values of the first image and the RGB three-channel gain coefficient of the first image in the third determination subunit, so as to obtain the first white balance image. The first acquisition subunit includes: a calculation subunit, an adjustment subunit, and a determination subunit; the calculation subunit is configured to obtain the RGB three-channel pixel values according to the product of the pixel RGB three-channel values of the first image and the RGB three-channel gain coefficient of the first image; the adjustment subunit is configured to adjust the original pixel RGB three-channel values of the first image according to the RGB three-channel pixel values obtained by the calculation subunit; the determination subunit is configured to determine the image adjusted by the adjustment subunit as the first white balance image.

[0173] The second acquisition subunit is configured to adjust the pixels of the second image according to the original pixel RGB three-channel values of the second image and the RGB three-channel gain coefficient of the second image in the fourth determination subunit, so as to obtain the second white balance image. The second acquisition subunit includes: a calculation subunit, an adjustment subunit, and a determination subunit; the calculation subunit is configured to obtain the RGB three-channel pixel values according to the product of the pixel RGB three-channel values of the second image and the RGB three-channel gain coefficient of the second image; the adjustment subunit is configured to adjust the original pixel RGB three-channel values of the second image according to the RGB three-channel pixel values obtained by the calculation subunit; the determination subunit is configured to determine the image adjusted by the adjustment subunit as the second white balance image.

[0174] For the detailed technical content involved in the specific implementation process of the first processing unit 203, reference may be made to the above step S103, which will not be repeated here.

[0175] A second processing unit 204, configured to perform image fusion processing on the first white balance image and the second white balance image to generate a target fusion image in which the target image elements are fused into the second image;

[0176] The second processing unit 204 may specifically include: an acquisition subunit and a processing subunit;

[0177] The acquisition subunit is configured to acquire a mask image of the first image; for specific reference to the content of step S104-1, which will not be repeated here.

[0178] The processing subunit is configured to perform image fusion processing on the first white balance image, the second white balance image, and the mask image by using a Laplacian pyramid fusion method to generate a target fusion image in which the target image elements are fused into the second image. Specifically, it may include: a first construction subunit, a second construction subunit, a fusion subunit, and a reconstruction subunit.

[0179] The first construction subunit is configured to respectively construct Gaussian pyramids for the first white balance image, the second white balance image, and the mask image to obtain a first Gaussian pyramid of the first white balance image, a second Gaussian pyramid of the second white balance image, and a third Gaussian pyramid of the mask image;

[0180] The second construction subunit is configured to respectively construct a first Laplacian pyramid corresponding to the first Gaussian pyramid, a second Laplacian pyramid corresponding to the second Gaussian pyramid, and a third Laplacian pyramid corresponding to the third Gaussian pyramid according to the first Gaussian pyramid, the second Gaussian pyramid, and the third Gaussian pyramid;

[0181] The fusion subunit is configured to respectively fuse the corresponding image layers in the first Laplacian pyramid, the second Laplacian pyramid, and the third Laplacian pyramid to generate a new Laplacian image pyramid;

[0182] The reconstruction subunit is configured to reconstruct the new Laplacian image pyramid to generate the target fusion image.

[0183] The reconstruction subunit may include: an upsampling subunit and a determination subunit.

[0184] The upsampling subunit is configured to perform upsampling on the top-layer image of the new Laplacian image pyramid in sequence to generate a sampled image;

[0185] The determining subunit is configured to determine the sampled image as the target fusion image.

[0186] In order to provide more image data to expand the image dataset, in this embodiment, it may further include: an enhancement unit configured to perform data enhancement processing on the first white balance image and the second white balance image to obtain a first enhanced image and a second enhanced image.

[0187] Specifically, the second processing unit may perform image fusion processing according to the first enhanced image of the first white balance image, the second enhanced image of the second white balance image, and the mask image in the enhancement unit, and generate a target fusion image in which the target image elements are fused into the second image.

[0188] Regarding the detailed technical content involved in the specific implementation process of the second processing unit 204, reference may be made to the above step S104, which will not be repeated here.

[0189] To increase the amount of image data in the image dataset and provide a large number of training sample data for the training of the image segmentation model to improve the training performance of the portrait segmentation model. Combining the above content, the present application further provides a method for expanding an image dataset. Please refer to Figure 3 and Figure 4 shown, Figure 3 is a flowchart of an embodiment of a method for expanding an image dataset provided by the present application; Figure 4 is a schematic diagram of the principle structure of an embodiment of a method for expanding an image dataset provided by the present application.

[0190] As Figure 3 shown, the embodiment of the method for expanding an image dataset of the present application includes:

[0191] Step S301: Obtain a target fusion image generated according to the above image processing method;

[0192] Regarding the specific implementation process of step S301, reference may be made to the above steps S101 - step S104, which will not be repeated here.

[0193] Step S302: Expand the image dataset according to the target fusion image.

[0194] The specific implementation process of step S302 may be to use the target fusion image as training data for the image segmentation model and store it in the image dataset.

[0195] According to the above content, in this embodiment, the portrait is used for the first image. Therefore, the image segmentation model in step S302 may be a portrait segmentation model.

[0196] The described image segmentation model can be applied to relevant application scenarios such as video conferencing, video live streaming, and portrait recognition. Through the embodiment of the method for expanding the image dataset provided by this application, different image data can be obtained from different datasets. After white balance processing of two different image data to narrow the environmental differences between them, image fusion is then performed through an image pyramid to generate a target fusion image. The target fusion image is provided as the image expansion data for the image dataset required by the image segmentation model. Since the target fusion image can avoid problems such as low quality, obvious boundaries, and large color differences in which the target image elements are fused into the background image due to the environmental differences of the original images, it can thus avoid the problem that the performance of the image segmentation model deteriorates due to the use of low-quality fused images, and improve the segmentation efficiency and quality of the image segmentation model.

[0197] The above is a specific description of an embodiment of the method for expanding an image dataset provided by this application. Corresponding to the foregoing embodiment of the method for expanding an image dataset, this application also provides an embodiment of an apparatus for expanding an image dataset. Please refer to Figure 5 , since the apparatus embodiment is basically similar to the method embodiment, the description is relatively simple. For related parts, refer to the partial description of the method embodiment. The apparatus embodiment described below is merely illustrative.

[0198] As Figure 5 shown, Figure 5 is a structural schematic diagram of an embodiment of an apparatus for expanding an image dataset provided by this application. This embodiment includes:

[0199] An acquisition unit 501, configured to acquire a target fusion image generated according to the above-mentioned image processing method; the specific content of the acquisition unit 501 refers to the specific description of steps S101 to S104 above, and will not be repeated here.

[0200] An expansion unit 502, configured to expand the image dataset according to the target fusion image.

[0201] The expansion unit 502 may specifically include a storage subunit, configured to store the target fusion image as training data of the image segmentation model into the image dataset. The specific implementation process of the expansion unit 501 can refer to the content of step S302 above.

[0202] The above is a description of an embodiment of an apparatus for expanding an image dataset provided by this application. The understanding of this apparatus embodiment can refer to the description of the corresponding method embodiment above, and will not be repeated here.

[0203] Combining the above content, this application also provides a method for training an image segmentation model. As Figure 6As shown Figure 6 Figure 6 is a flowchart of an embodiment of a method for training an image segmentation model provided by this application. This embodiment of the training method includes:

[0204] Step S601: Obtain image data according to the augmented image dataset in the above-provided method for augmenting the image dataset;

[0205] The purpose of step S601 is to obtain image data from the image dataset. The image dataset is the processed target fusion image obtained based on the image processing method provided in the above steps S101 to S104. The image dataset is augmented according to steps S301 to S302 through the target fusion image. The augmented image dataset has a larger amount of image data compared to the image dataset before augmentation. The augmented image dataset can provide a large number of image training parameters for the image segmentation model.

[0206] Step S602: Use the image data as training parameters of the image segmentation model and input them into the image segmentation model for training to obtain a trained image segmentation model; wherein, the trained image segmentation model can segment the foreground image or the background image from the image data to be segmented input into the image segmentation model.

[0207] In this embodiment, the foreground image can be a human body image, that is, the image segmentation model can segment the human body image or the background image, or both the human body image and the background image from the image data to be segmented.

[0208] Combined with the above content, this application also provides an image processing method for video conferencing, as Figure 7 shown. This embodiment of the image processing method for video conferencing includes:

[0209] Step S701: Obtain the video conferencing image in the video conference;

[0210] Step S702: Input the video conferencing image into the image segmentation model for learning to identify the human body image in the video conferencing image; wherein, the image segmentation model is a model trained using the image data in the image dataset obtained by using the above-provided method for augmenting the image dataset as training data;

[0211] Step S703: Blur or replace the image outside the range of the human body image area.

[0212] Combined with the above content, this application also provides a method for processing video data in video live streaming, as Figure 8 shown. This embodiment of the method for processing video data in video live streaming includes:

[0213] Step S801: Obtain the live video image in the video live stream;

[0214] Step S802: Input the live video image into an image segmentation model for learning to identify the human body image in the live video image; wherein, the image segmentation model is a model trained with the image data in the image dataset obtained by using the above-provided method for expanding the image dataset as training data;

[0215] Step S803: Add preset information to the image outside the range of the human body image area. The preset information can be displayed in a specified area of the live video image or in the area outside the human body image area of the live video image. And the preset information can change with the change of the range of the human body image area, that is to say, the image segmentation model can learn and identify the input live video image in real time and determine the area range of the human body image in real time.

[0216] Based on the above content, the present application also provides a computer storage medium for storing the data generated by the network platform and the program for processing the data generated by the corresponding network platform;

[0217] When the program is read and executed, it executes the steps of the image processing method provided by the present application as described above; or, when the program is read and executed, it executes the steps of the method for expanding the dataset provided by the present application as described above; or, when the program is read and executed, it executes the steps of the method for training the image segmentation model provided by the present application as described above; or, when the program is read and executed, it executes the steps of the method for processing the image in the video conference provided by the present application as described above; or, when the program is read and executed, it executes the steps of the method for processing the video data in the video live stream provided by the present application as described above.

[0218] As Figure 9 shown, Figure 9 is a schematic structural diagram of an embodiment of an electronic device provided by the present application. The embodiment of the electronic device includes: a processor 901 and a memory 902;

[0219] The memory 902 is used to store the program for processing the data generated by the network platform. When the program is read and executed by the processor 901, it executes the steps of the image processing method provided by the present application as described above; or, when the program is read and executed, it executes the steps of the method for expanding the image dataset provided by the present application as described above; or, when the program is read and executed, it executes the steps of the method for training the image segmentation model provided by the present application as described above; or, when the program is read and executed, it executes the steps of the method for processing the image in the video conference provided by the present application as described above; or, when the program is read and executed, it executes the steps of the method for processing the video data in the video live stream provided by the present application as described above.

[0220] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory.

[0221] The memory may include non-permanent memory in the form of computer-readable media, random access memory (RAM) and / or non-volatile memory such as read-only memory (ROM) or flash memory (flash RAM). Memory is an example of computer-readable media.

[0222] 1. Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can store information by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile discs (DVD) or other optical storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transitory medium that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory media such as modulated data signals and carrier waves.

[0223] 2. Those skilled in the art will appreciate that the embodiments of the present application may be provided as a method, a system, or a computer program product. Accordingly, the present application may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.

[0224] Although the present application is disclosed above in preferred embodiments, it is not intended to limit the present application. Any person skilled in the art can make possible changes and modifications without departing from the spirit and scope of the present application. Therefore, the protection scope of the present application should be determined by the scope defined by the claims of the present application.

Claims

1. An image processing method, characterized in that, Including: Obtain a first image including target image elements from a first data set; Obtain a second image from a second data set; Perform white balance processing on the first image and the second image to obtain a first white balance image corresponding to the first image and a second white balance image corresponding to the second image; Perform image fusion processing on the first white balance image and the second white balance image to generate a target fusion image in which the target image elements are fused into the second image, including: obtaining a mask image of the first image; Perform image fusion processing on the first white balance image, the second white balance image, and the mask image by using a Laplacian pyramid fusion method to generate a target fusion image in which the target image elements are fused into the second image.

2. The image processing method according to claim 1, wherein The performing white balance processing on the first image and the second image to obtain a first white balance image corresponding to the first image and a second white balance image corresponding to the second image includes: Determine the gray mean value of the first image according to the RGB three-channel mean value of the first image; Determine the gray mean value of the second image according to the RGB three-channel mean value of the second image; Determine the RGB three-channel gain coefficients of the first image according to the gray mean value of the first image; Determine the RGB three-channel gain system of the second image according to the gray mean value of the second image; Adjust the pixels of the first image according to the original pixel RGB three-channel values of the first image and the RGB three-channel gain coefficients of the first image to obtain the first white balance image; Adjust the pixels of the second image according to the original pixel RGB three-channel values of the second image and the RGB three-channel gain coefficients of the second image to obtain the second white balance image.

3. The image processing method according to claim 2, wherein The adjusting the pixels of the first image according to the original pixel RGB three-channel values of the first image and the RGB three-channel gain coefficients of the first image to obtain the first white balance image includes: Obtain RGB three-channel pixel values according to the product of the original pixel RGB three-channel values of the first image and the RGB three-channel gain coefficients of the first image; Adjust the original pixel RGB three-channel values of the first image according to the RGB three-channel pixel values; Determine the adjusted image as the first white balance image.

4. The image processing method according to claim 2, characterized in that, The adjusting the pixels of the second image according to the original pixel RGB three-channel values of the second image and the RGB three-channel gain coefficients of the second image to obtain the second white balance image includes: Obtain RGB three-channel pixel values according to the product of the original pixel RGB three-channel values of the second image and the RGB three-channel gain coefficients of the second image; Adjust the original pixel RGB three-channel values of the second image according to the RGB three-channel pixel values; Determine the adjusted image as the second white balance image.

5. The image processing method according to claim 1, wherein Performing image fusion processing on the first white balance image, the second white balance image, and the mask image by using Laplacian pyramid fusion to generate a target fusion image that fuses the target image elements into the second image, includes: Constructing Gaussian pyramids for the first white balance image, the second white balance image, and the mask image respectively to obtain a first Gaussian pyramid of the first white balance image, a second Gaussian pyramid of the second white balance image, and a third Gaussian pyramid of the mask image; Constructing a first Laplacian pyramid corresponding to the first Gaussian pyramid, a second Laplacian pyramid corresponding to the second Gaussian pyramid, and a third Laplacian pyramid corresponding to the third Gaussian pyramid according to the first Gaussian pyramid, the second Gaussian pyramid, and the third Gaussian pyramid respectively; Fusing the corresponding image layers in the first Laplacian pyramid, the second Laplacian pyramid, and the third Laplacian pyramid respectively to generate a new Laplacian image pyramid; Reconstructing the new Laplacian image pyramid to generate the target fusion image.

6. The image processing method according to claim 5, wherein The reconstructing the new Laplacian image pyramid to generate the target fusion image includes: Performing upsampling on the top-layer image of the new Laplacian image pyramid in sequence to generate a sampled image; and determining the sampled image as the target fusion image.

7. The image processing method according to claim 1, wherein It also includes: Performing data enhancement processing on the first white balance image and the second white balance image to obtain a first enhanced image and a second enhanced image; The performing image fusion processing on the first white balance image and the second white balance image to generate a target fusion image that fuses the target image elements into the second image includes: Performing image fusion processing on the first enhanced image of the first white balance image and the second enhanced image of the second white balance image to generate a target fusion image that fuses the target image elements into the second image.

8. An image processing apparatus, characterized in that, It includes: A first acquisition unit for acquiring a first image including target image elements from a first data set; A second acquisition unit for acquiring a second image from a second data set; A first processing unit for performing white balance processing on the first image and the second image to obtain a first white balance image corresponding to the first image and a second white balance image corresponding to the second image; A second processing unit for performing image fusion processing on the first white balance image and the second white balance image to generate a target fusion image that fuses the target image elements into the second image, including: acquiring a mask image of the first image; Performing image fusion processing on the first white balance image, the second white balance image, and the mask image by using Laplacian pyramid fusion to generate a target fusion image that fuses the target image elements into the second image.

9. A method for augmenting an image dataset, characterized in that, It includes: Acquiring a target fusion image generated by the image processing method according to any one of claims 1 to 7 above; Expanding the image data set according to the target fusion image.

10. The method for augmenting an image data set according to claim 9, wherein The expanding of the image dataset according to the target fusion image includes: Storing the target fusion image as training data of an image segmentation model into the image dataset.

11. An apparatus for augmenting an image data set, characterized in that, Including: An obtaining unit, configured to obtain a target fusion image generated by the image processing method according to any one of claims 1 to 6 above; An expanding unit, configured to expand the image dataset according to the target fusion image.

12. A training method for an image segmentation model, characterized in that, Including: Obtaining image data from the image dataset expanded by the method for expanding an image dataset according to any one of claims 9 to 10 above; Using the image data as training parameters of an image segmentation model, and inputting same into the image segmentation model for training to obtain a trained image segmentation model; wherein, the trained image segmentation model can segment a foreground image or a background image from the image data input into the image segmentation model.

13. An image processing method for video conferencing, characterized in that, Including: Obtaining a video conference image in a video conference; Inputting the video conference image into an image segmentation model for learning to identify a human body image in the video conference image; wherein, the image segmentation model is a model trained using the image data in the image dataset obtained by the method for expanding an image dataset according to any one of claims 9 to 10 above as training data; Blurring or replacing the image outside the range of the human body image area.

14. A method for processing video data in a video live broadcast, characterized in that, Including: Obtaining a live broadcast image in a video live broadcast; Inputting the live broadcast image into an image segmentation model for learning to identify a human body image in the live broadcast image; wherein, the image segmentation model is a model trained using the image data in the image dataset obtained by the method for expanding an image dataset according to any one of claims 9 to 10 above as training data; Adding preset information to the image outside the range of the human body image area.

15. A computer storage medium, configured to store data generated by a network platform and a program for processing the data generated by the network platform; When the program is read and executed, it executes the steps of the image processing method according to any one of claims 1 to 7; or executes the steps of the method for expanding an image dataset according to any one of claims 9 to 10; or executes the steps of the method for training an image segmentation model according to claim 12; or executes the steps of the method for processing video data in a video conference according to claim 13; or executes the steps of the method for processing video data in a video live broadcast according to claim 14.

16. An electronic device, including: A processor; A memory for storing a program for processing data generated by a network platform. When the program is read and executed by the processor, it performs the steps of the image processing method according to any one of claims 1 to 7; performs the steps of the method for augmenting an image data set according to any one of claims 9 to 10; or, performs the steps of the method for training an image segmentation model according to claim 12; or, performs the steps of the method for processing image data in a video conference according to claim 13; or, performs the steps of the method for processing video data in a video live broadcast according to claim 14.

Citation Information

Patent Citations

  • Distortion processing method and terminal

    CN107071274A

  • Sample generation method and device, electronic equipment and computer readable storage medium

    CN110889824A

  • Training set acquisition method and electric power facility detection method and device

    CN111428753A