Image processing, model training, body beautifying method, device and storage medium
By segmenting and deforming the image, and combining it with a trained completion model to complete the background, the background distortion problem caused by the body shaping effect is solved, and the overall image quality is improved.
Patent Information
- Application Number
- CN202210625411.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-02
- Publication Date
- 2026-01-23
- Estimated Expiration
- 2042-06-02
AI Technical Summary
Existing anti-background distortion technology performs poorly when dealing with complex backgrounds and cannot effectively correct background distortion caused by body shaping effects.
By segmenting the image of the target object, separating the background and foreground, performing deformation processing, and then using a trained completion model to fill in the missing areas of the foreground image, the image is synthesized to improve the realism and plausibility of the background image.
It effectively reduces background distortion caused by deformation processing, improves the overall image effect, and ensures the integrity and realism of the background area.
Smart Images

Figure CN115082384B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of vision, in particular to an image processing method, a model training method, a body beautifying method, equipment and a storage medium. BACKGROUND
[0002] In various short video and live broadcast products, body beautifying effects are common, such as slim waist effects, slim neck effects, swan neck effects, slim arm effects, leg length effects, etc. After implementing the body beautifying effect, the background will become distorted, which will make the overall image look very unrealistic and cause background distortion.
[0003] The existing anti-background distortion technology detects straight lines in the background, and then eliminates the straight line distortion caused by the body beautifying effect based on the detection result.
[0004] The existing solution can only correct the distortion of straight lines, and the correction effect is poor for complex backgrounds. SUMMARY
[0005] In view of the above problems, the present application is proposed to provide an image processing method, a model training method, a body beautifying method, equipment and a storage medium to solve the above problems or at least partially solve the above problems.
[0006] Therefore, in an embodiment of the present application, an image processing method is provided, which includes:
[0007] Segmenting a target object image to be processed to obtain a background image and a foreground image corresponding to the target object;
[0008] Performing morphing processing on the foreground image to obtain a morphed foreground image;
[0009] Using a trained first completion model to complete a blank area in the background image corresponding to the foreground image to obtain a completed background image;
[0010] Synthesizing the morphed foreground image and the completed background image to obtain a target image.
[0011] In another embodiment of the present application, a model training method is provided, which includes:
[0012] According to the contour of a sample object, a corresponding first contour region is determined in a first image;
[0013] Deleting the pixels of the first contour region in the first image to obtain a training sample;
[0014] Training a first completion model according to the training sample and the first image;
[0015] The first completion model is configured to complete the vacancy area corresponding to the foreground image in the background image to obtain a completed background image; the foreground image and the background image are obtained by segmenting a to-be-processed image; the foreground image contains a target object; the target object and the sample object belong to the same category.
[0016] In another embodiment of the present application, a body beautifying method is provided, which is suitable for a live broadcast client, and the method comprises the following steps.
[0017] In a live broadcast process, a live broadcast image collected by an anchor is obtained.
[0018] The live broadcast image is segmented to obtain a background image and a foreground image corresponding to the anchor.
[0019] The foreground image is subjected to body beautifying processing to obtain a body-beautified foreground image.
[0020] A first trained completion model is used to complete a vacancy area corresponding to the foreground image in the background image to obtain a completed background image.
[0021] The body-beautified foreground image and the completed background image are subjected to synthesis processing to obtain a target live broadcast image.
[0022] The target live broadcast image is sent to a live broadcast server.
[0023] In another embodiment of the present application, a body beautifying method is provided, which is suitable for a live broadcast client, and the live broadcast client is built-in with a software development kit provided by a body beautifying service provider; the live broadcast client is provided by a live broadcast service provider; the method comprises the following steps.
[0024] In a live broadcast process, a live broadcast image collected by an anchor is obtained.
[0025] The software development kit is called to realize the following steps: segmenting the live broadcast image to obtain a background image and a foreground image corresponding to the anchor; subjecting the foreground image to body beautifying processing to obtain a body-beautified foreground image; using a first trained completion model to complete a vacancy area corresponding to the foreground image in the background image to obtain a completed background image; and subjecting the body-beautified foreground image and the completed background image to synthesis processing to obtain a target live broadcast image.
[0026] The target live broadcast image is sent to a live broadcast server provided by the live broadcast service provider.
[0027] In another embodiment of the present application, an electronic device is provided. The electronic device comprises a memory and a processor, wherein
[0028] The memory is configured to store a program.
[0029] The processor is coupled to the memory and is configured to execute the program stored in the memory to implement the method in any of the preceding embodiments.
[0030] In another embodiment of the present application, a computer readable storage medium storing a computer program is provided, and the computer program can implement the method in any of the preceding embodiments when executed by a computer.
[0031] In the technical scheme provided by the embodiments of the present application, the foreground image and the background image are separated, the target object in the foreground image is processed for deformation separately, and the deformed foreground image is obtained. In this way, the background distortion problem caused by the deformation processing is avoided. In addition, the first completed model trained is used to complete the vacancy area corresponding to the foreground image in the background image to obtain the completed background image. Not only the hole problem of the background area in the synthesized image is avoided, but also the authenticity and rationality of the completed background image are improved. It can be seen that the technical scheme provided by the embodiments of the present application can reduce the background distortion while not affecting the purpose of the target object deformation, and improve the overall effect of the processed image. BRIEF DESCRIPTION OF DRAWINGS
[0032] In order to more clearly illustrate the technical scheme in the embodiments of the present application or the prior art, the drawings needed in the embodiment or the prior art description will be briefly introduced. Obviously, the drawings in the following description are some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.
[0033] Figure 1 Flowchart of the image processing method provided by an embodiment of the present application;
[0034] Figure 2 Flowchart of the model training method provided by an embodiment of the present application;
[0035] Figure 3 Flowchart of the body beautifying method provided by an embodiment of the present application;
[0036] Figure 4 Flowchart of the body beautifying method provided by another embodiment of the present application;
[0037] Figure 5 Example diagram of the image processing method provided by an embodiment of the present application;
[0038] Figure 6 Structural block diagram of the image processing device provided by an embodiment of the present application;
[0039] Figure 7 A structural block diagram of a model training apparatus provided by an embodiment of the present application is shown.
[0040] Figure 8 A structural block diagram of an electronic device provided by an embodiment of the present application is shown. DETAILED DESCRIPTION
[0041] In order to enable persons skilled in the art to better understand the scheme of the present application, the technical scheme in the embodiments of the present application will be described clearly and completely below according to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by persons skilled in the art without creative labor fall within the scope of protection of the present application.
[0042] In addition, in some of the processes described in the specification, claims, and accompanying drawings of the present application, a plurality of operations appearing in a specific order are included, which can be executed in the order appearing in the present text or in parallel. The serial numbers of the operations, such as 101, 102, etc., are only used to distinguish different operations, and the serial numbers themselves do not represent any execution order. In addition, these processes can include more or fewer operations, and the operations can be executed in sequence or in parallel. It should be noted that the descriptions of "first", "second", etc. in the present text are used to distinguish different messages, devices, modules, etc., and do not represent the order of precedence. In addition, "first" and "second" are different types.
[0043] Figure 1 A flowchart of an image processing method provided by an embodiment of the present application is shown. The execution subject of the method can be a client or a server. The client can be a hardware with an embedded program integrated on a terminal, an application software installed in the terminal, a tool software embedded in the operating system of the terminal, etc., which are not limited by the embodiments of the present application. The terminal can be any terminal device, including a mobile phone, a tablet computer, a vehicle-mounted terminal device, etc. The server can be a commonly used server, a cloud server, a virtual server, etc., which are not limited by the embodiments of the present application. As shown in the figure, the method comprises: Figure 1
[0044] 101, segmenting a to-be-processed image of a target object to obtain a background image and a foreground image corresponding to the target object.
[0045] 102, performing morphing processing on the foreground image to obtain a morphed foreground image.
[0046] 103. Completing the blank area corresponding to the foreground image in the background image by using the trained first completion model to obtain a completed background image.
[0047] 104. Synthesizing the deformed foreground image and the completed background image to obtain a target image.
[0048] In the above 101, the target object is included in the to-be-processed image. In actual applications, the target object can refer to different contents, and the target object can be a person, an animal, a plant, an article, etc. The target object belongs to the foreground in the to-be-processed image, and the area in the to-be-processed image except the target object is the background.
[0049] Note: The images mentioned in the embodiments of the present application can be RGB images.
[0050] In an example, an editing interface can be provided to a user by a client, and the to-be-processed image is displayed on the editing interface. The target object contour outlining operation performed by the user on the to-be-processed image in the editing interface is received, the contour of the target object is determined according to the target object contour outlining operation, and the to-be-processed image of the target object is segmented according to the contour of the target object to obtain a background image and a foreground image corresponding to the target object.
[0051] In another example, an image segmentation model can be used to segment the to-be-processed image of the target object to obtain a background image and a foreground image corresponding to the target object. Specifically, the step of segmenting the to-be-processed image of the target object to obtain a background image and a foreground image corresponding to the target object in the above 101 can be implemented as follows:
[0052] 1011. The target object in the to-be-processed image is segmented by using a trained image segmentation model to obtain a mask image corresponding to the target object.
[0053] 1012. The to-be-processed image is segmented according to the mask image to obtain the foreground image and the background image.
[0054] In the above 1011, the semantic segmentation process is a process of determining whether each pixel in the to-be-processed image belongs to the target object or the background.
[0055] The mask image can be a binary image. For example, the mask image includes a first region corresponding to the target object in the to-be-processed image and a second region corresponding to the background in the to-be-processed image, the pixel values of the first region can all be 1, and the pixel values of the second region can all be 0.
[0056] In 1022, all pixels of the target object are extracted from the to-be-processed image according to the mask image to obtain a foreground image; and all pixels of the target object are deleted from the to-be-processed image according to the mask image to obtain a background image.
[0057] The image segmentation model has been trained to a first preset convergence state. The first preset convergence state can be set according to actual needs, and embodiments of the present application do not make specific limitations thereto.
[0058] The image segmentation model can be a machine learning model, for example, a deep learning model. The internal implementation and training process of the image segmentation model can refer to prior art, and will not be described in detail here.
[0059] In 102, the foreground image is morphed to obtain a morphed foreground image, which is essentially morphing the target object in the foreground image to obtain a morphed foreground image.
[0060] The morphing process can include overall shrinking, overall enlargement, local shrinking, local enlargement, and the like.
[0061] In general, in the beauty scene, the morphing process can be a beauty treatment, which can include slimming the waist, slimming the neck, slimming the neck, slimming the arms, lengthening the legs, and the like.
[0062] In actual application, a morphing parameter corresponding to the target object can be obtained, and the target object in the foreground image is morphed according to the morphing parameter. In an example, the morphing parameter can be user-defined or default. The morphing parameter can include an adjustment direction and an adjustment intensity for adjusting the components of the target object.
[0063] In an example, the morphing process of the foreground image in 102 to obtain a morphed foreground image can be implemented by the following steps:
[0064] 1021, obtaining a preset beauty parameter.
[0065] 1022, according to the beauty parameter, the human body part of the person in the foreground image is beautified to obtain the morphed foreground image.
[0066] The body beautifying parameter can include an adjustment direction and an adjustment intensity for adjusting a body part of the person. For example, the adjustment direction of the calf in the morphing parameter can be to reduce the calf leg circumference, and the corresponding adjustment intensity can be 0.1 times, that is, the body beautifying parameter is to reduce the calf leg circumference by 0.1 times. For another example, the adjustment direction of the waist in the morphing parameter can be to reduce the waist circumference, and the corresponding adjustment intensity can be 0.2 times, that is, the body beautifying parameter is to reduce the waist circumference by 0.2 times.
[0067] In an example, a reference three-dimensional model corresponding to the target object can be acquired; and the morphing parameter corresponding to the target object is generated according to the reference three-dimensional model. The reference three-dimensional model can be designed according to actual needs.
[0068] In 103, the first completion model is trained to a second preset convergence state. The second preset convergence state can be set according to actual needs, and embodiments of the present application do not make specific limitations thereto.
[0069] In an example, a first image can be acquired; pixels in a partial region in the first image are deleted to obtain a training sample; the first image is used as a training label of the training sample; and the first completion model is trained according to the training sample and the training label thereof. In actual application, the number of the first images is multiple, and the position and shape of the partial region can be determined randomly.
[0070] The background image can be input into the trained first completion model to obtain a completed background image output by the first completion model.
[0071] The first completion model can be a machine learning model, for example, a deep learning model.
[0072] In 104, the morphed foreground image is superimposed on the completed background image to cover a partial region on the completed background image, to obtain a target image. The shape and area of the partial region are the same as the shape and area of the target object after morphing in the morphed foreground image.
[0073] In the technical scheme provided by the embodiments of the present application, the foreground image and the background image are separated, the target object in the foreground image is morphed separately to obtain a morphed foreground image, so that the background distortion problem caused by morphing is avoided. Moreover, the first completion model is trained to complete the vacancy region corresponding to the foreground image in the background image to obtain a completed background image, which not only avoids the hole problem of the background region in the synthesized image, but also improves the authenticity and rationality of the completed background image. It can be seen that the technical scheme provided by the embodiments of the present application can reduce the background distortion while not affecting the purpose of morphing the target object, and improve the overall effect of the processed image.
[0074] To improve the completion effect of the background image, the missing area corresponding to the foreground image in the background image can be first subjected to texture completion, and then the missing area corresponding to the foreground image in the background image is subjected to color completion based on the texture completion result. Specifically, the above method can further include:
[0075] 105. Texture completion is performed on the missing area corresponding to the foreground image in the background image by using the trained second completion model to obtain a texture completion result.
[0076] Correspondingly, the "completing the missing area corresponding to the foreground image in the background image by using the trained first completion model to obtain a completed background image" in 103 can include:
[0077] 1031. Color completion is performed on the missing area corresponding to the foreground image in the background image by using the trained first completion model according to the texture completion result to obtain a completed background image.
[0078] In 105, the second completion model has been trained to a third convergence state, where the third convergence state can be set according to actual needs, and the embodiments of the present application do not make specific limitations. The texture completion result includes the texture of the missing area corresponding to the foreground image in the background image. The second completion model generates an extended texture of the non-missing area in the foreground image corresponding area according to the texture distribution of the non-missing area (i.e., the background area) in the background image.
[0079] In an example, texture extraction can be performed on the background image to obtain a texture image to be completed; texture completion is performed on the missing area corresponding to the foreground image in the texture image to be completed by using the trained second completion model to obtain a completed texture image, that is, the texture image to be completed is input into the second completion model to obtain the completed texture image output by the second completion model. The texture completion result includes the completed texture image. The training process of the second completion model will be described in the following embodiments.
[0080] In 1031, the input of the first completion model can be determined according to the background image and the texture completion result.
[0081] In an example, an input texture image can be determined according to the texture completion result; the background image and the input texture image are used as the input of the first completion model. The input texture image can only include the texture of the missing area corresponding to the foreground image in the background image, or the input texture image can include not only the texture of the missing area corresponding to the foreground image in the background image, but also the texture of the non-missing area in the background image.
[0082] Specifically, when the texture completion result includes the completed texture image, the completed texture image can be taken as the input texture image.
[0083] The background image and the input texture image are input into the first completion model to perform color completion on the empty area in the background image corresponding to the foreground image by the first completion model, to obtain a completed background image.
[0084] In the embodiments of the present application, the second completion model first performs texture completion on the empty area in the background image, and then inputs the texture completion result into the first completion model. In this way, the second completion model can refer to information in the texture aspect when performing color completion on the empty area in the background image, that is, the process of color completion is guided by the texture, which can improve the accuracy and rationality of the completion of the background image.
[0085] The training process of the first completion model and the second completion model will be introduced below. Specifically, the method can further include:
[0086] 106. According to the contour of the sample object, the corresponding contour region in the first image is determined.
[0087] 107. The pixels of the contour region in the first image are deleted to obtain a training sample.
[0088] 108. According to the training sample and the first image, the first completion model and the second completion model are trained respectively.
[0089] In 106, the sample object and the target object belong to the same category, so that the contours of the sample object and the target object are similar, and the training sample made is more meaningful for training. For example, if the target object is a person category, the sample object is also a person category; for another example, if the target object is an animal category, the sample object is also an animal category. In an example, the contour of the sample object can be manually drawn by a user. For example, a user has certain artistic skills, and he can draw some contours of sample objects according to actual needs. For example, a user can automatically draw some contours of people in a beauty scene.
[0090] In another example, a second image can be obtained; and an image segmentation model is used to determine the contour of the sample object corresponding to the sample object in the second image. The image segmentation model can be the same as the image segmentation model in the above embodiments. The image segmentation model is used to perform semantic segmentation on the sample object in the second image to obtain a mask image corresponding to the sample object; and the contour of the sample object is determined according to the mask image corresponding to the sample object.
[0091] In 106, the contour of the sample object only determines the shape of the contour region, and the position of the contour region can be randomly selected or set by default (for example, the center position in the first image). The embodiments of the present application do not make specific limitations on this.
[0092] In 107, the pixel values of all pixels in the contour region in the first image can be set to a preset value, which can be set according to actual needs. In an example, the preset value can be 0 or 255. Note that the value range of the pixel value is generally [0, 255], and the value range of the preset value can also be [0, 255].
[0093] In an example, the step of “training the first completion model and the second completion model respectively according to the training sample and the first image” in 108 includes the following steps:
[0094] 1081, performing texture extraction on the first image to obtain a reference texture image.
[0095] 1082, inputting the training sample and the reference texture image into the first completion model to perform color completion on the missing region corresponding to the contour region in the training sample by the first completion model, to obtain a completed training sample.
[0096] In the training sample, the contour of the missing region is the same as that of the contour region, and the position of the missing region in the training sample is the same as that of the contour region in the first image.
[0097] 1083, optimizing the first completion model according to the difference between the completed training sample and the first image.
[0098] In 1083, the first completion model can be optimized according to the difference between the pixels of the completed region corresponding to the contour region in the completed training sample and the pixels of the contour region in the first image.
[0099] In the completed training sample, the contour of the completed region is the same as that of the contour region, and the position of the completed region in the completed training sample is the same as that of the contour region in the first image.
[0100] The parameters in the first completion model can be optimized by a gradient back propagation algorithm according to the difference. The specific implementation of the gradient back propagation algorithm can be referred to the prior art, which will not be described in detail here.
[0101] In another example, the step of "training the first completion model and the second completion model according to the training samples and the first image" in 108 comprises the following steps:
[0102] 1084, performing texture extraction on the training samples to obtain a texture image of a sample to be completed.
[0103] 1085, inputting the texture image of the sample to be completed into the second completion model to perform texture completion on a blank area corresponding to the contour area in the texture image of the sample to be completed by the second completion model, to obtain a completed texture image of the sample.
[0104] Wherein, the contour of the blank area in the texture image of the sample to be completed is the same as the contour of the contour area, and the position of the blank area in the texture image of the sample to be completed is the same as the position of the contour area in the first image.
[0105] 1086, optimizing the second completion model according to the difference between the completed texture image of the sample and the reference texture image.
[0106] In 1086, the second completion model can be optimized according to the difference between the pixels of the completed area corresponding to the contour area in the completed texture image of the sample and the pixels of the reference area corresponding to the contour area in the reference texture image.
[0107] Wherein, the contour of the completed area in the completed texture image of the sample is the same as the contour of the contour area, and the position of the completed area in the completed texture image of the sample is the same as the position of the contour area in the first image.
[0108] Wherein, the contour of the reference area in the reference texture image is the same as the contour of the contour area, and the position of the reference area in the reference texture image is the same as the position of the contour area in the first image.
[0109] The parameters in the second completion model can be optimized according to the above difference by a gradient back propagation algorithm. The specific implementation of the gradient back propagation algorithm can be referred to the prior art, which will not be described in detail here.
[0110] Figure 2A flowchart of a model training method provided by an embodiment of the present application is shown. The execution subject of the method can be a client or a server. The client can be a hardware with an embedded program integrated on a terminal, an application software installed in the terminal, a tool software embedded in the operating system of the terminal, etc., which are not limited in the embodiments of the present application. The terminal can be any terminal device including a mobile phone, a tablet computer, a vehicle-mounted terminal device, etc. The server can be a common server, a cloud server or a virtual server, etc., which are not limited in the embodiments of the present application. As shown in FIG. 23, the method comprises: Figure 2
[0111] 201, determining a corresponding first contour region in the first image according to the contour of the sample object.
[0112] 202, deleting the pixels of the first contour region in the first image to obtain a training sample.
[0113] 203, training a first completion model according to the training sample and the first image.
[0114] The first completion model is used to complete a vacancy region corresponding to a foreground image in a background image to obtain a completed background image. The background image and the foreground image are obtained by segmenting a to-be-processed image of a target object. The foreground image corresponds to the target object. The target object and the sample object belong to the same category.
[0115] The specific implementation of steps 201, 202 and 203 can be referred to the corresponding content in the above embodiments, which will not be repeated here.
[0116] In the technical solution provided by the embodiments of the present application, the sample object and the target object belong to the same category, which can ensure that the contours of the sample object and the target object are similar, so that the training sample made is more meaningful for training. Thus, by using the trained first completion model to complete the vacancy region corresponding to the foreground image in the background image to obtain the completed background image, the authenticity and rationality of the completed background image can be further improved.
[0117] Optionally, in 203, “training the first completion model according to the training sample and the first image”, the following steps can be used to implement:
[0118] 2031, performing texture extraction on the first image to obtain a reference texture image.
[0119] 2032. Input the training sample and the reference texture image into the first completion model, so that the first completion model can fill in the missing regions corresponding to the contour regions in the training sample with color, and obtain the completed training sample.
[0120] 2033. Optimize the first completion model based on the difference between the completed training sample and the first image.
[0121] The specific implementation of steps 2031, 2032 and 2033 can be found in the corresponding contents of the above embodiments, and will not be repeated here.
[0122] It should be noted that any steps in the method provided in this application that are not described in detail can be found in the corresponding content of the above embodiments, and will not be repeated here. Furthermore, the method provided in this application may include other parts or all of the steps in the above embodiments in addition to the steps described above; for details, please refer to the corresponding content of the above embodiments, and will not be repeated here.
[0123] In one feasible solution, in a live streaming scenario, the target object can be: the streamer; the image to be processed can be: the live streaming image; the deformation processing can be: body shaping processing; the deformed foreground image is also the body shaping foreground image; and the target image can be: the target live streaming image. Figure 3 A schematic flowchart of a body shaping method according to an embodiment of this application is shown. The executing entity of this body shaping method is a live streaming client. Figure 3 As shown, the method includes:
[0124] 301. During the live broadcast, acquire live images captured from the broadcaster;
[0125] 302. Segment the live stream image to obtain a background image and a foreground image corresponding to the broadcaster;
[0126] 303. Perform body shaping processing on the foreground image to obtain a body-shaping foreground image;
[0127] 304. Using the trained first completion model, fill in the missing regions corresponding to the foreground image in the background image to obtain the completed background image;
[0128] 305. The foreground image after body shaping and the background image after completion are synthesized to obtain the target live image;
[0129] 306. Send the target live image to the live streaming server.
[0130] In the live streaming process, the live streaming client calls the camera on the terminal device to continuously collect live streaming images, then processes each live streaming image according to steps 301-305 to obtain a corresponding target live streaming image, and continuously sends the generated target live streaming image to the live streaming server to form a live streaming video stream. The live streaming server sends the live streaming video stream to the relevant audience client for display.
[0131] The specific implementation process of steps 301-305 can be referred to the corresponding content in the above embodiments, which will not be repeated here.
[0132] It should be noted that the content of each step in the method provided by the embodiments of the present application which is not fully described can be referred to the corresponding content in the above embodiments, which will not be repeated here. In addition, the method provided by the embodiments of the present application can include other parts or all steps in the above embodiments in addition to the above steps, which can be referred to the corresponding content in the above embodiments, which will not be repeated here.
[0133] In an implementable solution, in a live streaming scenario, the target object can be specifically a host, the image to be processed can be specifically a live streaming image, the deformation processing can be specifically body beautifying processing, the foreground image after deformation, i.e., the foreground image after body beautifying processing, and the target image can be specifically a target live streaming image. Figure 4 A flowchart of a body beautifying processing method provided by another embodiment of the present application is shown. The method is applicable to a live streaming client, which is built-in with a software development kit provided by a body beautifying service provider, and is provided by a live streaming service provider. As shown in Figure 4 The method comprises:
[0134] 401. In the live streaming process, a live streaming image collected by a host is obtained.
[0135] 402. The software development kit is called to perform body beautifying processing on the live streaming image to obtain a target live streaming image.
[0136] The software development kit is used to: segment the live streaming image to obtain a background image and a foreground image corresponding to the host; perform body beautifying processing on the foreground image to obtain a foreground image after body beautifying processing; use a trained first completion model to complete a vacancy area in the background image corresponding to the foreground image to obtain a background image after completion; and perform synthesis processing on the foreground image after body beautifying processing and the background image after completion to obtain the target live streaming image.
[0137] 403. The target live streaming image is sent to a live streaming server provided by the live streaming service provider.
[0138] In the embodiment, the live streaming service provider and the beauty service provider are different providers. The beauty service provider provides the developed beauty function in the form of a software development kit (SDK) to the live streaming service provider. The live streaming service provider can embed the SDK in the live streaming client when developing the live streaming client, without needing to understand how the beauty function is implemented, thereby improving the development efficiency. The live streaming service provider can be an Internet enterprise, for example, an enterprise providing game live streaming services, an enterprise providing live streaming goods sales services, an enterprise providing network conferences, and the like.
[0139] During the live streaming process, the live streaming client can call a camera on the terminal device to continuously collect live streaming images, call the SDK to process each live streaming image to obtain a corresponding target live streaming image, and continuously send the generated target live streaming image to the live streaming server to form a live streaming video stream. The live streaming server sends the live streaming video stream to related audience clients for display.
[0140] The specific processing process of the SDK can be referred to the corresponding content in the above embodiments, which will not be described here.
[0141] It should be noted that the details of the steps in the method provided in the embodiments of the present application can be referred to the corresponding content in the above embodiments, which will not be described here. In addition, the method provided in the embodiments of the present application can include other parts or all steps in the above embodiments in addition to the above steps, which can be referred to the corresponding content in the above embodiments, which will not be described here.
[0142] In actual application, the beauty function can be embedded into the browser in the form of a plug-in in addition to being embedded into the live streaming client in the form of an SDK. The specific implementation can be set according to actual needs, which is not limited in the embodiments of the present application.
[0143] In the above embodiments, the SDK is provided by the beauty service provider which is different from the live streaming service provider. Of course, the SDK can also be developed by the live streaming service provider itself.
[0144] The technical solutions provided in the embodiments of the present application will be introduced below taking a live streaming scenario as an example:
[0145] The terminal device (for example: mobile phone, computer) of the host runs the live broadcast client, and the live broadcast client collects the live broadcast image of the host by calling the camera of the terminal device. Then, the SDK is called to realize: the live broadcast image is segmented to obtain a background image and a foreground image containing the host; the host in the foreground image is beautified (for example: waist slimming, head circumference reduction, neck slimming, leg lengthening), to obtain a beautified foreground image; the trained second completion model is used to complete the texture of the vacancy area corresponding to the foreground image in the background image to obtain a texture completion result; according to the texture completion result, the trained first completion model is used to complete the color of the vacancy area corresponding to the foreground image in the background image to obtain a completed background image; the completed background image and the beautified foreground image are synthesized to obtain a target live broadcast image. Finally, the live broadcast client sends the target live broadcast image to the live broadcast server for sending to the corresponding audience client for display.
[0146] The beautification function can be applied not only in the live broadcast scene but also in the image editing scene. The application of the technical solutions provided by the embodiments of the present application in the image editing scene will be introduced below. Figure 5 The application of the technical solutions provided by the embodiments of the present application in the image editing scene will be introduced below.
[0147] The user can import the to-be-processed image from the local album through the editing interface provided by the terminal device (for example, a mobile phone), and the mobile phone terminal can send the to-be-processed image to the server. After receiving the to-be-processed image, the server performs the following processing:
[0148] Step 501, the to-be-processed image is segmented to obtain a background image 51a and a foreground image 52a corresponding to a person.
[0149] Step 502, the person in the foreground image 52a is beautified to obtain a beautified foreground image 52b.
[0150] The beautification includes head circumference reduction, slim neck, etc.
[0151] In step 503, the vacancy area corresponding to the foreground image 52a in the background image 51a is completed to obtain a completed background image 51b.
[0152] In step 504, the completed background image 51b and the beautified foreground image 52b are synthesized to obtain a target image.
[0153] After obtaining the target image, the server can send the target image to the terminal device of the user for display by the terminal device of the user. In an example, the terminal device can display the target image and the to-be-processed image side by side.
[0154] In the embodiment, the step of the body beautifying processing is executed by the server, and of course, can also be executed locally by the terminal device.
[0155] Figure 6 A structure block diagram of an image processing apparatus provided by an embodiment of the present application is shown. As shown in the figure, the apparatus comprises: Figure 6
[0156] The segmentation module 601 is configured to segment a target object to obtain a background image and a foreground image corresponding to the target object.
[0157] The deformation module 602 is configured to perform deformation processing on the foreground image to obtain a deformed foreground image.
[0158] The completion module 603 is configured to use a trained first completion model to complete a blank area in the background image corresponding to the foreground image to obtain a completed background image.
[0159] The synthesis module 604 is configured to perform synthesis processing on the deformed foreground image and the completed background image to obtain a target image.
[0160] In the technical scheme provided by the embodiment of the present application, the foreground image and the background image are separated, the target object in the foreground image is processed for deformation separately to obtain a deformed foreground image, so that the problem of background distortion caused by deformation processing is avoided. In addition, the blank area in the background image corresponding to the foreground image is completed by using the trained first completion model to obtain a completed background image, so that the problem of holes in the background area of the synthesized image is avoided, and the authenticity and rationality of the completed background image are improved. It can be seen that the technical scheme provided by the embodiment of the present application can reduce the background distortion while not affecting the purpose of target object deformation, and improve the overall effect of the processed image.
[0161] Optionally, the completion module 603 is specifically configured to:
[0162] use a trained second completion model to complete the texture of the blank area in the background image corresponding to the foreground image to obtain a texture completion result;
[0163] use the trained first completion model to complete the color of the blank area in the background image corresponding to the foreground image according to the texture completion result to obtain a completed background image.
[0164] Optionally, the completion module 603 is specifically configured to:
[0165] extract the texture of the background image to obtain a texture image to be completed;
[0166] perform texture completion on the blank area corresponding to the foreground image in the to-be-completed texture image by using the trained second completion model to obtain a completed texture image;
[0167] The texture completion result includes the completed texture image.
[0168] Optionally, the apparatus further includes:
[0169] The determining module is configured to determine a corresponding contour region in the first image according to the contour of the sample object.
[0170] The deleting module is configured to delete pixels of the contour region in the first image to obtain a training sample.
[0171] The training module is configured to train the first completion model and the second completion model respectively according to the training sample and the first image.
[0172] Optionally, the training module is specifically configured to:
[0173] perform texture extraction on the first image to obtain a reference texture image;
[0174] input the training sample and the reference texture image into the first completion model, so that the first completion model performs color completion on a blank area corresponding to the contour region in the training sample to obtain a completed training sample;
[0175] optimize the first completion model according to a difference between the completed training sample and the first image.
[0176] Optionally, the training module is specifically configured to:
[0177] perform texture extraction on the training sample to obtain a to-be-completed sample texture image;
[0178] input the to-be-completed sample texture image into the second completion model, so that the second completion model performs texture completion on a blank area corresponding to the contour region in the to-be-completed sample texture image to obtain a completed sample texture image;
[0179] optimize the second completion model according to a difference between the completed sample texture image and the reference texture image.
[0180] Optionally, the determining module is further configured to:
[0181] obtain a second image;
[0182] determine the contour of the sample object in the second image by using an image segmentation model.
[0183] Optionally, the segmentation module 601 is specifically used for:
[0184] performing semantic segmentation on the target object in the to-be-processed image by using the trained image segmentation model, to obtain a mask image corresponding to the target object;
[0185] segmenting the to-be-processed image according to the mask image, to obtain the foreground image and the background image.
[0186] Optionally, the target object includes a person.
[0187] Optionally, the morphing module 602 is specifically used for:
[0188] obtaining a preset body beautifying parameter;
[0189] performing body beautifying processing on a human body part of the person in the foreground image according to the body beautifying parameter, to obtain the morphed foreground image.
[0190] It should be noted that the image processing apparatus provided in the above embodiments can implement the technical solutions described in the above method embodiments, and the principles of implementation of the above modules can be referred to the corresponding content in the above method embodiments, which will not be described herein.
[0191] Figure 7 A structural block diagram of a model training apparatus provided in an embodiment of the present application is shown. As shown in the figure, Figure 7 the apparatus includes:
[0192] A determination module 701 is configured to determine a corresponding first contour region in a first image according to a contour of a sample object.
[0193] A deletion module 702 is configured to delete pixels of the first contour region in the first image, to obtain a training sample.
[0194] A training module 703 is configured to train a first completion model according to the training sample and the first image.
[0195] The first completion model is configured to complete a blank region corresponding to a foreground image in a background image, to obtain a completed background image; the background image and the foreground image are obtained by segmenting a to-be-processed image of a target object; the foreground image corresponds to the target object; and the target object and the sample object belong to the same category.
[0196] In the technical scheme provided by the embodiment of the present application, the sample object and the target object belong to the same category, so that the contours of the sample object and the target object are similar, and the training sample is more meaningful for training. In this way, the first completion model trained is used to complete the blank area corresponding to the foreground image in the background image, so that the authenticity and rationality of the completed background image are further improved.
[0197] Optionally, the training module 703 is specifically configured to:
[0198] perform texture extraction on the first image to obtain a reference texture image;
[0199] input the training sample and the reference texture image into the first completion model, so that the first completion model performs color completion on the blank area corresponding to the contour region in the training sample, and obtains a completed training sample;
[0200] According to the difference between the completed training sample and the first image, the first completion model is optimized.
[0201] It should be noted that the model training apparatus provided by the above embodiments can implement the technical solutions described in the above method embodiments, and the principles of the implementation of the above modules can be referred to the corresponding content in the above method embodiments, which will not be described here.
[0202] Figure 8 A structural schematic diagram of an electronic device provided by an embodiment of the present application is shown. As shown in Figure 8 The electronic device includes a memory 1101 and a processor 1102. The memory 1101 can be configured to store various data to support operations on the electronic device. Examples of the data include instructions for operating any application or method on the electronic device. The memory 1101 can be realized by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk.
[0203] The memory 1101 is configured to store a program;
[0204] The processor 1102 is coupled to the memory 1101 and is configured to execute the program stored in the memory 1101 to implement the image processing method, body beautifying method or model training method provided by the above method embodiments.
[0205] Further, asFigure 8 As shown, the electronic device also includes a communication component 1103, a display 1104, a power component 1105, an audio component 1106, and other components. Figure 8 The electronic device is only schematically shown with some components, and is not meant to limit the electronic device to only including Figure 8 the components shown.
[0206] Correspondingly, the embodiments of the present application further provide a computer readable storage medium storing a computer program, which can implement the steps or functions of the image processing method, the body beautifying processing method or the model training method provided by each method embodiment when the computer program is executed by a computer.
[0207] The device embodiments described above are only schematic, and the units shown as separate components can or can not be physically separate, and the components shown as units can or can not be physical units, i.e., can be located in one place, or can be distributed on multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the embodiments. Those skilled in the art can understand and implement without creative labor.
[0208] From the above description of the embodiments, those skilled in the art can clearly understand that the embodiments can be implemented by means of software plus necessary universal hardware platforms, and of course can also be implemented by hardware. Based on such understanding, the above technical solutions, essentially or in other words, the part that contributes to the prior art, can be embodied in the form of a software product, which can be stored in a computer readable storage medium, such as a ROM / RAM, a magnetic disk, an optical disk, etc., and includes a number of instructions to make a computer device (which can be a personal computer, a server, or a network device, etc.) execute the methods described in each embodiment or some part of the embodiments.
[0209] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present application, and not to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacements to some technical features; and these modifications or replacements do not make the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. An image processing method, wherein, include: The image of the target object is segmented to obtain a background image and a foreground image corresponding to the target object; The foreground image is deformed to obtain a deformed foreground image; The background image is subjected to texture extraction to obtain the texture image to be completed; Using the trained second completion model, texture completion is performed on the missing regions corresponding to the foreground image in the texture image to be completed, resulting in a completed texture image. Based on the completed texture image, the trained first completion model is used to fill in the missing areas of the foreground image in the background image with color, so as to obtain the completed background image. The deformed foreground image and the completed background image are combined to obtain the target image.
2. The method according to claim 1, wherein, Also includes: Based on the contour of the sample object, the corresponding contour region is determined in the first image; The sample object and the target object belong to the same category; Delete pixels in the contour region from the first image to obtain training samples; The first completion model and the second completion model are trained based on the training samples and the first image, respectively.
3. The method according to claim 2, wherein, Based on the training samples and the first image, the first completion model and the second completion model are trained respectively, including: Texture extraction is performed on the first image to obtain a reference texture image; The training sample and the reference texture image are input into the first completion model so that the first completion model can fill in the missing regions corresponding to the contour regions in the training sample with color to obtain the completed training sample. The first completion model is optimized based on the difference between the completed training sample and the first image.
4. The method according to claim 3, wherein, Training the first completion model and the second completion model based on the training samples and the first image, respectively, further includes: Texture extraction is performed on the training samples to obtain the texture image of the sample to be completed; The texture image of the sample to be completed is input into the second completion model, so that the second completion model can complete the texture of the missing region corresponding to the contour region in the texture image of the sample to be completed, and obtain the completed sample texture image. The second completion model is optimized based on the difference between the completed sample texture image and the reference texture image.
5. The method according to claim 2, wherein, Also includes: Obtain the second image; The outline of the sample object is determined in the second image using an image segmentation model.
6. The method according to any one of claims 1 to 5, wherein, The image to be processed is segmented to obtain a background image and a foreground image corresponding to the target object, including: Using a trained image segmentation model, semantic segmentation is performed on the target object in the image to be processed to obtain a mask image corresponding to the target object; Based on the mask image, the image to be processed is segmented to obtain the foreground image and the background image.
7. The method according to any one of claims 1 to 5, wherein, The target objects include: people.
8. The method according to claim 7, wherein, The foreground image is deformed to obtain a deformed foreground image, including: Obtain preset body shaping parameters; Based on the body shaping parameters, body shaping is performed on the human body parts of the person in the foreground image to obtain the deformed foreground image.
9. A model training method, wherein, include: Based on the contour of the sample object, the corresponding first contour region is determined in the first image; Delete pixels from the first contour region in the first image to obtain training samples; Training the first completion model based on the training samples and the first image includes: The training samples and the reference texture image are input into the first completion model, so that the first completion model performs color completion on the missing regions corresponding to the contour regions in the training samples to obtain the completed training samples; wherein, the reference texture image is obtained by extracting texture from the first image; The first completion model is optimized based on the difference between the completed training sample and the first image. The texture image of the sample to be completed is input into the second completion model, so that the second completion model can complete the texture of the missing regions corresponding to the contour regions in the texture image of the sample to be completed, thereby obtaining the completed sample texture image; wherein, the texture image of the sample to be completed is obtained by extracting texture from the training samples; Based on the difference between the completed sample texture image and the reference texture image, the second completion model is optimized; wherein, the first completion model and the second completion model are used to: complete the missing regions corresponding to the foreground image in the background image to obtain the completed background image; the background image and the foreground image are obtained by segmenting the image to be processed of the target object; the foreground image corresponds to the target object; the target object and the sample object belong to the same category.
10. The method according to claim 9, wherein, Training the first completion model based on the training samples and the first image includes: Texture extraction is performed on the first image to obtain a reference texture image.
11. A body shaping method, applicable to live streaming clients, wherein, The method includes: During the live stream, capture live images of the streamer; The live stream image is segmented to obtain a background image and a foreground image corresponding to the broadcaster; The foreground image is subjected to body shaping processing to obtain a body-shaping foreground image; The background image is subjected to texture extraction to obtain the texture image to be completed; Using the trained second completion model, texture completion is performed on the missing regions corresponding to the foreground image in the texture image to be completed, resulting in a completed texture image. Based on the completed texture image, the trained first completion model is used to fill in the missing areas of the foreground image in the background image with color, so as to obtain the completed background image. The foreground image after body shaping and the background image after completion are synthesized to obtain the target live image; The target live image is sent to the live streaming server.
12. A body shaping method, applicable to a live streaming client, wherein the live streaming client has a built-in software development kit provided by a body shaping service provider; The live streaming client is provided by the live streaming service provider; wherein... The method includes: During the live stream, capture live images of the streamer; The software development kit is invoked to perform body shaping processing on the live stream image to obtain the target live stream image; The target live image is sent to the live server provided by the live service provider. The software development kit (SDK) is used to: segment the live stream image to obtain a background image and a foreground image corresponding to the streamer; perform body shaping on the foreground image to obtain a body-shaping foreground image; extract texture from the background image to obtain a texture image to be completed; use a trained second completion model to complete the texture in the texture image to be completed, corresponding to the foreground image, to obtain a completed texture image; based on the completed texture image, use a trained first completion model to complete the color in the background image, corresponding to the foreground image, to obtain a completed background image; and synthesize the body-shaping foreground image and the completed background image to obtain the target live stream image.
13. An electronic device, wherein, include: Memory and processor, among which, The memory is used to store programs; The processor, coupled to the memory, is configured to execute the program stored in the memory to implement the method of any one of claims 1 to 12.
14. A computer-readable storage medium storing a computer program, wherein, When the computer program is executed by a computer, it can implement the method of any one of claims 1 to 12.
Citation Information
Patent Citations
Human body image body beautifying method and device, equipment and medium
CN114298941A