A live class portrait slimming processing method and device and electronic equipment

By real-time collection and processing of portraits and blackboard videos in live classes, the teacher's image can be slimmed down, the problem of improving the teacher's image in live classes can be solved, and students' enthusiasm and concentration in class can be improved.

CN112561790BActive Publication Date: 2025-10-10ZUOYEBANG EDUCATION TECH (BEIJING) CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202011539638.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-12-23
Publication Date
2025-10-10
Estimated Expiration
2040-12-23

AI Technical Summary

Technical Problem

How to slim down teachers in live classes, improve their image, and thus increase students' enthusiasm and concentration in class.

Method used

Real-time video images of the first scene containing portraits and the second scene containing blackboard writing in the live class are collected. The portrait images are extracted through image segmentation technology, the slimming position information is located, and the positions of key parts are determined using a deep learning model. The slimmed-down portrait images are processed by image distortion and fused with the blackboard video images. The Gaussian aliasing algorithm is used to process the pasting boundaries, and the lighting distribution is adjusted in combination with the graphics lighting model.

Benefits of technology

The teacher's portrait in the live class can be slimmed down in real time without deformation, which improves the teacher's image, enhances students' learning attention and classroom atmosphere, and improves learning effects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN112561790B_ABST
    Figure CN112561790B_ABST
Patent Text Reader

Abstract

The present application belongs to the technical field of live class, and provides a live class portrait slimming processing method and device and electronic equipment, the method comprising: respectively collecting a first scene video picture containing a portrait and a second scene video picture containing a board book in a live class; extracting a portrait picture in the first scene video picture; positioning slimming position information of the portrait picture; performing slimming processing on the slimming position information to obtain a slimmed portrait picture; fusing the slimmed portrait picture with the second scene video picture; and outputting the fused video picture. The present application can realize real-time slimming of a live class portrait and ensure the effect of non-deformation of a board book, effectively improving the teacher image in a live class, thereby improving the student's enthusiasm and concentration in class.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of online education technology, and is particularly applicable to online live class technology. More specifically, it relates to a method, device, electronic device and computer-readable medium for processing portrait slimming in live classes. Background Art

[0002] With the development of Internet technology and people's increasing emphasis on education, Internet live classes are popular among students and parents because they are guided by famous teachers and are cheaper than traditional courses.

[0003] Live internet classes typically require students to watch the lectures and complete exercises via internet-connected devices such as tablets or smartphones. During live classes, teachers interact with students via online video. A positive teacher image in live videos can attract students and increase their enthusiasm and focus. Therefore, how to streamline the teacher's image and enhance their image in live classes has become a pressing issue for live class videos. Summary of the Invention

[0004] (1) Technical issues to be resolved

[0005] The present invention aims to solve the technical problem of how to slim down teachers in live classes, improve their image, and thus increase students' enthusiasm and concentration in class.

[0006] (2) Technical solution

[0007] In order to solve the above technical problems, one aspect of the present invention provides a method for processing a slimming portrait in a live class, the method comprising the following steps:

[0008] Collect the first scene video images containing portraits and the second scene video images containing blackboard writing in the live class respectively;

[0009] Extracting a portrait image from the video image of the first scene;

[0010] Locating the slimming position information of the portrait image;

[0011] Performing slimming processing on the slimming position information to obtain a slimming portrait image;

[0012] Merging the slimmed-down portrait image with the second scene video image;

[0013] Output the fused video image.

[0014] According to a preferred embodiment of the present invention, the locating of the weight loss position information of the portrait image includes:

[0015] Extracting key part position information and key part height information of the portrait image;

[0016] Determining standard position information of key parts according to the key part height information;

[0017] The key part position information is compared with the key part standard position information to obtain the weight loss position information that needs to be lost.

[0018] According to a preferred embodiment of the present invention, the portrait image is input into a deep learning model to obtain the position information of key parts of the portrait image.

[0019] According to a preferred embodiment of the present invention, before inputting the portrait image into the deep learning model to obtain the key part position information of the portrait image, the method further includes:

[0020] Get portrait images from live history classes as a sample set;

[0021] Marking the key parts of the sample set images;

[0022] A deep learning model is trained based on the sample set and the key part position information.

[0023] According to a preferred embodiment of the present invention, image distortion is used to map the pixels of the weight loss position information into a preset range to obtain a weight loss portrait image.

[0024] According to a preferred embodiment of the present invention, the step of fusing the weight-loss portrait image with the second scene video image includes:

[0025] Paste the slimming portrait image to a position other than the blackboard in the second scene video image;

[0026] A Gaussian aliasing algorithm is used to process the pasting boundary between the portrait image and the second scene video image.

[0027] According to a preferred embodiment of the present invention, after fusing the weight-loss portrait image with the second scene video image, the method further includes:

[0028] The fused image is processed using the graphics lighting model.

[0029] According to a preferred embodiment of the present invention, the method further comprises:

[0030] When a scene switching instruction is received, the second scene video picture is switched to the corresponding scene picture.

[0031] According to a preferred embodiment of the present application, the method further comprises:

[0032] When the key knowledge prompt instruction is received, the key knowledge information is prompted.

[0033] The second aspect of the present application provides a live class portrait slimming processing device, the device comprises:

[0034] The acquisition module is configured to acquire a first scene video picture containing a portrait and a second scene video picture containing a blackboard picture in the live class, respectively.

[0035] The extraction module is configured to extract a portrait picture in the first scene video picture.

[0036] The positioning module is configured to position slimming position information of the portrait picture.

[0037] The slimming processing module is configured to perform slimming processing on the slimming position information to obtain a slimmed portrait picture.

[0038] The fusion module is configured to fuse the slimmed portrait picture with the second scene video picture.

[0039] The output module is configured to output the fused video picture.

[0040] The third aspect of the present application provides an electronic device comprising a processor and a memory, wherein the memory is configured to store a computer executable program, and when the computer executable program is executed by the processor, the processor executes the method.

[0041] The fourth aspect of the present application further provides a computer readable medium storing a computer executable program, and when the computer executable program is executed, the method is implemented.

[0042] (III) Beneficial effects

[0043] The present application acquires at least two scene video pictures of a live class in real time, wherein one scene video picture is composed of a first scene video picture containing a portrait, and the other scene video picture is composed of a second scene video picture containing a blackboard; the portrait picture in the first scene video picture is extracted; the slimming position information of the portrait picture is positioned; the slimming position information is subjected to slimming processing to complete the individual slimming of the portrait in the live class; the slimmed portrait picture is fused with the second scene video picture; and the fused video picture is outputted. Thus, the real-time slimming of the portrait in the live class and the effect of ensuring the non-distortion of the blackboard are realized, the teacher image in the live class is effectively improved, and thus the student's enthusiasm and concentration in class are improved.

[0044] The present invention adopts the Gaussian aliasing algorithm to process the pasting boundary between the portrait image and the second scene video image, which can effectively alleviate the mosaic phenomenon that appears at the pasting boundary between the portrait image and the second scene video image due to inconsistent lighting, and improve the quality of the fused video image.

[0045] When the present invention receives a scene switching instruction, it switches the second scene video picture to the corresponding scene picture, which can further improve students' learning attention and enhance the classroom atmosphere.

[0046] When receiving a key knowledge prompt instruction, the present invention prompts key knowledge information, deepens students' attention and memory of key knowledge, and improves learning effects. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] Figure 1 It is a flow chart of a method for processing a slimming portrait in a live class according to the present invention;

[0048] Figure 2 is a schematic diagram of locating the weight loss position information of the portrait image according to the present invention;

[0049] Figure 3a This is a schematic diagram of the present invention showing a live class in real time including a slimmed-down teacher and blackboard writing;

[0050] Figure 3b is Figure 3a Schematic diagram of automatically switching the video picture of the second scene after receiving a scene switching instruction during the live broadcast;

[0051] Figure 4 It is a schematic diagram of a key display method of the present invention;

[0052] Figure 5 This is a structural diagram of a device for processing slimming portraits in live classes according to the present invention;

[0053] Figure 6 This is a schematic structural diagram of an electronic device according to an embodiment;

[0054] Figure 7 is a schematic diagram of a computer-readable recording medium according to an embodiment of the present invention. DETAILED DESCRIPTION

[0055] In the introduction of specific embodiments, the detailed description of the structure, performance, effect or other features is intended to enable those skilled in the art to fully understand the embodiments. However, this does not preclude those skilled in the art from implementing the present invention with a technical solution that does not include the aforementioned structure, performance, effect or other features under specific circumstances.

[0056] The flowcharts in the accompanying drawings are merely illustrative of the process flow and do not necessarily include all of the content, operations, and steps in the flowcharts, nor do they necessarily imply that all of the steps in the flowcharts must be executed in the order shown. For example, some of the steps in the flowcharts may be separated, some may be combined or partially combined, and so on. The execution order shown in the flowcharts may be changed according to actual circumstances without departing from the spirit of the present invention.

[0057] Frames in the accompanying drawings Figure 1 The term "functional entity" generally refers to a functional entity and does not necessarily correspond to a physically independent entity. That is, these functional entities may be implemented in software, in one or more hardware modules or integrated circuits, or in different networks and / or processing unit devices and / or microcontroller devices.

[0058] The same reference numerals in the accompanying drawings represent the same or similar elements, components or parts, and thus repeated descriptions of the same or similar elements, components or parts may be omitted below. It should also be understood that although the first, second, third and other numbered adjectives may be used herein to describe various devices, elements, components or parts, these devices, elements, components or parts should not be limited by these adjectives. In other words, these adjectives are only used to distinguish one from another. For example, the first device may also be called the second device, but this does not deviate from the essential technical solution of the present invention. In addition, the terms "and / or" and "and / or" refer to all combinations including any one or more of the listed items.

[0059] The present invention is primarily used to slim down teachers in online live classes. Since online live broadcasts are real-time scenes, not static single images, high speed requirements are required, ensuring real-time processing. Furthermore, during online live classes, teachers write on the blackboard on the display screen. Slimming down only the teacher's image can cause deformation and distortion of the blackboard writing in the slimmed-down areas. Therefore, the present invention's method for slimming down portraits in live classes needs to ensure both real-time performance and the preservation of distortion in the blackboard writing after slimming down. To this end, the present invention captures at least two scene videos of a live class in real time, one of which consists of a first scene video image containing a portrait, and the other of which consists of a second scene video image containing the blackboard writing. The method extracts a portrait image from the first scene video image, locates the slimming position information of the portrait image, and then slims down the slimming position information to achieve a separate slimming of the portrait in the live class. The slimmed-down portrait image is then merged with the second scene video image, and the resulting fused video image is output. This achieves real-time slimming of the portrait in the live class while preserving distortion in the blackboard writing.

[0060] In a specific embodiment, the present invention extracts key part position information and key part height information from a portrait image; determines standard key part position information based on the key part height information based on a standard height distribution; and compares the key part position information with the key part standard position information to obtain weight loss position information. The key part positions of a portrait image refer to positions requiring weight loss, including but not limited to the waist and face. Correspondingly, the key part position information of the portrait image can be waist width, face width, etc. The key part position information of the portrait image can be obtained specifically through a pre-trained deep learning model.

[0061] The present invention uses image distortion to map the pixels of the weight-loss position information to a preset range to obtain a weight-loss portrait image, wherein the preset range is a weight-loss range set according to the height of the portrait.

[0062] The present invention pastes the slimmed-down portrait image into the second scene video image, excluding the blackboard text. A Gaussian blending algorithm is then used to process the border between the portrait image and the second scene video image. This effectively mitigates the mosaic effect that occurs at the border between the portrait image and the second scene video image due to lighting inconsistencies, thereby improving the quality of the fused video image. Furthermore, a graphics-based lighting model is used to process the fused image, further eliminating any lighting inconsistencies between the portrait image and the second scene video image, thereby enhancing the authenticity of the fused image.

[0063] When the present invention receives a scene switching instruction, it switches the second scene video picture to the corresponding scene picture, which can further improve students' learning attention and enhance the classroom atmosphere.

[0064] When receiving a key knowledge prompt instruction, the present invention prompts key knowledge information, deepens students' attention and memory of key knowledge, and improves learning effects.

[0065] In order to make the objectives, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with specific embodiments and with reference to the accompanying drawings.

[0066] In the introduction of specific embodiments, the detailed description of the structure, performance, effect or other features is intended to enable those skilled in the art to fully understand the embodiments. However, this does not preclude those skilled in the art from implementing the present invention with a technical solution that does not include the aforementioned structure, performance, effect or other features under specific circumstances.

[0067] The flowchart in the drawing is only an example of flow demonstration, does not represent that all the contents, operations and steps in the flowchart must be included in the scheme of the present application, and does not represent that the execution order shown in the drawing must be executed. For example, some operations / steps in the flowchart can be decomposed, some operations / steps can be combined or partially combined, etc. The execution order shown in the flowchart can be changed according to the actual situation without departing from the inventive concept of the present application.

[0068] The block in the drawing Figure 1 Generally represents a functional entity, and does not necessarily correspond to a physically independent entity. That is, these functional entities can be implemented in software form, or in one or more hardware modules or integrated circuits, or in different network and / or processing unit devices and / or microcontroller devices.

[0069] The same reference signs in the various drawings represent the same or similar elements, components or parts, so that the repeated description of the same or similar elements, components or parts can be omitted hereinafter. It should also be understood that although the first, second, third, etc. representative adjectives may be used herein to describe various devices, elements, components or parts, these devices, elements, components or parts should not be limited by these adjectives. That is, these adjectives are only used to distinguish one from another. For example, a first device can also be referred to as a second device without departing from the essential technical solution of the present application. In addition, the terms "and / or", "and / or" mean all combinations of one or more of the listed items.

[0070] Figure 1 is a flowchart of a live class portrait slimming processing method of the present application, as Figure 1 shown, the method comprises the following steps:

[0071] S1, respectively collecting a first scene video picture containing a portrait in a live class and a second scene video picture containing a blackboard;

[0072] In order to slim the portrait in the live class while ensuring the real-time performance of the live class and the effect that the blackboard is not distorted after slimming, the present application needs to collect the scenes containing the portrait and the scenes containing the blackboard in the live class in real time. Specifically, two cameras can be used to shoot simultaneously, one for real-time collection of a first scene video picture containing a portrait in a live class, and the other for real-time collection of a second scene video picture containing a blackboard in a live class. For example, a first camera is used to collect the real scene of a teacher giving a live class, and a second camera is used to collect the pure blackboard scene in the live class.

[0073] To ensure real-time slimming of portraits in live classes, the video images are collected in real time from the live class and included in the video playing simultaneously with the class. To balance speed and effectiveness of the slimming process, every frame of the video playing simultaneously with the class can be selected as the video image, or every few frames of the video playing simultaneously with the class can be selected as the video image.

[0074] S2. Extracting a portrait image from the first scene video image;

[0075] Specifically, image segmentation technology can be used to extract the teacher's image from the first scene video image. Image segmentation refers to the process of dividing an image into several regions with similar properties. It can be applied to scene object segmentation, human background segmentation, face and body parsing, 3D reconstruction and other technologies. Image segmentation can be roughly divided into three categories: graph theory-based methods, pixel clustering-based methods, and deep semantics-based methods. The present invention preferably uses a deep speech-based method to extract the teacher's image from the first scene video image.

[0076] S3, locating the weight loss position information of the portrait image;

[0077] Exemplarily, the present invention extracts key part position information and key part height information from a portrait image; based on the standard height and weight distribution, determines the key part standard position information according to the key part height information; compares the key part position information with the key part standard position information to obtain the slimming position information required for slimming.

[0078] Among them, the key parts of the portrait image refer to the positions that need to be slimmed down. In the present invention, waist slimming can be achieved by reducing the teacher's waist width, and face slimming can also be achieved by reducing the teacher's face width. Therefore, the key parts of the portrait image include but are not limited to: waist, face, etc. The height of the key parts of the portrait image corresponds to the height of the position that needs to be slimmed down. For example, if the waist needs to be slimmed down, the key part height refers to the portrait height. If the face needs to be slimmed down, the key part height refers to the face height. Correspondingly, the key part position information of the portrait image can be the waist width, and the key part height information in the portrait image is the portrait height information; the key part position information of the portrait image can also be the face width, and the key part height information in the portrait image is the face height information.

[0079] In the present invention, the key parts positions of the portrait image can be specifically obtained through a pre-trained deep learning model. When training the deep learning model, you can first obtain portrait images in historical live classes from historical data or public data sets as a sample set; mark the key parts position information (such as waist width, face width, etc.) in the sample set images; train the deep learning model based on the sample set and the key parts position information. Furthermore, you can also input the portrait image to be verified into the trained deep learning model to obtain the predicted key parts position information, compare the actual key parts position information of the video frame image to be verified with the predicted key parts position information, and calculate the loss function; if the loss function is less than the preset value, the deep learning model is determined as the final trained deep learning model. Among them, the deep learning model can be specifically implemented using network structures such as CNN, Hourglass, Attention, Transform, LSTM, etc.

[0080] Similarly, the height information of key parts in the portrait image can also be obtained by a trained deep learning model. The present invention will configure a correspondence table based on the standard height and fatness distribution in advance based on the first scene video image coordinate system. The correspondence table contains the standard position information of key parts corresponding to the key part height information. The standard position information of the key parts of the current portrait can be determined based on the extracted key part height information and the above correspondence. By comparing the key part position information extracted from the current portrait with the key part standard position information, the slimming position information that needs to be slimmed down can be obtained. Among them, the slimming position information includes: the coordinates of the slimming part (such as waist, face, etc.), and the slimming length (that is, the difference between the key part position information and the key part standard position information in the first scene video image coordinate system).

[0081] Take the example of slimming down the teacher’s waist in a live class. Figure 2 As shown, in this step, the waist width D and the teacher height H in the teacher image 21 in the first scene video picture are obtained through the trained deep learning model, and the standard waist width D corresponding to the teacher height H is found according to the corresponding relationship. K , then compare the teacher's waist width D with the standard waist width D K , determine the coordinates of the slimming part and the slimming length.

[0082] S4, performing slimming processing on the slimming position information to obtain a slimming portrait image;

[0083] The present invention uses image distortion to map the pixels of the weight loss position information to a preset range, thereby obtaining a weight loss portrait image. Specifically, based on the coordinate-pixel conversion relationship of the first scene video image, the weight loss position coordinates are converted to weight loss position pixels, and the weight loss length is converted to the corresponding weight loss pixels. Then, through image distortion, the pixels of the weight loss position are mapped to the weight loss pixel range to obtain the weight loss portrait image. Image distortion involves mapping each pixel in the image to a new position according to a certain pattern, essentially determining the new pixel coordinates x and y.

[0084] S5, integrating the weight-loss portrait image with the second scene video image;

[0085] Specifically, the present invention pastes the slimmed-down portrait image into the second scene video image, excluding the blackboard text. A Gaussian blending algorithm is then used to process the border between the portrait image and the second scene video image. This effectively mitigates the mosaic effect that can occur at the border between the portrait image and the second scene video image due to inconsistent lighting, thereby improving the quality of the fused video image.

[0086] The Gaussian aliasing algorithm includes the following steps:

[0087] S11. Construct the corresponding Gaussian residual pyramids (the number of layers is the preset value level) based on the images L and R to be fused, and retain the top image of the Gaussian pyramid downsampling (the smallest image, level+1 layer):

[0088] The Gaussian pyramid construction method is as follows, taking picture L as an example:

[0089] (1) Perform Gaussian downsampling on the image L to obtain downL, which can be implemented by using the pyrDown() function in OpenCV. Then perform Gaussian upsampling on downL to obtain upL, which can be implemented by the pyrUp() function in OpenCV.

[0090] (2) Calculate the residual between the image L and upL, and obtain a residual image lapL0, which is the image at the lowest end of the Gaussian pyramid.

[0091] (3) Continue steps (1) and (2) for downL, and continuously calculate the residual maps lapL1, lap2, lap3, ..., lapN. This will result in a series of residual maps, which are the Gaussian residual pyramid.

[0092] (4) There are level images in the Gaussian residual pyramid. Keep the Gaussian downsampled image topL of level+1 for later use.

[0093] S12. Sampling under the binary mask mask constructs a Gaussian pyramid, also implemented using pyrDown(), with a total of level+1 layers.

[0094] S13. Using the mask images at each level of the mask pyramid, merge the images of the corresponding levels of the Gaussian residual pyramids of images L and R into a single image. This yields a merged Gaussian residual pyramid. Simultaneously, using the topmost mask, merge the topL and topR retained in step S11 into topLR.

[0095] S14. Take topLR as the image at the top of the pyramid, use the pyrUp() function to perform Gaussian upsampling on topLR to obtain upTopLR, and add upTopLR to the image of the corresponding layer of the residual pyramid merged in step S13 to reconstruct the image of this layer.

[0096] S15. Repeat step S14 until the 0th layer, that is, the image at the lowest end of the pyramid, namely blendImg, is reconstructed and output.

[0097] Furthermore, in the present invention, the light distribution of the slimmed-down portrait image and the second scene video image may differ. To improve the authenticity of the fused image, it is necessary to eliminate the difference in light distribution between the slimmed-down portrait image and the second scene video image. Therefore, this step can further utilize a graphics-based light model to process the fused image.

[0098] Among them, the illumination model in graphics is used to describe the relationship between the light source and the illumination of the object surface in a three-dimensional scene. The present invention obtains the illumination distribution of the weight-loss portrait image and the second scene video image respectively by solving the illumination model parameters, and then corrects the illumination distribution of the weight-loss portrait image according to the illumination distribution of the second scene video image, so that the illumination distribution of the two is consistent. Obviously, the illumination distribution of the second scene video image can also be corrected according to the illumination distribution of the weight-loss portrait image, so that the illumination distribution of the two is consistent. Alternatively, the illumination distribution of the weight-loss portrait image set and the second scene video image can also be adjusted to a pre-set optimal illumination intensity at the same time, so that the illumination distribution of the two is consistent, which is not specifically limited by the present invention.

[0099] Specifically, there are many lighting models in graphics. Taking the Phong Lighting Model as an example, the main structure of the Phong Lighting Model consists of three components: ambient, diffuse, and specular lighting.

[0100] Among them, Ambient Lighting uses an ambient lighting constant to simulate that in a dark environment, objects still have some light (moon, distant light). The ambient lighting C can be calculated by the following formula amb :

[0101]

[0102] Where: m amb is the ambient light component of the material, which is always equal to the diffuse component. amb is the ambient lighting value for the entire scene.

[0103] Diffuse lighting simulates the directional impact of light on an object. Diffuse lighting obeys Lambert's law: the reflected light intensity is proportional to the cosine of the angle between the normal vector and the light. Using dot product to calculate the cosine, the diffuse lighting C can be calculated using the following formula: diff :

[0104]

[0105] Where: n is the surface normal vector, l is the unit vector pointing to the light source, m diff is the scattering color of the material, that is, the color of the object recognized by most people, S diff is the scattered color of the light source, m gls The glossiness of the material, also known as the Phong exponent.

[0106] Specular lighting simulates the bright spots that appear on shiny objects. Specular lighting C can be calculated using the following formula: spec :

[0107]

[0108] Where: θ is the angle between r and v, given by r·v, which describes the orientation of the mirror image; v points towards the observer; r is the "mirror" vector. gls is the glossiness of the material, also known as the Phong index. m spec The reflected color of the material controls the intensity of the light spot. spec is the specular color of the light source.

[0109] The above formula can be used to calculate the ambient lighting, diffuse lighting, and specular lighting of the weight-loss portrait image, as well as the ambient lighting, diffuse lighting, and specular lighting of the second scene video image. Combined with the light attenuation (Attenuation), the following light attenuation formula is obtained:

[0110]

[0111] Among them: A is ambient lighting, D is diffuse lighting, S is specular lighting, k is the coefficient of the corresponding parameter, and a0, a1 and a2 are attenuation parameters.

[0112] By adjusting the values ​​of the three parameters a0, a1 and a2 of the weight-loss portrait image and / or the second scene video image, different light intensity attenuation effects can be achieved, thereby making the lighting distribution of the weight-loss portrait image and the second scene video image consistent.

[0113] S6. Output the fused video image.

[0114] The present invention captures at least two scene videos of a live class in real time, wherein one scene video is composed of a first scene video image containing a portrait, and the other scene video is composed of a second scene video image containing blackboard writing. The present invention extracts the portrait image from the first scene video image, locates the slimming position information of the portrait image, and then performs slimming processing on the slimming position information to achieve the slimming of the portrait in the live class. The slimmed-down portrait image is then fused with the second scene video image, and the fused video image is output. This achieves the effect of slimming the portrait in the live class in real time while ensuring that the blackboard writing does not deform.

[0115] In order to further improve students' learning attention and enhance classroom atmosphere, the present invention switches the second scene video picture to the corresponding scene picture when receiving a scene switching instruction. Figure 3a In the live class, the video of the teacher 31 and the blackboard picture 32 after slimming down is displayed in real time. When explaining the ancient poem "The desert is filled with solitary smoke, and the long river is filled with round sunset" in the live class, the teacher can issue a scene switching instruction by clicking on the screen of the live device, and then Figure 3b , automatically switching the blackboard writing 32 in the live class to the scene of sunset in the Gobi Desert.

[0116] Furthermore, in order to deepen students' attention and memory of key knowledge and improve learning effects. The present invention prompts key knowledge information when receiving key knowledge prompt instructions. Among them, the key knowledge prompt instruction can be obtained by detecting the teacher's touch on the detection device (such as a touch screen, physical buttons, etc. set on the live broadcast equipment). The method of prompting key knowledge can be to play the key knowledge at a volume greater than the normal volume, or to display the key knowledge on the live broadcast screen in a key display manner, or it can be a combination of the two. Figure 4The key display method may be to display a cartoon image 11 (such as a rocking bear) in the upper right corner of the live screen 10, and to display the key knowledge XXX below the cartoon image 11. The key display method may also be to display the key knowledge in an enlarged color font on the live screen, which is not specifically limited in the present invention.

[0117] Figure 5 This is a structural diagram of a device for processing a slimming portrait of a live class according to the present invention. Figure 5 As shown, the device includes:

[0118] The acquisition module 51 is used to respectively acquire a first scene video image containing a portrait and a second scene video image containing a blackboard image in the live class;

[0119] An extraction module 52 is configured to extract a portrait image from the first scene video image;

[0120] A positioning module 53 is used to locate the weight loss position information of the portrait image;

[0121] A slimming processing module 54 is used to perform slimming processing on the slimming position information to obtain a slimming portrait image;

[0122] A fusion module 55 is configured to fuse the weight-loss portrait image with the second scene video image;

[0123] The output module 56 is used to output the fused video image.

[0124] In a specific embodiment, the positioning module 53 includes:

[0125] An extraction module, configured to extract key part position information and key part height information of the portrait image;

[0126] A determination module, configured to determine standard position information of key parts according to the key part height information;

[0127] The comparison module is used to compare the key part position information with the key part standard position information to obtain the weight loss position information that needs to be lost.

[0128] Specifically, the portrait image is input into a deep learning model to obtain the position information of key parts of the portrait image. The device also includes:

[0129] The acquisition module is used to obtain portrait images in the history live class as a sample set;

[0130] A labeling module, used to label the key position information of the sample set images;

[0131] A training module is used to train a deep learning model based on the sample set and the key part position information.

[0132] In a specific embodiment, the weight loss processing module 54 uses image distortion to map the pixels of the weight loss position information to a preset range to obtain a weight loss portrait image.

[0133] The fusion module 55 includes:

[0134] A pasting module, configured to paste the slimmed-down portrait image to a position other than the blackboard writing in the second scene video image;

[0135] The sub-processing module is used to process the pasting boundary between the portrait image and the second scene video image by using a Gaussian aliasing algorithm.

[0136] Furthermore, the device further comprises:

[0137] The graphics processing module is used to process the fused image using the graphics lighting model.

[0138] In a specific embodiment, the device further comprises:

[0139] The switching module 57 is configured to switch the second scene video picture to a corresponding scene picture when a scene switching instruction is received.

[0140] The prompt module 58 is used to prompt key knowledge information when receiving a key knowledge prompt instruction.

[0141] The present invention captures at least two scene videos of a live class in real time, wherein one scene video is composed of a first scene video image containing a portrait, and the other scene video is composed of a second scene video image containing blackboard writing. The present invention extracts the portrait image from the first scene video image, locates the slimming position information of the portrait image, and then performs slimming processing on the slimming position information to achieve the slimming of the portrait in the live class. The slimmed-down portrait image is then fused with the second scene video image, and the fused video image is output. This achieves the effect of slimming the portrait in the live class in real time while ensuring that the blackboard writing does not deform.

[0142] Those skilled in the art will appreciate that the modules in the above device embodiments may be distributed in the device as described, or may be modified accordingly and distributed in one or more devices different from the above embodiments. The modules in the above embodiments may be combined into one module or further split into multiple submodules.

[0143] Figure 6It is a structural diagram of an electronic device according to an embodiment of the present invention. The electronic device includes a processor and a memory. The memory is used to store a computer executable program. When the computer program is executed by the processor, the processor executes a method for processing portrait slimming in a live class.

[0144] like Figure 6 As shown, the electronic device is implemented as a general-purpose computing device. The processor may be one or multiple processors working in concert. The present invention also does not exclude distributed processing, meaning that the processors may be dispersed across different physical devices. The electronic device of the present invention is not limited to a single entity but may also be the sum of multiple physical devices.

[0145] The memory stores a computer executable program, typically a machine-readable code, which can be executed by the processor to enable the electronic device to perform the method of the present invention, or at least some of the steps in the method.

[0146] The memory includes a volatile memory, such as a random access memory unit (RAM) and / or a cache memory unit, and may also be a non-volatile memory, such as a read-only memory unit (ROM).

[0147] Optionally, in this embodiment, the electronic device further includes an I / O interface for exchanging data with an external device. The I / O interface may represent one or more of several types of bus structures, including a storage unit bus or storage unit controller, a peripheral bus, an accelerated graphics port, a processing unit, or a local bus using any of a variety of bus structures.

[0148] It should be understood that Figure 6 The electronic device shown is merely an example of the present invention. The electronic device of the present invention may also include elements or components not shown in the above examples. For example, some electronic devices also include display units such as screens, and some electronic devices also include human-computer interaction elements such as buttons and keyboards. As long as the electronic device can execute a computer-readable program stored in its memory to implement the method of the present invention or at least some of the steps of the method, it is considered an electronic device covered by the present invention.

[0149] Figure 7 FIG is a schematic diagram of a computer readable recording medium according to an embodiment of the present invention. Figure 7As shown, a computer executable program is stored in a computer-readable recording medium, and when the computer executable program is executed, the above-mentioned live class portrait slimming processing method of the present invention is implemented. The computer-readable storage medium may include a data signal propagated in the baseband or as part of a carrier, which carries a readable program code. This propagated data signal can take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The readable storage medium can also be any readable medium other than a readable storage medium, which can send, propagate or transmit a program for use by or in combination with an instruction execution system, device or device. The program code contained on the readable storage medium can be transmitted using any appropriate medium, including but not limited to wireless, wired, optical cable, RF, etc., or any suitable combination of the above.

[0150] The program code for performing the operations of the present invention may be written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Java, C++, and the like, as well as conventional procedural programming languages ​​such as "C" or similar programming languages. The program code may be executed entirely on the user computing device, partially on the user device, as a stand-alone software package, partially on the user computing device and partially on a remote computing device, or entirely on a remote computing device or server. In cases involving a remote computing device, the remote computing device may be connected to the user computing device via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computing device (e.g., via the Internet using an Internet service provider).

[0151] Through the above description of the embodiments, it is easy for those skilled in the art to understand that the present invention can be implemented by hardware capable of executing specific computer programs, such as the system of the present invention, and the electronic processing unit, server, client, mobile phone, control unit, processor, etc. contained in the system. The present invention can also be implemented by a vehicle containing at least a part of the above-mentioned system or components. The present invention can also be implemented by computer software that executes the method of the present invention, such as control software executed by the microprocessor, electronic control unit, client, server, etc. of the live broadcast device. However, it should be noted that the computer software that executes the method of the present invention is not limited to being executed by one or a specific hardware entity. It can also be implemented in a distributed manner by unspecified specific hardware. For computer software, the software product can be stored in a computer-readable storage medium (which can be a CD-ROM, USB flash drive, mobile hard disk, etc.), or it can be distributed and stored on the network, as long as it can enable the electronic device to execute the method according to the present invention.

[0152] The specific embodiments described above further illustrate the objectives, technical solutions, and beneficial effects of the present invention. It should be understood that the present invention is not inherently related to any specific computer, virtual device, or electronic device, and various general-purpose devices can also implement the present invention. The above description is only a specific embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention shall be included in the scope of protection of the present invention.

Claims

1. A method for processing slimming portraits in live classes, characterized in that: The method comprises the following steps: Real-time video capture of the first scene containing portraits and the second scene containing blackboard writing in the live class; Extracting a portrait image from the video image of the first scene; Extracting key part position information and key part height information from the portrait image, including: inputting the portrait image into a deep learning model to obtain the key part position information and key part height information of the portrait image; determining standard key part position information based on the key part height information based on a standard height distribution, comparing the key part position information with the key part standard position information, and locating slimming position information of a part in the portrait image that needs to be slimmed down, the slimming position information including coordinates of the slimming part and the slimming length; The coordinates of the slimming part are converted into slimming part pixels and the slimming length is converted into corresponding slimming part pixels according to the coordinate-pixel conversion relationship of the first scene video image, and the slimming part pixels are mapped into a preset slimming pixel range using an image distortion method to perform slimming processing on the slimming position information, thereby obtaining a slimming portrait image; The slimmed-down portrait image is pasted to a position other than the blackboard writing in the second scene video image, and the pasting boundary of the slimmed-down portrait image and the second scene video image is merged with the second scene video image by using a Gaussian aliasing algorithm, including: taking the slimmed-down portrait image and the second scene video image as images to be fused, respectively constructing a corresponding Gaussian residual pyramid for each image to be fused and retaining the top image of each Gaussian residual pyramid downsampled; using the binary mask image of each layer of the Gaussian residual pyramid constructed by sampling under a binary mask to merge the images of the corresponding layers of the Gaussian residual pyramid of the image to be fused into one image; performing Gaussian upsampling on the image after merging the top image, and adding the image obtained by Gaussian upsampling to an image of the corresponding layer to reconstruct the image of the layer, repeating the reconstruction until the lowest image is reconstructed and outputting it; Output the fused video image.

2. The method for processing slimming portraits in live classes according to claim 1, characterized in that: Before inputting the portrait image into the deep learning model to obtain the key part position information of the portrait image, the method further includes: Get portrait images from live history classes as a sample set; Marking the key parts of the sample set images; A deep learning model is trained based on the sample set and the key part position information.

3. The method for processing a slimming portrait in a live class according to any one of claims 1-2, characterized in that: After fusing the slimmed-down portrait image with the second scene video image, the method further includes: The fused image is processed using a graphics illumination model to eliminate illumination inconsistencies between the portrait image and the second scene video image.

4. The method for processing a slimming portrait in a live class according to any one of claims 1-2, characterized in that: Also includes: When a scene switching instruction is received, switching the second scene video picture to the corresponding scene picture; and / or, When a key knowledge prompt instruction is received, key knowledge information is prompted.

5. The method for processing slimming portraits in live classes according to claim 3, characterized in that: Also includes: When a scene switching instruction is received, switching the second scene video picture to the corresponding scene picture; and / or, When a key knowledge prompt instruction is received, key knowledge information is prompted.

6. A device for processing slimming portraits in live classes, characterized in that: The device comprises: The acquisition module is used to respectively acquire in real time the first scene video image containing a portrait and the second scene video image containing a blackboard image in the live class; An extraction module, configured to extract a portrait image from the video image of the first scene; a positioning module for extracting key part position information and key part height information from the portrait image, comprising: inputting the portrait image into a deep learning model to obtain the key part position information and key part height information of the portrait image; determining standard key part position information based on the key part height information based on a standard height distribution, comparing the key part position information with the key part standard position information, and locating weight loss position information of the person in the portrait image that needs to be reduced; A slimming processing module, configured to perform slimming processing on the slimming position information to obtain a slimming portrait image; a fusion module, configured to paste the slimmed-down portrait image to a position other than the blackboard writing in the second scene video image, and to fuse the slimmed-down portrait image with the second scene video image by using a Gaussian blending algorithm to separate the pasting boundary between the slimmed-down portrait image and the second scene video image; The output module is used to output the fused video image.

7. An electronic device comprising a processor and a memory, wherein the memory is used to store a computer executable program, wherein: When the computer-implemented program is executed by the processor, the processor performs the method according to any one of claims 1 to 5.

8. A computer-readable medium storing a computer-executable program, characterized in that: When the computer executable program is executed, the method according to any one of claims 1 to 5 is implemented.

Citation Information

Patent Citations

  • Stature adjustment method and device based on image data and computing equipment

    CN107977927A