Denoising method, device and equipment for video conference portrait zooming

By using pre-trained models to identify the scaling scale features and noise feature maps in video conferences, and dynamically adjusting the denoising intensity, the problem of poor denoising effect during portrait scaling in video conferences is solved, achieving more efficient denoising and detail retention.

CN120013797APending Publication Date: 2025-05-16YEALINK (XIAMEN) NETWORK TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510063806.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-15
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

In the prior art, it is difficult to dynamically adjust the denoising intensity when zooming in portraits in video conferences, resulting in noise amplification meetings for image amplification, and the details of portraits are erased as noise when image is reduced.

Method used

By using a pretrained model in a video conference, the scaling scale features of the first frame image are recognized in real time, and the noise feature map is obtained in combination with the features of the previous frame image, thereby dynamically adjusting the denoising intensity.

Benefits of technology

Dynamic adjustment of the denoising intensity during portrait scaling in video conferencing is achieved, which improves the denoising effect, retains important details, and enhances the correlation between continuous frames and improves stability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120013797A_ABST
    Figure CN120013797A_ABST
Patent Text Reader

Abstract

The invention discloses a denoising method, device and equipment for video conference portrait zooming. The method comprises the following steps: inputting a first frame image which is currently acquired in real time into a pre-trained first model, so that the first model identifies a zoom scale feature of the first frame image and a first feature of a previous frame image of the first frame image, and obtains a first feature of the first frame image; the first feature is a noise feature map extracted from an input image through the first model; and superposing the first frame image with the first feature of the first frame image to obtain a denoised image of the first frame image. According to the method and the device, the denoising intensity of the image after portrait zooming can be dynamically adjusted in the video conference scene, and the denoising effect in the video conference scene is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of image processing, and in particular to a denoising method, device and equipment for portrait zooming in video conferences. Background Art

[0002] During a video conference, it is usually necessary to scale the portraits of the people in each current video frame, such as scaling up the portrait of the speaker or scaling down the portrait of the person who is not speaking.

[0003] In order to achieve the above functions, some existing technologies use hardware control to control the camera to zoom and achieve portrait zooming. However, due to cost constraints, the above technologies sometimes cannot use zoom lenses to achieve automatic zooming of portraits, and need to use fixed-focus lenses for digital zooming, which causes the noise in the enlarged image to be amplified at the same time. Some technologies use deep learning video denoising methods, but the current methods dynamically adjust the denoising intensity under different noise intensities. If the same set of parameters is used, it is difficult to apply it to close-range scenes and long-range scenes at the same time. When the parameter settings are unreasonable, it will result in the noise not being removed when the image is enlarged or the portrait details being erased as noise when the image is reduced.

[0004] Therefore, how to dynamically adjust the denoising intensity when zooming in on-screen portraits in video conferences to improve the denoising effect is a technical problem that needs to be solved at present. Summary of the invention

[0005] The present application provides a denoising method, device and equipment for portrait zooming in video conferencing to solve the problem of how to dynamically adjust the denoising intensity when zooming in the portrait in the video conferencing. Improving the denoising effect is a technical problem that needs to be solved at present.

[0006] In order to solve the above technical problems, in a first aspect, an embodiment of the present application provides a denoising method for portrait zooming in a video conference, comprising:

[0007] Inputting the first frame image currently collected in real time into the pre-trained first model, so that the first model can identify the scaling feature of the first frame image, and obtain the first feature of the first frame image by combining the first feature of the previous frame image of the first frame image; the first feature is a noise feature map extracted from the input image by the first model;

[0008] The first frame image is superimposed with the first feature of the first frame image to obtain a denoised image of the first frame image.

[0009] Compared with the prior art, the embodiments of the present application have the following beneficial effects: since the zoom scale of the first frame image currently collected in real time will affect the setting of the denoising intensity of the subsequent denoising process, the zoom scale of the first frame image is firstly identified through the first model to obtain the zoom scale feature, and further according to the zoom scale feature, the extraction process of the noise feature map is regulated, thereby ensuring that when the noise feature map obtained subsequently is used to denoise the first frame image, the noise is accurately denoised according to the zoom scale of the first frame image, thereby improving the accuracy of picture denoising during video conferencing.

[0010] In some embodiments of the first aspect of the present application, enabling the first model to identify the scaling feature of the first frame image and combining the first feature of a previous frame image of the first frame image to obtain the first feature of the first frame image includes:

[0011] The first model includes a first branch, a second branch and a first convolutional layer;

[0012] Extracting a second feature of the first frame image through the first convolutional layer;

[0013] extracting the scaled feature from the second feature through the first branch;

[0014] The second branch combines the zoom scale feature, the second feature, and the first feature of the previous frame image to obtain the first feature of the first frame image.

[0015] Compared with the prior art, the above embodiment has the following beneficial effects: by setting two branches in the first model for parallel data processing, and identifying the zoom scale of the first frame image currently input into the first model through the first branch, obtaining the zoom scale feature, and dynamically adjusting the denoising intensity of the first frame image according to the zoom scale identified by the first branch through the second branch, not only the efficiency of data processing is improved, but also the accuracy of denoising is improved.

[0016] In some embodiments of the first aspect of the present application, extracting the scaling feature from the second feature through the first branch includes:

[0017] Inputting the second feature into a second convolutional layer of the same scale as the first convolutional layer to obtain a third feature;

[0018] The third feature is sequentially input into a plurality of third convolutional layers to compress the size of the third feature and obtain the scaled feature.

[0019] Compared with the prior art, the above embodiment has the following beneficial effects: after obtaining the third feature by performing shallow feature extraction on the second feature through the second convolutional layer, the third feature size is compressed again through several third convolutional layers, and the depth of the feature map corresponding to the third feature is increased, so that the first branch focuses on the extraction of the portrait scaling information in the first frame image, thereby improving the accuracy of the scaling feature and further improving the accuracy of the scaling evaluation of the first frame image.

[0020] In some embodiments of the first aspect of the present application, the acquiring, by the second branch, the first feature of the first frame image by combining the zoom scale feature, the second feature, and the first feature of the previous frame image includes:

[0021] Inputting the second feature into a plurality of residual modules in sequence to obtain a fourth feature;

[0022] Inputting the fourth feature and the scaling feature into a feature control module simultaneously to obtain a fifth feature;

[0023] The fifth feature is concatenated with the first feature of the previous frame image and the concatenated features are sequentially inputted into a plurality of residual modules to obtain a sixth feature;

[0024] The sixth feature and the scaling feature are simultaneously input into the feature control module to obtain the first feature of the first frame image.

[0025] Compared with the prior art, the above embodiment has the following beneficial effects: since the second branch is mainly used to denoise the first frame image according to the scaling features obtained by the first branch, it undertakes the main computing tasks. Through multi-layer convolution and residual modules, the efficiency of feature extraction and processing can be improved, thereby solving the real-time dynamic denoising problem in the video conferencing scenario; further, through the feature control module, the first branch is connected to the second branch, so that when the second branch performs the denoising task, it combines the scaling features provided by the first branch to dynamically adjust its own denoising strength, improve the denoising effect, and retain important details in the first frame image; at the same time, when the second branch processes the current frame noise, it refers to the noise feature map extracted from the previous frame image, and helps to enhance the correlation between consecutive frames through the transmission of timing information, so that the model is more stable when performing denoising tasks in the video conferencing scenario.

[0026] In some embodiments of the first aspect of the present application, comprising:

[0027] The residual module includes a plurality of fourth convolutional layers;

[0028] The feature control module includes a fifth convolutional layer and a plurality of residual modules;

[0029] Inputting the scaled feature into the fifth convolutional layer to obtain a seventh feature;

[0030] Inputting the seventh feature into each of the residual modules in the feature control module respectively, and obtaining corresponding first parameters respectively;

[0031] According to each of the first parameters and the fourth feature or the sixth feature, an output of the corresponding feature control module is obtained.

[0032] Compared with the prior art, the above embodiment has the following beneficial effects: a first parameter for adjusting the denoising intensity is generated according to the scaling feature of the first branch, and the denoising intensity of the fourth feature or the sixth feature extracted from the first frame image is further adjusted according to the first parameter, thereby realizing dynamic adjustment of the denoising intensity, improving the denoising effect, and retaining important details in the first frame image.

[0033] In some embodiments of the first aspect of the present application, the pre-training step of the first model includes:

[0034] Performing data preprocessing on the first video data that has not been de-noised to obtain a training data set;

[0035] Constructing corresponding first loss functions for the first branch and the second branch respectively, and obtaining a second loss function of the first model according to each of the first loss functions;

[0036] Pre-train the first model according to the second loss function and the training data set.

[0037] Compared with the prior art, the above embodiment has the following beneficial effects: by designing different loss functions for the first branch and the second branch respectively, and optimizing the model pre-training process based on these loss functions, the parameters in the first model can be adjusted more finely, thereby improving the model's ability to learn scaling features and noise features, while avoiding the optimization conflicts caused by a single loss function, making the model obtained by the final training more efficient and accurate.

[0038] In some embodiments of the first aspect of the present application, performing data preprocessing on the first video data that has not been denoised to obtain a training data set includes:

[0039] Processing the first video data by time series mean filtering to obtain a second frame of image;

[0040] The second frame image is copied a first preset number of times to obtain second video data having the same number of frames as the first video data;

[0041] Data enhancement is performed on the first video data and the second video data to obtain a training data set.

[0042] Compared with the prior art, the above embodiment has the following beneficial effects: by preprocessing the un-denoised video data through temporal mean filtering and data enhancement technology to generate a training data set, not only the diversity of the training data is increased, but also the first model can better learn the temporal changes in the video, thereby helping to improve the generalization ability of the model, effectively deal with noise interference in different video scenes, and improve the denoising effect.

[0043] In some embodiments of the first aspect of the present application, constructing corresponding first loss functions for the first branch and the second branch respectively, and obtaining a second loss function of the first model according to each of the first loss functions, includes:

[0044] The loss function of the first model is specifically:

[0045]

[0046] in, is the loss value corresponding to the first model; L2(R,R_true) is the first loss function of the first branch, representing the L2 norm of R and R_true, R is the scaling scale corresponding to the scaling scale feature obtained through the first branch; R_true is the true scaling scale corresponding to the test data; λ1 is the weight coefficient; L1(Y,GT) is the first loss function of the second branch, representing the L1 norm of Y and GT, Y is the denoised image obtained according to the noise feature map, and GT is the true denoised image corresponding to the test data.

[0047] In a second aspect, an embodiment of the present application further provides a denoising device for portrait zooming in a video conference, comprising: a noise extraction module and a denoising module;

[0048] The noise extraction module is used to input the first frame image currently collected in real time into the pre-trained first model, so that the first model can identify the scaling feature of the first frame image, and obtain the first feature of the first frame image in combination with the first feature of the previous frame image of the first frame image; the first feature is a noise feature map extracted from the input image by the first model;

[0049] The denoising module is used to superimpose the first frame image with the first feature of the previous frame image to obtain a denoised image of the first frame image.

[0050] In a third aspect, the present application also provides a terminal device, comprising a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, wherein the processor implements the above-mentioned denoising method for portrait zooming in video conferencing when executing the computer program. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] Figure 1 A flowchart of a denoising method for portrait zooming in a video conference provided in some embodiments of the present application;

[0052] Figure 2 A schematic diagram of the structure of a first model provided in some embodiments of the present application;

[0053] Figure 3 A schematic diagram of the structure of a denoising device for portrait zooming in a video conference provided in some embodiments of the present application. DETAILED DESCRIPTION

[0054] In order to solve the denoising problem in the context of portrait zooming, some existing technologies use hardware control to control the camera to zoom and achieve portrait zooming. However, due to cost constraints, the above technologies sometimes cannot use zoom lenses to achieve automatic zooming of portraits, and need to use fixed-focus lenses for digital zooming, which causes the noise in the enlarged image to be amplified at the same time. Some technologies use deep learning video denoising methods, but the current methods dynamically adjust the denoising intensity under different noise intensities. If the same set of parameters is used, it is difficult to apply it to close-range scenes and long-range scenes at the same time. When the parameter settings are unreasonable, it will result in the noise not being removed when the image is enlarged or the portrait details being erased as noise when the image is reduced.

[0055] In order to solve the above technical problems, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0056] Embodiment 1

[0057] Please refer to Figure 1 , a denoising method for portrait zooming in a video conference provided in an embodiment of the present application, including S10 to S20, specifically:

[0058] S10: Input the first frame image currently captured in real time into the pre-trained first model, so that the first model can recognize the scaling feature of the first frame image, and obtain the first feature of the first frame image in combination with the first feature of the previous frame image of the first frame image; the first feature is a noise feature map extracted from the input image by the first model.

[0059] refer to Figure 2, is a schematic diagram of a first model provided in some embodiments of the present application, wherein the first model is a single-loop structure model, wherein LQ i-1 is the first frame image, LQ i is the next frame image of the first frame image. It can be seen that LQ i-1 Some features of LQ are concatenated with i Some features of the video are spliced ​​together, and so on, to complete the transmission of the entire video link information.

[0060] Further, in some embodiments of the present application, enabling the first model to identify the scaling feature of the first frame image and combining the first feature of the previous frame image of the first frame image to obtain the first feature of the first frame image includes:

[0061] The first model includes a first branch, a second branch and a first convolutional layer;

[0062] Extracting a second feature of the first frame image through the first convolutional layer;

[0063] extracting the scaled feature from the second feature through the first branch;

[0064] The second branch combines the zoom scale feature, the second feature, and the first feature of the previous frame image to obtain the first feature of the first frame image.

[0065] Preferably, reference Figure 2 In some embodiments of the present application, the first model includes two branches, in which LQ i-1 Before inputting the first model and entering the two branches, a first convolution layer is first passed to extract the shallowest features and obtain the second features; wherein the second features are shared features of the first branch and the second branch, and are used for synchronous transmission in the first branch and the second branch.

[0066] Preferably, reference Figure 2 In some embodiments of the present application, the first branch performs three convolution operations to obtain a scale feature (R_feature), and the scale feature is further input to the fully connected layer to obtain a noise scale factor R, where the noise scale factor R can represent LQ i-1 In addition, the purpose of obtaining the scaling factor R through the fully connected layer here includes facilitating the subsequent loss value calculation of the first branch.

[0067] Preferably, reference Figure 2In some embodiments of the present application, after the first branch obtains the zoom scale feature, the zoom scale feature is input into two feature control modules (Mod Block) in the second branch, and combined with the second feature and the previous frame image LQ i-2 The first feature of the first frame image is obtained by taking the noise feature map of the first frame image. The second feature is Figure 2 LQ i-1 The feature data first input to the residual module (Res Block) in the second branch is the shallowest feature extracted through the first convolutional layer; the previous frame image LQ i-2 The first characteristic is LQ i-2 The noise feature map output by the last feature control module of the second branch.

[0068] It can be seen from the above embodiments that the present application sets two branches in the first model for parallel data processing, and identifies the zoom scale of the first frame image currently input into the first model through the first branch to obtain the zoom scale feature, and dynamically adjusts the denoising intensity of the first frame image according to the zoom scale identified by the first branch through the second branch, which not only improves the efficiency of data processing, but also improves the accuracy of denoising.

[0069] Further, in some embodiments of the present application, extracting the scaling feature from the second feature through the first branch includes:

[0070] Inputting the second feature into a second convolutional layer of the same scale as the first convolutional layer to obtain a third feature;

[0071] The third feature is sequentially input into a plurality of third convolutional layers to compress the size of the third feature and obtain the scaled feature.

[0072] Preferably, reference Figure 2 In some embodiments of the present application, the first branch extracts the scaling feature by the following operation: the second feature is first input into the second convolution layer with the same scale as the first convolution layer, the shared feature is converted into the third feature exclusive to the first branch, and then the third feature is sequentially subjected to two convolution operations with a step length of 2 to further compress the size of the third feature, while increasing the feature map depth to obtain the scaling feature. The above operation can make the first branch focus on the extraction of the scaling information of the portrait in the first frame image, thereby improving the accuracy of the scaling feature and further improving the accuracy of the scaling evaluation of the first frame image.

[0073] Further, in some embodiments of the present application, the step of obtaining the first feature of the first frame image by combining the zoom scale feature, the second feature, and the first feature of the previous frame image through the second branch includes:

[0074] Inputting the second feature into a plurality of residual modules in sequence to obtain a fourth feature;

[0075] Inputting the fourth feature and the scaling feature into a feature control module simultaneously to obtain a fifth feature;

[0076] The fifth feature is concatenated with the first feature of the previous frame image and the concatenated features are sequentially inputted into a plurality of residual modules to obtain a sixth feature;

[0077] The sixth feature and the scaling feature are simultaneously input into the feature control module to obtain the first feature of the first frame image.

[0078] Preferably, reference Figure 2 In some embodiments of the present application, the second branch consists of two parts. The first part is the shallow feature extraction part before the splicing operation, through which the fourth feature and the fifth feature are obtained; the second part is the deep feature extraction part after the splicing operation, through which the sixth feature and the noise feature map are obtained.

[0079] Preferably, reference Figure 2 In some embodiments of the present application, when performing shallow feature extraction, the fourth feature and the fifth feature are mainly obtained through the following operations: first, the second feature is input into N residual modules (Res Block) composed of two convolutional layers to perform preliminary shallow feature extraction and obtain the fourth feature; then, the fourth feature is input together with the scaling feature into the first feature control module (Mod Block) of the second branch to control the shallow fourth feature and obtain the fifth feature.

[0080] Preferably, reference Figure 2 In some embodiments of the present application, when performing deep feature extraction, the sixth feature and the noise feature map are mainly obtained by the following operations: LQ i-1 Zhong and LQ i-2 The first characteristic, namely LQ i-2 The feature data obtained by splicing the noise feature map is input again into a convolution layer with the same scale as the first convolution layer for deep feature extraction; then the feature is input into N residual modules consisting of two convolution layers to perform further deep feature extraction to obtain the sixth feature; finally, the sixth feature is input into the second feature control module of the second branch together with the scaled feature to control the deep sixth feature to obtain the final noise feature map, that is, the first feature of the first frame image.

[0081] Preferably, in some embodiments of the present application, if the first frame image is the first frame image, the first feature of the previous frame image is set to 0.

[0082] Preferably, in some embodiments of the present application, the first frame image can be an image captured by any image acquisition device such as any camera or camera, and single-frame image denoising is achieved through the first model of the present application; in addition, the first frame image can also be a video frame image in the video data captured by any of the above-mentioned image acquisition devices, and the video frames in the video data are sequentially input into the first model of the present application to achieve dynamic and real-time denoising of the video data.

[0083] Since the second branch is mainly used to denoise the first frame image according to the scaling features obtained by the first branch, it undertakes the main computing tasks. Through multi-layer convolution and residual modules, the efficiency of feature extraction and processing can be improved, thereby solving the real-time dynamic denoising problem in video conferencing scenarios; further, through the feature control module, the first branch is connected to the second branch, so that when the second branch performs the denoising task, it combines the scaling features provided by the first branch to dynamically adjust its own denoising strength, improve the denoising effect, and retain important details in the first frame image; at the same time, when the second branch processes the current frame noise, it refers to the noise feature map extracted from the previous frame image, and helps to enhance the correlation between consecutive frames through the transmission of timing information, so that the model is more stable when performing denoising tasks in video conferencing scenarios.

[0084] Furthermore, in some embodiments of the present application, it also includes:

[0085] The residual module includes a plurality of fourth convolutional layers;

[0086] The feature control module includes a fifth convolutional layer and a plurality of residual modules;

[0087] Inputting the scaled feature into the fifth convolutional layer to obtain a seventh feature;

[0088] Inputting the seventh feature into each of the residual modules in the feature control module respectively, and obtaining corresponding first parameters respectively;

[0089] According to each of the first parameters and the fourth feature or the sixth feature, an output of the corresponding feature control module is obtained.

[0090] Preferably, reference Figure 2 In some embodiments of the present application, the residual module is composed of two fourth convolutional layers and a splicing module, and the input of the residual module is added to the features output by the two fourth convolutional layers to obtain the final output result (i.e., the fourth feature or the sixth feature).

[0091] Preferably, reference Figure 2In some embodiments of the present application, the feature control module is composed of a fifth convolutional layer and two residual modules. After receiving the scaled feature, the fifth convolutional layer outputs the seventh feature. Then, the seventh feature is input into the two residual modules respectively to obtain the first parameters gamma and beta. Gamma is used to scale the input feature, while beta is used to offset the input feature, so that the first model can adjust the denoising ability according to different noise intensities or magnifications, making the denoising process more refined. Finally, the feature control result is obtained by the formula output = gamma * input + beta, and input is the fourth feature or the sixth feature of the input. Through the feature control module, the dynamic adjustment of the denoising intensity is realized, the denoising effect is improved, and the important details in the first frame image are retained.

[0092] S20: Superimposing the first frame image and the noise feature map to obtain a denoised image of the first frame image.

[0093] Preferably, in some embodiments of the present application, after obtaining the noise characteristic map, the first frame image LQ i-1 Superimpose with the noise feature map to obtain the first frame image LQ i-1 The corresponding denoised image Y i-1 .

[0094] Preferably, reference Figure 2 In some embodiments of the present application, the superposition operation of the first frame image and the noise feature map can be embedded in the first model to achieve LQ i-1 Input to the first model, directly output LQ through the first model i-1 The corresponding denoised image Y i-1 .

[0095] Furthermore, in some embodiments of the present application, the pre-training step of the first model includes:

[0096] Performing data preprocessing on the first video data that has not been de-noised to obtain a training data set;

[0097] Constructing corresponding first loss functions for the first branch and the second branch respectively, and obtaining a second loss function of the first model according to each of the first loss functions;

[0098] Pre-train the first model according to the second loss function and the training data set.

[0099] Further, in some embodiments of the present application, the step of performing data preprocessing on the first video data that has not been denoised to obtain a training data set includes:

[0100] Processing the first video data by time series mean filtering to obtain a second frame of image;

[0101] The second frame image is copied a first preset number of times to obtain second video data having the same number of frames as the first video data;

[0102] Data enhancement is performed on the first video data and the second video data to obtain a training data set.

[0103] Preferably, in some embodiments of the present application, the training data set can be obtained by the following steps: turn off the denoising function of the image signal processor (ISP) in the video conference camera device, collect a number of first video data under different lighting conditions and different conference scenes through the ISP, and use the first video data as a low-quality image (Low Quality Image, LQ), each LQ has the same number of video frames n. Further, according to the following steps, use the collected LQ to obtain a high-quality image (Ground Truth, GT) as a positive sample in the training data set, that is, a real denoised image:

[0104] (1) For the static position, the temporal mean filter is used to process the LQ, remove the dynamic noise, and obtain the single-frame GT (i.e., the second frame image);

[0105] (2) Copying the single frame GT n times to obtain the second video data GT of the same length as LQ, so that LQ and GT form paired video frames of the same length;

[0106] (3) Perform data enhancement on the paired video frames according to the following formula:

[0107] Image_sub=[RS(RC(Crop,Resize))] Image

[0108] Among them, Crop is a random cropping operation; Resize is a random up-sampling and down-sampling operation; RC(*) represents the random selection of one or two of the above two operations; RS(*) represents the random sorting of different operations; Image refers to the paired video frames of LQ and GT; Image_sub refers to the sub-image of the paired video frames of LQ and GT after data enhancement. Because it is necessary to ensure that the data enhancement method of the same paired fragment is the same, Image is used instead of the two symbols LQ and GT.

[0109] Preferably, in some embodiments of the present application, when performing data enhancement on the paired video frames, it is necessary to simultaneously record the image scale scaling factor corresponding to the Resize operation as the noise scale factor R_true output by the subsequent first model.

[0110] Preferably, in some embodiments of the present application, constructing corresponding first loss functions for the first branch and the second branch respectively, and obtaining the second loss function of the first model according to each of the first loss functions, includes:

[0111] The loss function of the first model is specifically:

[0112]

[0113] in, is the loss value corresponding to the first model; L2(R,R_true) is the first loss function of the first branch, representing the L2 norm of R and R_true, R is the scaling scale corresponding to the scaling scale feature obtained through the first branch; R_true is the true scaling scale corresponding to the test data; λ1 is the weight coefficient; L1(Y,GT) is the first loss function of the second branch, representing the L1 norm of Y and GT, Y is the denoised image obtained according to the noise feature map, and GT is the true denoised image corresponding to the test data.

[0114] From the loss function of the first model mentioned above, it can be seen that L2(R, R_true) is designed to optimize the scaling feature by optimizing the noise scale factor output by the first branch in the first model, and L1(Y, GT) is designed to optimize the denoising effect of the second branch in the first model. This application designs different loss functions for the first branch and the second branch respectively, and optimizes the model pre-training process based on these loss functions, so as to more finely adjust the parameters in the first model, thereby improving the model's ability to learn scaling features and noise features, while also avoiding the optimization conflict caused by a single loss function, making the model obtained by the final training more efficient and accurate.

[0115] In summary, the denoising method for portrait zooming in video conferencing provided in an embodiment of the present application has the following beneficial effects: since the zoom scale of the first frame image currently collected in real time will affect the setting of the denoising intensity of the subsequent denoising process, the zoom scale of the first frame image is first identified by the first model to obtain the zoom scale feature, and further according to the zoom scale feature, the extraction process of the noise feature map is regulated, thereby ensuring that when the noise feature map obtained subsequently is used to denoise the first frame image, the noise is accurately denoised according to the zoom scale of the first frame image, thereby improving the accuracy of picture denoising during video conferencing.

[0116] Embodiment 2

[0117] refer to Figure 3 , a denoising device for portrait zooming in a video conference provided in an embodiment of the present application, includes: a noise extraction module 11 and a denoising module 12.

[0118] Furthermore, in some embodiments of the present application, the noise extraction module 11 is used to input the first frame image currently collected in real time into the pre-trained first model, so that the first model can recognize the scaling feature of the first frame image, and obtain the first feature of the first frame image in combination with the first feature of the previous frame image of the first frame image; the first feature is the noise feature map extracted from the input image by the first model; the denoising module 12 is used to superimpose the first frame image with the first feature of the first frame image to obtain a denoised image of the first frame image.

[0119] Further, in some embodiments of the present application, the first model is used to identify the scaling feature of the first frame image, and to obtain the first feature of the first frame image in combination with the first feature of the previous frame image of the first frame image, including: the first model includes a first branch, a second branch, and a first convolutional layer; the second feature of the first frame image is extracted through the first convolutional layer; the scaling feature is extracted from the second feature through the first branch; and the first feature of the first frame image is obtained by combining the scaling feature, the second feature, and the first feature of the previous frame image through the second branch.

[0120] Furthermore, in some embodiments of the present application, extracting the scaled feature from the second feature through the first branch includes: inputting the second feature into a second convolutional layer of the same scale as the first convolutional layer to obtain a third feature; and inputting the third feature into several third convolutional layers in sequence to compress the size of the third feature to obtain the scaled feature.

[0121] Furthermore, in some embodiments of the present application, the first feature of the first frame image is obtained by combining the zoom scale feature, the second feature and the first feature of the previous frame image through the second branch, including: inputting the second feature into several residual modules in sequence to obtain a fourth feature; inputting the fourth feature and the zoom scale feature into a feature control module at the same time to obtain a fifth feature; splicing the fifth feature with the first feature of the previous frame image and inputting them into several residual modules in sequence to obtain a sixth feature; inputting the sixth feature and the zoom scale feature into the feature control module at the same time to obtain the first feature of the first frame image.

[0122] Furthermore, in some embodiments of the present application, it includes: the residual module includes several fourth convolutional layers; the feature control module includes a fifth convolutional layer and several residual modules; the scaled feature is input into the fifth convolutional layer to obtain a seventh feature; the seventh feature is respectively input into each residual module in the feature control module to obtain corresponding several first parameters; according to each first parameter and the fourth feature or the sixth feature, the output of the corresponding feature control module is obtained.

[0123] Furthermore, in some embodiments of the present application, the pre-training step of the first model includes: performing data preprocessing on the first video data that has not been denoised to obtain a training data set; constructing corresponding first loss functions for the first branch and the second branch, respectively, and obtaining the second loss function of the first model according to each of the first loss functions; and pre-training the first model according to the second loss function and the training data set.

[0124] Furthermore, in some embodiments of the present application, the data preprocessing is performed on the first video data that has not been denoised to obtain a training data set, including: processing the first video data through time series mean filtering to obtain a second frame image; copying the second frame image a first preset number of times to obtain second video data with the same number of frames as the first video data; and performing data enhancement on the first video data and the second video data to obtain a training data set.

[0125] Further, in some embodiments of the present application, constructing corresponding first loss functions for the first branch and the second branch respectively, and obtaining the second loss function of the first model according to each of the first loss functions, includes:

[0126] The loss function of the first model is specifically:

[0127]

[0128] in, is the loss value corresponding to the first model; L2(R,R_true) is the first loss function of the first branch, representing the L2 norm of R and R_true, R is the scaling scale corresponding to the scaling scale feature obtained through the first branch; R_true is the true scaling scale corresponding to the test data; λ1 is the weight coefficient; L1(Y,GT) is the first loss function of the second branch, representing the L1 norm of Y and GT, Y is the denoised image obtained according to the noise feature map, and GT is the true denoised image corresponding to the test data.

[0129] It can be understood that the above-mentioned device item embodiments correspond to the method item embodiments of the present invention. The denoising device for video conferencing portrait scaling provided by the embodiments of the present invention can implement any method item embodiment of the present invention, that is, the denoising method for video conferencing portrait scaling provided in Embodiment 1.

[0130] In summary, the denoising device for portrait zooming in video conferencing provided in an embodiment of the present application has the following beneficial effects: since the zoom scale of the first frame image currently collected in real time will affect the setting of the denoising intensity of the subsequent denoising process, the zoom scale of the first frame image is first identified by the first model to obtain the zoom scale feature, and further according to the zoom scale feature, the extraction process of the noise feature map is regulated, thereby ensuring that when the noise feature map obtained subsequently is used to denoise the first frame image, the noise is accurately denoised according to the zoom scale of the first frame image, thereby improving the accuracy of picture denoising during video conferencing.

[0131] Embodiment 3

[0132] Based on the above-mentioned embodiment of the denoising method for video conferencing portrait zooming, another embodiment of the present application provides a denoising terminal device for video conferencing portrait zooming. The denoising terminal device for video conferencing portrait zooming includes a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor. When the processor executes the computer program, the denoising method for video conferencing portrait zooming of any embodiment of the present application is implemented.

[0133] Exemplarily, in this embodiment, the computer program may be divided into one or more modules, which are stored in the memory and executed by the processor to complete the present application. The one or more modules may be a series of computer program instruction segments capable of completing specific functions, which are used to describe the execution process of the computer program in the denoising device for video conferencing portrait zooming.

[0134] The denoising device for zooming portraits in video conferences can be a computing device such as a desktop computer, a notebook, a PDA, a cloud server, etc. The denoising terminal device for zooming portraits in video conferences can include, but is not limited to, a processor and a memory.

[0135] The processor may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may be any conventional processor, etc. The processor is the control center of the denoising device for video conference portrait zooming, and uses various interfaces and lines to connect the various parts of the denoising device for video conference portrait zooming. The memory may be used to store the computer program and / or module, and the processor implements various functions of the denoising device for video conference portrait zooming by running or executing the computer program and / or module stored in the memory, and calling the data stored in the memory. The memory may mainly include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application required for at least one function, etc.; the data storage area may store data created according to the use of the mobile phone, etc. In addition, the memory may include high-speed random access memory and may also include non-volatile memory, such as a hard disk, a memory, a plug-in hard disk, a smart memory card (Smart Media Card, SMC), a secure digital (Secure Digital, SD) card, a flash card (Flash Card), at least one disk storage device, a flash memory device, or other volatile solid-state storage devices.

[0136] Embodiment 4

[0137] Based on the above-mentioned embodiment of the denoising method for video conferencing portrait zooming, another embodiment of the present application provides a storage medium, wherein the storage medium includes a stored computer program, wherein when the computer program is running, the device where the storage medium is located is controlled to execute the denoising method for video conferencing portrait zooming of any embodiment of the present application.

[0138] In this embodiment, the storage medium is a computer-readable storage medium, and the computer program includes computer program code, which may be in source code form, object code form, executable file or some intermediate form, etc. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signal, telecommunication signal and software distribution medium, etc. It should be noted that the content contained in the computer-readable medium may be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, computer-readable media do not include electric carrier signals and telecommunication signals.

[0139] The specific embodiments described above further describe the purpose, technical solutions and beneficial effects of the present application in detail. It should be understood that the above description is only a specific embodiment of the present application and is not intended to limit the scope of protection of the present application. It is particularly pointed out that for those skilled in the art, any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application should be included in the scope of protection of the present application.

Claims

1. A denoising method for portrait zooming in video conferencing, characterized in that: include: Inputting the first frame image currently collected in real time into the pre-trained first model, so that the first model can identify the scaling feature of the first frame image, and obtain the first feature of the first frame image by combining the first feature of the previous frame image of the first frame image; the first feature is a noise feature map extracted from the input image by the first model; The first frame image is superimposed with the first feature of the first frame image to obtain a denoised image of the first frame image.

2. A denoising method for portrait zooming in video conference as claimed in claim 1, characterized in that: The method of enabling the first model to identify the scaling feature of the first frame image and combining the first feature of a previous frame image of the first frame image to obtain the first feature of the first frame image includes: The first model includes a first branch, a second branch and a first convolutional layer; Extracting a second feature of the first frame image through the first convolutional layer; extracting the scaled feature from the second feature through the first branch; The second branch combines the zoom scale feature, the second feature, and the first feature of the previous frame image to obtain the first feature of the first frame image.

3. A denoising method for portrait zooming in video conference as claimed in claim 2, characterized in that: The extracting the scaling feature from the second feature through the first branch includes: Inputting the second feature into a second convolutional layer of the same scale as the first convolutional layer to obtain a third feature; The third feature is sequentially input into a plurality of third convolutional layers to compress the size of the third feature and obtain the scaled feature.

4. A denoising method for portrait zooming in video conference as claimed in claim 2, characterized in that: The step of obtaining the first feature of the first frame image by combining the zoom scale feature, the second feature, and the first feature of the previous frame image through the second branch includes: Inputting the second feature into a plurality of residual modules in sequence to obtain a fourth feature; Inputting the fourth feature and the scaling feature into a feature control module simultaneously to obtain a fifth feature; The fifth feature is concatenated with the first feature of the previous frame image and the concatenated features are sequentially inputted into a plurality of residual modules to obtain a sixth feature; The sixth feature and the scaling feature are simultaneously input into the feature control module to obtain the first feature of the first frame image.

5. A denoising method for portrait zooming in video conference as claimed in claim 4, characterized in that: include: The residual module includes a plurality of fourth convolutional layers; The feature control module includes a fifth convolutional layer and a plurality of residual modules; Inputting the scaled feature into the fifth convolutional layer to obtain a seventh feature; Inputting the seventh feature into each of the residual modules in the feature control module respectively, and obtaining corresponding first parameters respectively; According to each of the first parameters and the fourth feature or the sixth feature, an output of the corresponding feature control module is obtained.

6. A denoising method for portrait zooming in video conference as claimed in claim 2, characterized in that: The pre-training step of the first model includes: Performing data preprocessing on the first video data that has not been de-noised to obtain a training data set; Constructing corresponding first loss functions for the first branch and the second branch respectively, and obtaining a second loss function of the first model according to each of the first loss functions; Pre-train the first model according to the second loss function and the training data set.

7. A denoising method for portrait zooming in video conference as claimed in claim 6, characterized in that: The performing data preprocessing on the first video data that has not been de-noised to obtain a training data set includes: Processing the first video data by time series mean filtering to obtain a second frame of image; The second frame image is copied a first preset number of times to obtain second video data having the same number of frames as the first video data; Data enhancement is performed on the first video data and the second video data to obtain a training data set.

8. A denoising method for portrait zooming in video conference as claimed in claim 6, characterized in that: The constructing corresponding first loss functions for the first branch and the second branch respectively, and obtaining the second loss function of the first model according to the first loss functions, comprises: The loss function of the first model is specifically: in, is the loss value corresponding to the first model; L2(R,R_true) is the first loss function of the first branch, representing the L2 norm of R and R_true, R is the scaling scale corresponding to the scaling scale feature obtained through the first branch; R_true is the true scaling scale corresponding to the test data; λ1 is the weight coefficient; L1(Y,GT) is the first loss function of the second branch, representing the L1 norm of Y and GT, Y is the denoised image obtained according to the noise feature map, and GT is the true denoised image corresponding to the test data.

9. A denoising device for portrait zooming in video conference, characterized in that: include: Noise extraction module and denoising module; The noise extraction module is used to input the first frame image currently collected in real time into the pre-trained first model, so that the first model can identify the scaling feature of the first frame image, and obtain the first feature of the first frame image in combination with the first feature of the previous frame image of the first frame image; the first feature is a noise feature map extracted from the input image by the first model; The denoising module is used to superimpose the first frame image with the first feature of the first frame image to obtain a denoised image of the first frame image.

10. A terminal device, characterized in that: The invention comprises a processor, a memory and a computer program stored in the memory and configured to be executed by the processor, wherein when the processor executes the computer program, a denoising method for portrait zooming in a video conference is implemented as described in any one of claims 1 to 8.