Video conference background blurring method and system based on deep learning

By combining image semantic understanding and prior knowledge, the portrait mask of the current frame is optimized by using the Alpha mask of the previous frame, the problem of unsatisfactory edge segmentation and unstable timing in video conference background blur is solved, and a more natural background blur effect and a better user experience is achieved.

CN120378571AActive Publication Date: 2025-07-25SHENZHEN MINRRAY IND CORP LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510860557.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-25
Publication Date
2025-07-25
Estimated Expiration
2045-06-25

AI Technical Summary

Technical Problem

The existing video conference background blur method based on deep learning is not ideal in the segmentation accuracy of portrait edge areas, resulting in stiff and unnatural blur effect, and the timing of the video stream causes background blur effect to flicker or jump, affecting the user experience.

Method used

By combining image edge refinement processing module with image semantic understanding ability and prior knowledge, the Alpha mask of the previous frame has been refined as a reference, the portrait mask of the current frame is deeply optimized to generate a current frame with clearer edges, richer details and more time-consistent Alpha mask for background blur processing.

Benefits of technology

Effectively smooth and correct portrait edges, reduce flickering and jumping of background blur effects, and enhance visual reality and user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120378571A_ABST
    Figure CN120378571A_ABST
Patent Text Reader

Abstract

The invention relates to the field of image intelligent processing, and particularly discloses a video conference background blurring method and system based on deep learning, and the method comprises the steps: firstly carrying out the rapid preliminary positioning of a portrait region of a current video frame, so as to meet the real-time processing demands; then, a unique edge refining and stabilizing mechanism is introduced: the mechanism does not independently process each frame, but intelligently uses the refined Alpha mask of the previous frame as an important reference, and combines the initial segmentation result of the current frame. According to the method, a portrait mask of a current frame is deeply optimized through an image edge refining processing module combining image semantic understanding ability and prior knowledge, a current frame Alpha mask which is clearer in edge, richer in details and more coherent in time is finally generated, and then a current frame background blurred image is generated by using the mask graph.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of intelligent image processing, and more specifically, to a method and system for virtualizing the background of video conferences based on deep learning. Background Art

[0002] With the popularization of remote work and online communication, video conferencing has become an indispensable part of daily life and work. During video conferencing, users often hope to protect personal privacy, maintain a professional image, or reduce the interference of the background environment on the meeting. Therefore, the background virtualization function has been widely favored. By blurring the background behind the user, the visual focus can be concentrated on the speaker, effectively improving the experience and concentration of video conferencing. Traditional background virtualization methods usually rely on specific hardware devices, such as depth cameras or green screen backgrounds, which are often not available in the daily use scenarios of ordinary users, limiting the wide application of these methods. Therefore, it is of great practical significance and application value to develop a background virtualization solution based on ordinary cameras and implemented through software algorithms.

[0003] Currently, there are already some video conference background virtualization solutions based on deep learning. They usually use a human portrait semantic segmentation network (such as models based on architectures like U-Net, DeepLab, etc.) to identify the human portrait area in the video frame, generate a binary human portrait mask (mask), and then apply blur processing to the background area based on this mask. Although these methods can achieve background virtualization to a certain extent, they still face many challenges and defects in practical applications. First, the segmentation accuracy of the existing methods in the edge area of the human portrait is often not ideal. Especially when dealing with complex situations such as hair strands, clothing wrinkles, finger gaps, and when the color of the person is similar to the background color, it is easy to produce edge sawtooth, detail loss, or incorrect segmentation, resulting in a rigid and unnatural virtualization effect. Second, the video stream has temporality. If the segmentation results between consecutive frames lack stability, even a small jitter or inconsistency will cause the background virtualization effect to flicker or jump in the visual sense, seriously affecting the user experience. Therefore, how to improve the fineness of the background virtualization edge and the temporal stability while ensuring real-time performance is a difficult problem that needs to be solved urgently in the current technology.

[0004] Therefore, an optimized video conference background virtualization solution is expected. Summary of the Invention

[0005] To solve the above technical problems, the present application is proposed. Embodiments of the present application provide a method and system for virtualizing the background of a video conference based on deep learning. First, it performs a preliminary localization of the fast portrait area on the current video frame to meet the requirements of real-time processing. Subsequently, a unique edge refinement and stabilization mechanism is introduced: this mechanism does not process each frame independently, but intelligently uses the Alpha mask that has been refined in the previous frame as an important reference and combines it with the preliminary segmentation result of the current frame. Through an image edge refinement processing module that combines image semantic understanding ability and prior knowledge, it deeply optimizes the portrait mask of the current frame, and finally generates a current frame Alpha mask with clearer edges, richer details, and more temporal coherence. Then, this mask image is used to generate the current frame background virtualized image. This design can effectively smooth and correct the portrait edges, reduce the flickering and jumping of the background virtualization effect, and thus improve the visual realism and user experience while ensuring performance.

[0006] According to one aspect of the present application, there is provided a method for virtualizing the background of a video conference based on deep learning, which includes:

[0007] Collecting raw video conference video stream data through a camera and extracting the current frame image of the video conference from the raw video conference video stream data;

[0008] Performing portrait semantic segmentation on the current frame image of the video conference to obtain a current frame portrait rough segmentation mask image of the video conference;

[0009] Obtaining the previous frame image of the current frame image of the video conference to obtain a previous frame image of the video conference;

[0010] Performing an Alpha mask on the previous frame image of the video conference to obtain a previous frame Alpha mask image;

[0011] Passing the current frame portrait rough segmentation mask image of the video conference and the previous frame Alpha mask image through an image edge refinement processing module based on prior information to obtain a current frame Alpha mask image;

[0012] Generating a blurred background image based on the current frame image of the video conference;

[0013] Using the current frame Alpha mask image as a weight, using the current frame image of the video conference as the foreground and using the blurred background image as the background, and performing Alpha blending to obtain a current frame background virtualized image of the video conference.

[0014] According to another aspect of the present application, there is provided a system for virtualizing the background of a video conference based on deep learning, which includes:

[0015] A current video conference frame acquisition module, configured to collect original video conference video stream data through a camera, and extract a current video conference frame image from the original video conference video stream data;

[0016] A human portrait semantic segmentation module, configured to perform human portrait semantic segmentation on the current video conference frame image to obtain a current video conference frame human portrait rough segmentation mask image;

[0017] A previous conference frame image acquisition module, configured to acquire a previous frame image of the current video conference frame image to obtain a previous video conference frame image;

[0018] A previous frame Alpha mask module, configured to perform Alpha masking on the previous video conference frame image to obtain a previous frame Alpha mask image;

[0019] An image edge refinement processing module, configured to perform image edge refinement processing on the current video conference frame human portrait rough segmentation mask image and the previous frame Alpha mask image through an image edge refinement processing module based on prior information to obtain a current frame Alpha mask image;

[0020] A blurred background image generation module, configured to generate a blurred background image based on the current video conference frame image;

[0021] A conference background blurring image generation module, configured to use the current frame Alpha mask image as a weight, use the current video conference frame image as a foreground, and use the blurred background image as a background, and perform Alpha blending to obtain a current video conference frame background blurring image.

[0022] Compared with the prior art, a video conference background blurring method and system based on deep learning provided by the present application first performs a preliminary positioning of a fast human portrait area on the current video frame to meet the requirements of real-time processing. Subsequently, a unique edge refinement and stabilization mechanism is introduced: instead of independently processing each frame, this mechanism intelligently uses the Alpha mask that has been refined in the previous frame as an important reference and combines it with the preliminary segmentation result of the current frame. Through an image edge refinement processing module that combines image semantic understanding ability and prior knowledge, the human portrait mask of the current frame is deeply optimized, and finally a current frame Alpha mask with clearer edges, richer details and more temporal coherence is generated, and then this mask image is used to generate the current frame background blurring image. This design can effectively smooth and correct the human portrait edges, reduce the flicker and jump of the background blurring effect, and thus improve the visual realism and user experience while ensuring performance. Description of the Drawings

[0023] The above and other objects, features, and advantages of the present application will become more apparent by describing the embodiments of the present application in more detail with reference to the accompanying drawings. The accompanying drawings are used to provide a further understanding of the embodiments of the present application and constitute a part of the specification. They are used to explain the present application together with the embodiments of the present application and do not constitute a limitation to the present application. In the drawings, the same reference numerals generally represent the same components or steps.

[0024] Figure 1 It is a flowchart of a method for background blurring in a video conference based on deep learning according to an embodiment of the present application;

[0025] Figure 2 It is a schematic diagram of data flow of a method for background blurring in a video conference based on deep learning according to an embodiment of the present application;

[0026] Figure 3 It is a flowchart of obtaining the current frame Alpha mask image by performing image edge refinement processing on the current frame portrait rough segmentation mask image and the previous frame Alpha mask image of the video conference through an image edge refinement processing module based on prior information according to an embodiment of the present application;

[0027] Figure 4 It is a flowchart of obtaining a semantically optimized representation vector of the current frame portrait in a video conference by performing an image semantic joint encoder based on prior information-assisted optimization on the semantic feature map of the previous frame image and the semantic feature vector of the previous frame Alpha mask in a method for background blurring in a video conference based on deep learning according to an embodiment of the present application;

[0028] Figure 5 It is a block diagram of a system for background blurring in a video conference based on deep learning according to an embodiment of the present application. Detailed implementation manners

[0029] Next, exemplary embodiments according to the present application will be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments of the present application. It should be understood that the present application is not limited by the exemplary embodiments described herein.

[0030] As shown in the present application and the claims, unless the context clearly indicates otherwise, words such as "a", "an", "one", and / or "the" are not specifically singular and may also include plural. Generally speaking, the terms "comprising" and "including" only indicate the inclusion of the clearly identified steps and elements, and these steps and elements do not constitute an exclusive list. The method or device may also include other steps or elements.

[0031] Although this application makes various references to certain modules in the system according to embodiments of this application, any number of different modules may be used and run on a user terminal and / or a server. The modules are merely illustrative, and different aspects of the system and method may use different modules.

[0032] Flowcharts are used in this application to illustrate the operations performed by the system according to embodiments of this application. It should be understood that the operations above or below do not necessarily have to be executed precisely in order. Instead, various steps may be processed in reverse order or simultaneously as needed. Also, other operations may be added to these processes, or one or more steps may be removed from these processes.

[0033] Hereinafter, example embodiments according to this application will be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all embodiments of this application. It should be understood that this application is not limited by the example embodiments described herein.

[0034] The concept of this technical solution aims to construct a set of efficient and natural video conferencing background blurring solutions through deep learning to solve problems such as rough portrait edge processing, detail loss, and unstable blurring effects in video sequences in the prior art. Its core idea is to adopt a strategy that combines fast preliminary segmentation and refined temporal optimization. Specifically, first, a quick preliminary positioning of the portrait area in the current video frame is performed to meet the requirements of real-time processing. Subsequently, the key technical innovation lies in introducing a unique edge refinement and stabilization mechanism: This mechanism does not process each frame independently, but intelligently uses the Alpha mask that has been refined in the previous frame as an important reference and combines it with the preliminary segmentation result of the current frame. Through an "image edge refinement processing module" with built-in image semantic understanding ability and prior knowledge, the portrait mask of the current frame is deeply optimized. This module performs high-level semantic feature fusion and representation learning on the rough information of the current frame and the fine information of the previous frame, and finally generates an Alpha mask of the current frame with clearer edges, richer details, and more temporal coherence. Furthermore, based on the Alpha mask of the current frame, the current frame image of the video conference, and the blurred background image, the video conference background blurring process is comprehensively performed. This design can effectively smooth and correct the portrait edge, reduce the flicker and jump of the background blurring effect, thereby improving the visual realism and user experience while ensuring performance.

[0035] In the technical solution of this application, a deep learning-based video conferencing background blurring method is proposed. Figure 1 It is a flowchart of the deep learning-based video conferencing background blurring method according to an embodiment of this application. Figure 2Schematic diagram of data flow for the deep learning-based video conferencing background blurring method according to an embodiment of the present application. As Figure 1 and Figure 2 shown, the deep learning-based video conferencing background blurring method according to an embodiment of the present application includes the steps of: S100, collecting original video conferencing video stream data through a camera, and extracting the current frame image of the video conference from the original video conferencing video stream data; S200, performing human portrait semantic segmentation on the current frame image of the video conference to obtain a rough segmentation mask image of the current frame human portrait of the video conference; S300, obtaining the previous frame image of the current frame image of the video conference to obtain a previous frame image of the video conference; S400, performing an Alpha mask on the previous frame image of the video conference to obtain a previous frame Alpha mask image; S500, passing the rough segmentation mask image of the current frame human portrait of the video conference and the previous frame Alpha mask image through an image edge refinement processing module based on prior information to obtain a current frame Alpha mask image; S600, generating a blurred background image based on the current frame image of the video conference; S700, using the current frame Alpha mask image as a weight, using the current frame image of the video conference as the foreground and the blurred background image as the background, and performing Alpha blending to obtain a background-blurred image of the current frame of the video conference.

[0036] Specifically, in step S100, original video conferencing video stream data is collected through a camera, and the current frame image of the video conference is extracted from the original video conferencing video stream data. It should be understood that the background blurring technology essentially separates the foreground human portrait from the background and blurs the background area for a single frame image, while the input of a video conference is a continuous dynamic video stream. Therefore, it is necessary to first obtain discrete static image frames one by one from this continuous video stream as the basic unit for subsequent deep learning model processing. Each frame image carries the user's image and background information at a specific moment and is the direct object for performing human portrait segmentation, feature extraction, and background blurring operations.

[0037] More specifically, in a specific example of the present application, first, the camera device is initialized and accessed, which is usually achieved through the application programming interface (API) provided by the operating system or a standardized multimedia framework to establish a communication link with the physical camera. Second, once the camera is activated, it starts continuously outputting raw video data to form a video stream; the system needs to capture this data stream in real time. Then, since the raw video stream data is often encoded and compressed (such as in H.264, VP9, etc. formats), it needs to be decoded by the corresponding decoder to restore it into a series of uncompressed or low-compressed image frame data, and each frame is usually represented as a pixel matrix. Finally, from the decoded sequence of consecutive image frames, according to the frame rate of the video (such as 25 frames per second or 30 frames per second) or according to the processing requirements, an image frame at a specific moment is selected as the "current video conference frame image" for subsequent processing by the deep learning module.

[0038] Specifically, in step S200, semantic segmentation of the human figure in the current video conference frame image is performed to obtain a rough segmentation mask image of the human figure in the current video conference frame. It should be understood that in order to accurately apply a blurring effect to the background without affecting the clarity of the foreground human figure, the system must first understand the image content at the pixel level and clarify which pixels belong to the human figure and which belong to the background. Therefore, in the technical solution of the present application, semantic segmentation of the human figure in the current video conference frame image is performed to obtain a rough segmentation mask image of the human figure in the current video conference frame. Semantic segmentation of the human figure can assign a class label (such as "human" or "background") to each pixel in the image, thereby generating a rough segmentation mask image of the human figure in the current video conference frame that identifies the human figure area. This mask image directly guides the subsequent background blurring process and Alpha blending operation, and is the "blueprint" for distinguishing the foreground from the background.

[0039] More specifically, in the embodiments of the present application, performing human portrait semantic segmentation on the current frame image of the video conference to obtain a rough segmentation mask image of the current frame of the video conference human portrait, including: passing the current frame image of the video conference through a human portrait semantic segmentation module based on MobileNet to obtain the rough segmentation mask image of the current frame of the video conference human portrait. Specifically, first, the acquired "current frame image of the video conference" is preprocessed, which may include adjusting its size to the input size expected by the MobileNet model (such as 224x224 or a size with a specific aspect ratio), and performing pixel value normalization operations (such as scaling pixel values from the range of 0-255 to the range of 0-1 or -1 to 1). Subsequently, the preprocessed image data is fed into a semantic segmentation network based on the MobileNet architecture. MobileNet, as a lightweight convolutional neural network, its core lies in the adoption of depthwise separable convolutions, which can significantly reduce the number of computational parameters and the amount of operations while maintaining good feature extraction capabilities, thus being very suitable for performing human portrait segmentation in resource-constrained devices or video conference scenarios that require real-time response. This MobileNet backbone network is responsible for extracting the deep semantic features of the image, and usually, a decoder structure (such as an upsampling path similar to U-Net or a module that uses dilated convolutions to maintain the feature map resolution and dilation, such as the ASPP module in the DeepLab series) is connected on top of it to gradually restore the spatial resolution and perform pixel-level classification. The network finally outputs a segmentation map of the same size as the input image (or the adjusted input size), and the value of each pixel point on the map indicates the probability that the point belongs to the human portrait or directly is a class label (such as 1 representing a person and 0 representing the background). This output map is the "rough segmentation mask image of the current frame of the video conference human portrait", which defines the general outline and area of the human portrait and provides a basis for subsequent refined processing and background blurring. This preliminary segmentation result based on a lightweight model can provide a good starting point for subsequent more refined edge processing while meeting the real-time requirements.

[0040] Specifically, in step S300 and step S400, the previous frame image of the current frame image of the video conference is obtained to obtain the pre-frame image of the video conference, and an Alpha mask is applied to the pre-frame image of the video conference to obtain the pre-frame Alpha mask image. It should be understood that since the content change between adjacent frames in a video is usually gradual, the processing result of the previous frame, especially its refined portrait mask, contains valuable prior information about the portrait edge, details, and morphology. By introducing the previous frame Alpha mask of the pre-frame image of the video conference, it can provide strong guidance for the refinement of the portrait segmentation edge of the current frame, helping to correct possible jitter, discontinuity, or detail loss problems in the rough segmentation of the current frame, thereby generating a more smooth background blurring effect in the time dimension and a more natural visual effect. If the previous frame information is not utilized and each frame is independently segmented and blurred, the accumulation of tiny segmentation errors may lead to unnatural flickering or jumping at the edge of the background blurring area in the final output video, affecting the user experience.

[0041] More specifically, in a specific example of the present application, first, regarding "obtaining the previous frame image of the current frame image of the video conference to obtain the video conference previous frame image", when the system processes a continuous video stream, it needs to maintain a buffer that contains at least the previous frame image. When the system processes the Nth frame (i.e., the current frame), the (N - 1)th frame (i.e., the previous frame image) has been captured and stored in the previous processing cycle. This process usually occurs after the video capture module and manages the frame sequence through simple data structures such as queues or pointers to ensure that the data of the immediately adjacent previous frame image is accessible when processing the current frame. Second, regarding "performing an Alpha mask on the video conference previous frame image to obtain the previous frame Alpha mask image", the "performing an Alpha mask" here is to obtain or generate a refined Alpha mask corresponding to the previous frame image. In the entire continuous processing flow, when the "video conference previous frame image" (i.e., the (N - 1)th frame) is processed as the "current frame", the system has already performed a series of operations on it, such as full-body segmentation, edge refinement, and Alpha blending, and finally generated the "current frame Alpha mask image" of this frame. The "current frame Alpha mask image" that has been refined for the (N - 1)th frame is used as the "previous frame Alpha mask image" when processing the Nth frame. Therefore, the "performing an Alpha mask" here does not mean re-executing a complete segmentation and mask generation process on the previous frame image, but rather directly using or retrieving the highly optimized Alpha mask result produced by the system when processing the previous frame. The implementation of this step is essentially a mechanism for state preservation and transfer: at the end of each processing cycle, the generated "current frame Alpha mask image" is temporarily stored so that it can be called by the subsequent "image edge refinement processing module based on prior information" as the "previous frame Alpha mask image" in the next processing cycle (when processing the next frame). If the system is started for the first time or processes the first frame of the video stream, and there is no "previous frame Alpha mask image" at this time, a default fully transparent or fully opaque mask can be used, or the refinement step based on the previous frame information can be skipped, and only the rough segmentation result of the current frame is relied on. This method of using the refined Alpha mask of the previous frame is equivalent to introducing a high-quality temporal prior for the edge refinement of the current frame, which is itself the product of the previous round of deep learning model processing and optimization and contains a more accurate description of the human body edge.

[0042] Specifically, in step S500, the rough segmentation mask image of the portrait of the current frame of the video conference and the Alpha mask image of the previous frame are processed through an image edge refinement processing module based on prior information to obtain the Alpha mask image of the current frame. It should be understood that simply relying on the rough segmentation mask image of the portrait of the current frame of the video conference (for example, the result quickly generated by a lightweight network such as MobileNet) is often difficult to balance speed and accuracy, especially in the edge area of the portrait, such as hair, clothing contours and other details, it is easy to have jagged, blurred or wrong segmentation, which directly affects the naturalness and realism of the final blur effect; at the same time, the independent segmentation of each frame in the video stream may cause the mask to lack continuity in the time series, which is manifested as jitter or flickering of the blurred edge, seriously reducing the user experience. Therefore, in order to make up for the lack of accuracy of the rough segmentation and introduce information of the time dimension to enhance the stability of the segmentation, in the technical solution of the present application, the rough segmentation mask image of the portrait of the current frame of the video conference and the Alpha mask image of the previous frame are further processed through an image edge refinement processing module based on prior information to obtain the Alpha mask image of the current frame. The previous frame Alpha mask image carries the portrait edge information that has been refined at the previous moment, constituting valuable "prior information". This step is performed precisely to solve the problem of insufficient precision of the coarse segmentation mask and time discontinuity, and a special module is used to fuse the preliminary judgment of the current frame with the high-quality result of the previous frame. Specifically, through the processing of the image edge refinement processing module based on prior information, it can intelligently determine which areas of the "current frame portrait coarse segmentation mask image" need to be "attracted" and corrected by the more accurate edge information in the "previous frame Alpha mask image". It does not simply superimpose or average the two masks, but based on prior knowledge (such as edge continuity, smoothness assumptions, and the inherent characteristics of the human body structure), it selectively adopts and fuses the fine edge details from the previous frame, while suppressing the noise and errors that may be introduced by the coarse segmentation of the current frame. This process can be seen as a feature decoupling and targeted enhancement of the rough mask of the current frame. The goal is to make the final generated Alpha mask of the current frame not only achieve pixel-level accuracy within the current frame, but also show good neighborhood correlation universality in the transition with the previous frame, thereby promoting the extraction of contextual joint semantics between the character edge and background blur during subsequent Alpha blending to achieve harmonious unity.

[0043] Figure 3 The present invention is a flowchart of the method for blurring the background of a video conference based on deep learning according to an embodiment of the present application, which processes the coarse segmentation mask image of the portrait in the current frame of the video conference and the Alpha mask image in the previous frame through an image edge refinement processing module based on prior information to obtain the Alpha mask image in the current frame. Figure 3As shown, the method for background blurring in a video conference based on deep learning according to an embodiment of the present application, step S500, includes: S510, extracting video conference semantic features from the rough segmentation mask image of the current frame portrait in the video conference to obtain a semantic feature map of the previous frame image of the video conference; S520, extracting semantic features of the previous frame Alpha mask image from the previous frame Alpha mask image to obtain a semantic feature vector of the previous frame Alpha mask; S530, passing the semantic feature map of the previous frame image of the video conference and the semantic feature vector of the previous frame Alpha mask through an image semantic joint encoder assisted by prior information optimization to obtain a semantic optimization representation vector of the current frame portrait in the video conference; S540, generating the current frame Alpha mask image based on the semantic optimization representation vector of the current frame portrait in the video conference.

[0044] Specifically, in step S510, video conference semantic features are extracted from the rough segmentation mask image of the current frame portrait in the video conference to obtain a semantic feature map of the previous frame image of the video conference. It should be understood that when directly using the pixel-level "rough segmentation mask image of the current frame portrait in the video conference" and the "previous frame Alpha mask image" for fusion and refinement, the dimension of information representation is relatively low, and it is difficult to capture the deep structural and content associations between the two, especially when dealing with complex situations such as inaccurate edges of the rough segmentation mask, the existence of noise, or a large difference from the fine mask of the previous frame. In order to achieve more intelligent and robust edge refinement, it is necessary to elevate the rough segmentation mask to the semantic feature level, so that it contains higher-level abstract information, such as general features of the portrait's contour, posture, key parts, etc., rather than just the attribution of pixel points. Therefore, in the technical solution of the present application, it is necessary to extract video conference semantic features from the rough segmentation mask image of the current frame portrait in the video conference to obtain a semantic feature map of the previous frame image of the video conference. More specifically, in the embodiment of the present application, the rough segmentation mask image of the current frame portrait in the video conference is passed through a video conference semantic feature extractor based on a depthwise separable convolutional neural network model to obtain the semantic feature map of the previous frame image of the video conference.

[0045] It is worth mentioning that here, the depthwise separable convolutional neural network model is suitable for real-time video processing scenarios due to its computational efficiency, and can reduce the model complexity while ensuring a certain feature extraction ability. This feature extractor learns to map the input rough mask image to a multi-channel feature map, where each channel corresponds to different semantic attributes or spatial structure features of the human figure. This semantic feature map of the previous frame image in the video conference will serve as an important input for the subsequent image semantic joint encoder, which is used to interact and fuse with the semantic features extracted from the "previous frame Alpha mask image" in a higher dimension, so as to more profoundly understand the connection and difference between the current frame's rough segmentation result and the previous frame's fine result. In this way, the system no longer solely relies on pixel-level matching or simple filtering operations for edge refinement, but can integrate information at the semantic level, enabling it to more intelligently utilize the prior information of the previous frame to guide the optimization of the current frame's mask. For example, even if there is a large deviation in the edge of the current frame's rough segmentation, its semantic features may still imply that the area belongs to a certain part of the human body (such as an arm). Combining with the fine semantic features of the corresponding area in the previous frame's Alpha mask, the joint encoder can more accurately judge and correct the edge of the current frame. This helps to overcome the uncertainty brought by rough segmentation, and further improve the stability and realism of the background blurring effect.

[0046] Specifically, in step S520, the semantic features of the previous frame Alpha mask image are extracted from the previous frame Alpha mask image to obtain the previous frame Alpha mask semantic feature vector. It should be understood that although the "previous frame Alpha mask image" itself is already a refined high-quality result, in order to perform deeper and more effective fusion and comparison with the semantic features of the "coarse segmentation mask image of the current frame of the video conference" in the subsequent "image semantic joint encoder", this pixel-level fine mask also needs to be converted into an abstract and global semantic representation. That is to say, directly comparing or fusing two masks from different sources (one rough and one fine; one current and one past) at the pixel level may be difficult to capture their similarities and differences in macroscopic structure and contour morphology, while semantic features can provide this higher-dimensional understanding. Therefore, in the technical solution of this application, the semantic features of the previous frame Alpha mask image are further extracted from the previous frame Alpha mask image to obtain the previous frame Alpha mask semantic feature vector. More specifically, in the embodiment of this application, the previous frame Alpha mask image is passed through a semantic feature extractor for the previous frame Alpha mask image based on the ViT model to obtain the previous frame Alpha mask semantic feature vector. By using the "semantic feature extractor for the previous frame Alpha mask image based on the ViT (Vision Transformer) model" to achieve this purpose, the powerful ability of the ViT model in capturing global dependencies and long-range features of images is utilized. ViT regards an image as a series of image patches and uses the self-attention mechanism to learn the relationships between these patches, so as to effectively extract the overall shape and key semantic information from the entire "previous frame Alpha mask image" and compress it into a fixed-length previous frame Alpha mask semantic feature vector. This feature vector will serve as the "essence summary" of the high-quality segmentation result of the previous frame, carrying strong prior knowledge about the expected shape of the ideal person mask and ready to interact intelligently with the information of the current frame.

[0047] Specifically, in step S530, the video conference previous frame image semantic feature map and the previous frame Alpha mask semantic feature vector are passed through an image semantic joint encoder assisted and optimized based on prior information to obtain a video conference current frame portrait semantic optimized representation vector. It should be understood that in the previous steps, preliminary semantic features (i.e., the "video conference previous frame image semantic feature map", representing the current frame information) have been extracted from the rough segmentation mask of the current frame, and high-quality semantic features (i.e., the "previous frame Alpha mask semantic feature vector", representing historical high-quality prior information) have been extracted from the refined Alpha mask of the previous frame. However, these two types of features come from different sources and have different precisions. Simply splicing or averaging them for fusion is difficult to fully utilize their respective advantages and effectively suppress noise. Therefore, the fusion of these two different information sources requires an intelligent and selective interaction mechanism rather than a blind combination. Based on this, in the technical solution of this application, the video conference previous frame image semantic feature map and the previous frame Alpha mask semantic feature vector are further passed through an image semantic joint encoder assisted and optimized based on prior information to obtain a video conference current frame portrait semantic optimized representation vector. This video conference current frame portrait semantic optimized representation vector should not only accurately reflect the true semantic boundary of the portrait in the current video frame but also fully draw on the advantages of the previous frame segmentation to maintain temporal continuity and stability. Through the processing of the image semantic joint encoder assisted and optimized based on prior information, complex non-linear transformations and in-depth interaction modeling can be performed on the two input semantic features. Let the "previous frame Alpha mask semantic feature vector" have an "attracting" or "correcting" effect on the local semantics in the "video conference previous frame image semantic feature map", so that the current frame features that are more consistent and reliable with the high-quality prior of the previous frame are enhanced, while the inconsistent or noisy features are suppressed. In this way, the encoder learns how to extract the most refined and reliable combined information from the preliminary semantics of the current frame and the fine semantics of the previous frame. Subsequently, the "current frame Alpha mask image" generated based on this video conference current frame portrait semantic optimized representation vector will exhibit excellent performance in terms of the fineness of the portrait edge, the retention of details such as hair, and the temporal stability (reducing flicker and jitter) during the dynamic process, making the final video conference background blurring effect more natural and professional, and greatly improving the user experience.

[0048] Figure 4 A flowchart of obtaining a video conference current frame portrait semantic optimized representation vector by passing the video conference previous frame image semantic feature map and the previous frame Alpha mask semantic feature vector through an image semantic joint encoder assisted and optimized based on prior information for the deep learning-based video conference background blurring method according to an embodiment of the present application. As Figure 4As shown, the method for virtual background of video conferencing based on deep learning according to an embodiment of the present application, step S530 includes: S531, performing feature decoupling on the semantic feature map of the previous frame image of the video conferencing to obtain a set of local semantic feature encoding vectors of the previous frame image of the video conferencing; S532, based on the feature class magnetic attraction effect between the previous frame Alpha mask semantic feature vector and each local semantic feature encoding vector in the set of local semantic feature encoding vectors of the previous frame image of the video conferencing, determining a fast interaction set of local semantic feature encoding vectors of the previous frame image of the video conferencing; S533, inputting the previous frame Alpha mask semantic feature vector and the fast interaction set of local semantic feature encoding vectors of the previous frame image of the video conferencing into a multi-modal multi-scale joint encoder based on an LSTM-Transformer hybrid architecture to obtain a local context semantic joint encoding vector of the current frame - previous frame of the video conferencing and a global context semantic joint encoding vector of the current frame - previous frame of the video conferencing; S534, fusing the local context semantic joint encoding vector of the current frame - previous frame of the video conferencing and the global context semantic joint encoding vector of the current frame - previous frame of the video conferencing to obtain a portrait semantic optimization representation vector of the current frame of the video conferencing.

[0049] More specifically, step S531, performing feature decoupling on the semantic feature map of the previous frame image of the video conferencing to obtain a set of local semantic feature encoding vectors of the previous frame image of the video conferencing, which is expressed by the formula as:

[0050]

[0051]

[0052] Among them, is the semantic feature map of the previous frame image of the video conferencing, represents a one-dimensional convolutional layer or a fully connected layer, represents the block position index, represents feature decoupling, represents feature flattening processing, is the set of local semantic feature encoding vectors of the previous frame image of the video conferencing, are respectively the 1st, 2nd, th, and th local semantic feature encoding vectors in the set of local semantic feature encoding vectors of the previous frame image of the video conferencing.

[0053] It should be understood that since the semantic feature map of the previous frame of the video conference is a relatively holistic feature representation that includes spatial dimension information. However, in order to perform more detailed and selective interaction and fusion with the global "semantic feature vector of the previous frame Alpha mask" representing the fine segmentation result of the previous frame, it is necessary to decompose this holistic semantic feature map of the previous frame of the video conference into smaller and more representative local units. Without decoupling, only a rough global-to-global comparison can be performed, and it is difficult to achieve differential processing and fine correction of different regions of the current frame's rough segmentation mask. Each local semantic feature encoding vector of the previous frame of the video conference image carries the semantic information of the "semantic feature map of the previous frame of the video conference image" (i.e., the semantic representation of the current frame's rough segmentation mask) in a specific local region. Through this decoupling, for example, by dividing and encoding along the channel dimension or spatial dimension, the complex spatial structure feature map can be transformed into a set of discrete local feature vectors with a fixed dimension. This is done so that the global "semantic feature vector of the previous frame Alpha mask" can interact with these feature vectors representing different local regions of the current frame one by one, thereby quantifying the "feature imitation magnetic attraction effect value" or association strength between them. This decoupling process is a prerequisite for realizing subsequent selective information fusion and fine adjustment, laying a foundation for judging which local regions of the current frame's rough segmentation are more "attractive" or consistent with the high-quality prior of the previous frame.

[0054] Correspondingly, according to an embodiment of the present application, in step S532, based on the feature class magnetic attraction effect between the semantic feature vector of the previous frame Alpha mask and each local semantic feature encoding vector of the previous frame of the video conference image in the set, a fast interaction set of the local semantic feature encoding vectors of the previous frame of the video conference image is determined, including: calculating the class magnetic attraction characteristic parameters between the semantic feature vector of the previous frame Alpha mask and each local semantic feature encoding vector of the previous frame of the video conference image in the set to obtain a set of class magnetic attraction characteristic parameters of the current frame - previous frame of the video conference; based on the comparison between the set of class magnetic attraction characteristic parameters of the current frame - previous frame of the video conference and a preset threshold, extracting the fast interaction set of the local semantic feature encoding vectors of the previous frame of the video conference image from the set of local semantic feature encoding vectors of the previous frame of the video conference image.

[0055] More specifically, calculating the class magnetic attraction characteristic parameters between the semantic feature vector of the previous frame Alpha mask and each local semantic feature encoding vector of the previous frame of the video conference image in the set to obtain a set of class magnetic attraction characteristic parameters of the current frame - previous frame of the video conference, which is expressed by the formula as:

[0056]

[0057] ;

[0058]

[0059] Among them, represents the two-norm of the vector, is the semantic feature vector of the previous frame Alpha mask, is a trainable permutation weight matrix, is a trainable modulation weight vector, is and is the current frame-previous frame class magnetic attraction characteristic parameter between represents the maximum value in the extraction vector, represents the minimum value in the extraction vector, represents the average value of the calculation vector, is 's autocorrelation characteristic parameter, is 's autocorrelation characteristic parameter.

[0060] It should be understood that the semantic feature vector of the previous frame Alpha mask represents the global semantic fingerprint of the previous frame's fine segmentation result, and the set of local semantic feature encoding vectors of the previous frame image of the video conference represents the semantic descriptions of the current frame's rough segmentation result in different local regions. To achieve true intelligent fusion and optimization, rather than simply imposing the previous frame information on the current frame, the system must first evaluate the association strength or similarity between each local semantic feature of the current frame and the global semantic prior of the previous frame. Without this quantitative evaluation, subsequent fusion optimization will be blind and may not be able to effectively distinguish which local information in the current frame's rough segmentation is reliable and which needs to be corrected by the previous frame's fine information. Therefore, in the technical solution of this application, the class magnetic attraction characteristic parameter between each local semantic feature encoding vector of the semantic feature vector of the previous frame Alpha mask and the set of local semantic feature encoding vectors of the previous frame image of the video conference is further calculated. This class magnetic attraction characteristic parameter aims to quantify the "attraction" or semantic correlation degree between the semantic information of each local region of the current frame's rough segmentation mask and the "semantic feature vector of the previous frame Alpha mask" representing the high-quality segmentation result. In this scenario, a high "class magnetic attraction characteristic parameter" means that the rough segmentation result of this local region of the current frame is highly consistent with the previous frame's fine segmentation result in terms of semantics and has a high credibility; conversely, a low parameter value may indicate that there are large deviations in the rough segmentation of this local region or it does not conform to the ideal state and needs to draw more on the previous frame information for adjustment. Therefore, generating this set of parameters is the key prerequisite for subsequent selective fusion and weighted optimization, and it provides a basis for intelligent decision-making.

[0061] More specifically, based on the comparison between the set of the current-frame - previous-frame magnet-like characteristic parameters of the video conference and a preset threshold, a fast interaction set of the local semantic feature encoding vectors of the previous-frame image of the video conference is extracted from the set of the local semantic feature encoding vectors of the previous-frame image of the video conference, which is expressed by the formula:

[0062] ;

[0063] ;

[0064] wherein, is a predetermined threshold, is the fast interaction set of the local semantic feature encoding vectors of the previous-frame image of the video conference, are respectively the 1st, 2nd, and th local semantic feature encoding vectors of the previous-frame image of the video conference in the fast interaction set of the local semantic feature encoding vectors of the previous-frame image of the video conference.

[0065] It should be understood that after calculating the "magnet-like characteristic parameters" between each local semantic region of the current frame and the high-quality global semantics of the previous frame, it is clear that not all local regions of the current-frame rough segmentation mask are equally important or have the same credibility for subsequent optimization. Some regions may be highly consistent with the previous frame and of high quality, while other regions may have significant differences from the previous frame or contain significant noise due to motion, occlusion, or the limitations of the rough segmentation itself. If all local semantic features are sent to the subsequent complex "image semantic joint encoder" for processing without selection, it will not only greatly increase the computational burden and affect real-time performance, but also may interfere with the optimization process due to the introduction of too much low-quality or irrelevant information, and even reduce the final segmentation accuracy. Therefore, in the technical solution of this application, further based on the comparison between the set of the current-frame - previous-frame magnet-like characteristic parameters of the video conference and a preset threshold, a fast interaction set of the local semantic feature encoding vectors of the previous-frame image of the video conference is extracted from the set of the local semantic feature encoding vectors of the previous-frame image of the video conference. This fast interaction set only contains those local semantic feature vectors whose corresponding "magnet-like characteristic parameters" are higher than the preset threshold, that is, those local regions of the current frame that are initially judged to have a strong "attraction" or high correlation with the high-quality mask of the previous frame in terms of semantics. By setting a threshold for screening, the purpose is to eliminate those local semantic information with weak prior relevance to the previous frame, which may contain noise or errors, so as to focus the attention resources of the subsequent optimization module on the regions that are initially judged to be the most relevant, ensuring that the joint encoder can concentrate on processing the features that are most likely to benefit from the previous-frame information and contribute to the high-quality segmentation result.

[0066] Accordingly, according to an embodiment of the present application, in step S533, inputting the fast interaction set of the previous frame Alpha mask semantic feature vector and the local semantic feature encoding vector of the video conference previous frame image into a multi-modal multi-scale joint encoder based on an LSTM-Transformer hybrid architecture to obtain a video conference current frame-previous frame local context semantic joint encoding vector and a video conference current frame-previous frame global context semantic joint encoding vector includes: performing universal correlation correction on each local semantic feature encoding vector of the video conference previous frame image in the fast interaction set of the local semantic feature encoding vectors of the video conference previous frame image to obtain a fast interaction set of corrected local semantic feature encoding vectors of the video conference previous frame image; inputting the previous frame Alpha mask semantic feature vector and the fast interaction set of the corrected local semantic feature encoding vectors of the video conference previous frame image into a multi-modal multi-scale joint encoder based on an LSTM-Transformer hybrid architecture to obtain a video conference current frame-previous frame local context semantic joint encoding vector and a video conference current frame-previous frame global context semantic joint encoding vector.

[0067] More specifically, perform universal correlation correction on each local semantic feature encoding vector of the video conference previous frame image in the fast interaction set of the local semantic feature encoding vectors of the video conference previous frame image to obtain a fast interaction set of corrected local semantic feature encoding vectors of the video conference previous frame image. It should be understood that the local semantic feature encoding vectors of the video conference previous frame image in the fast interaction set of the local semantic feature encoding vectors of the video conference previous frame image in the global cross-modal interaction field are substantially in the near neighbor region centered on the previous frame Alpha mask semantic feature vector . Then, while ensuring the near neighbor association characteristics between the local semantic feature encoding vector of the video conference previous frame image and the previous frame Alpha mask semantic feature vector , it is also expected that the local semantic feature encoding vectors of the video conference previous frame image can have proximity association universality, thereby facilitating the extraction of subsequent context fusion semantics.

[0068] In the first step, for each corresponding class magnetic attraction characteristic parameter , calculate its difference from the threshold , and integrate on the interval of to reflect the global system attribute of all vectors through the boundary attribute of each vector , that is, indicating the individual critical difference change It has universality that is immune to the interference of microscopic elements globally:

[0069]

[0070] Subsequently, using the universality-related structure of the polynomial solution of a single perturbation as the global perturbation, correct the local semantic feature encoding vector of each pre-frame image of the video conference :

[0071]

[0072] Among them, represents the th local semantic feature encoding vector of the corrected pre-frame image of the video conference.

[0073] That is, while keeping the local semantic feature encoding vector of each pre-frame image of the video conference as an autonomous distribution, determine its universality-related structure by means of algebraic equations, so as to realize the near-neighbor association universality between the local semantic feature encoding vectors of the pre-frame images of the video conference and boost the extraction of the subsequent context-fused semantics.

[0074] More specifically, input the fast interaction set of the pre-frame Alpha mask semantic feature vector and the local semantic feature encoding vector of the corrected pre-frame image of the video conference into a multi-modal multi-scale joint encoder based on the LSTM-Transformer hybrid architecture to obtain the local context semantic joint encoding vector of the current frame - pre-frame of the video conference and the global context semantic joint encoding vector of the current frame - pre-frame of the video conference, which is expressed by the formula:

[0075]

[0076]

[0077] Among them, is the local context semantic joint encoding vector of the current frame - pre-frame of the video conference, is the global context semantic joint encoding vector of the current frame - pre-frame of the video conference, is the multi-head attention mechanism.

[0078] It should be understood that through the foregoing steps, the global fingerprint representing the high-quality segmentation result of the previous frame (i.e., the "previous frame Alpha mask semantic feature vector") and a set of semantic information representing the key local regions of the current frame, which have been screened and corrected for general relevance (i.e., the "fast interaction set of the local semantic feature encoding vectors of the previous frame image in the video conference after correction") have been obtained. At this time, although the local features of the current frame have been preliminarily optimized and screened, to achieve the most accurate and robust segmentation of the current frame portrait, it is still necessary to perform deep and multi-dimensional interaction and fusion on these two information flows - the global information representing past precise experience and the local information representing the current optimized but still to be further integrated information. Simply splicing them or performing shallow fusion is difficult to fully explore the complex dependencies and complementarities between them, and thus cannot maximize the quality of the current frame segmentation mask. Therefore, in the technical solution of this application, by using the "multi-modal multi-scale joint encoder based on the LSTM-Transformer hybrid architecture" to deeply process the two input features, two key outputs are generated: one is the "local context semantic joint encoding vector of the current frame - previous frame in the video conference", which can finely capture the relationship between the local regions of the current frame, its context, and the prior of the previous frame; the other is the "global context semantic joint encoding vector of the current frame - previous frame in the video conference", which can integrate all information from a more macroscopic perspective to form a comprehensive understanding of the overall shape and temporal consistency of the current frame portrait. LSTM (Long Short-Term Memory Network) is good at capturing local and ordered context dependencies in sequential data. Even if the input "fast interaction set" itself may not be strictly ordered, LSTM can learn the short-term associations between its local features; while the core self-attention mechanism of Transformer can effectively capture the long-range dependencies between any elements within the sequence, achieve global context awareness, and powerfully model the complex interaction between the global prior of the "previous frame Alpha mask semantic feature vector" and each "corrected local semantic feature encoding vector". Through this hybrid architecture, the encoder aims to simultaneously utilize the local fine modeling ability of LSTM and the global and cross-modal (here referring to the information between the current frame and the previous frame) interaction modeling ability of Transformer to perform complex non-linear transformations and interaction modeling on the screened and corrected key information, learn the deep semantic associations between modalities (or between the previous and current frames in the time dimension), and design to output encoding results reflecting different granularity interaction information at the same time. These vectors will be used as the direct basis for generating the final "current frame Alpha mask image" in the subsequent process, which can improve the accuracy of the mask, reduce edge jaggedness and artifacts, effectively suppress the flickering and jumping of the background blurring effect, so that the background blurring in the video conference is both natural and soft, and stable and smooth, greatly enhancing the user's visual experience and professional image.

[0079] More specifically, according to an embodiment of the present application, in step S534, the local context semantic joint encoding vector of the current frame - previous frame of the video conference and the global context semantic joint encoding vector of the current frame - previous frame of the video conference are fused to obtain the semantic optimization representation vector of the current frame portrait of the video conference, which is represented by the formula:

[0080]

[0081] Wherein, and are the fusion weight matrix and the fusion bias vector respectively, represents vector concatenation, represents the semantic optimization representation vector of the current frame portrait of the video conference.

[0082] It should be understood that although in the previous step, two highly valuable but different - focused semantic representations have been successfully extracted and generated from different perspectives - one focuses on local details and context (local context semantic joint encoding vector), and the other focuses on the overall structure and global consistency (global context semantic joint encoding vector). However, although the information carried by these two vectors is rich, if we want to form the most comprehensive and accurate final semantic judgment of the current frame portrait and generate a high - quality Alpha mask based on this, it is necessary to effectively integrate the information of these two different scales. Therefore, the local context semantic joint encoding vector of the current frame - previous frame of the video conference and the global context semantic joint encoding vector of the current frame - previous frame of the video conference are further fused to obtain the semantic optimization representation vector of the current frame portrait of the video conference. This final semantic optimization representation vector of the current frame portrait of the video conference aims to achieve the most accurate and robust depiction of the semantic boundary of the current frame portrait by organically combining the accuracy of local details and the stability of the global structure. This fusion is not a simple splicing or averaging, but a more intelligent weighted combination, attention mechanism fusion, or other advanced feature fusion techniques, aiming to maximize the advantages of the two input vectors while suppressing the possible individual noises or biases they may have, so as to obtain a highly refined semantic optimization representation that is most instructive for the subsequent Alpha mask generation.

[0083] Specifically, in step S540, the current frame Alpha mask image is generated based on the semantic optimization representation vector of the portrait of the current frame of the video conference. It should be understood that although the "semantic optimization representation vector of the portrait of the current frame of the video conference" contains all the high-level abstract understandings about the portrait contour, details, and correlation with the previous frame, it itself is not a pixel-level mask that can be directly used for image processing. It needs to be "decoded" or "translated" into an image of the same size as the original video frame, with each pixel indicating the probability (i.e., Alpha value) of belonging to the foreground (portrait) or background. Therefore, in the technical solution of the present application, the current frame Alpha mask image is further generated. In particular, in an embodiment of the present application, based on the semantic optimization representation vector of the portrait of the current frame of the video conference, the current frame Alpha mask image is generated, including: passing the semantic optimization representation vector of the portrait of the current frame of the video conference through an Alpha mask image generator based on AIGC to obtain the current frame Alpha mask image. The "AIGC (AI Generated Content)" here uses advanced generative artificial intelligence technologies, such as complex decoder networks, the generator part of generative adversarial networks (GANs), or the reverse process of diffusion models. Its purpose is not just to simply convert vectors into images, but to use the powerful generation and detail restoration capabilities of the AIGC model to ensure that the generated Alpha mask is not only consistent with the instructions of the semantic vector in terms of macroscopic contours, but also achieves extremely high fidelity and refinement in microscopic details (such as hair edges, translucent areas, complex clothing texture edges, etc.). This is intended to overcome the problems of rough edges and loss of details that may arise from traditional segmentation methods, and generate an Alpha channel map that is almost consistent with the real physical occlusion relationship, with smooth transitions and fine boundaries.

[0084] Specifically, in step S600, a blurred background image is generated based on the current frame image of the video conference. It should be understood that if a preset static blurred image or an image unrelated to the current video content is directly used as the background, it is very easy to cause problems such as color imbalance and lighting mismatch, resulting in a stiff and unnatural synthesis effect, which seriously affects the user experience. By blurring the current frame image of the video conference itself to create a background, it can ensure that the blurred background is consistent with the foreground character in terms of color, brightness, and ambient light reflection. This is a key prerequisite for achieving a seamless and natural background blur effect. Even if the foreground character is accurately segmented, a blurred background that is inconsistent with the current scene will greatly reduce the overall visual effect.

[0085] More specifically, in a specific example of the present application, image filtering technology is used. First, a complete "current frame image of a video conference" is obtained as input. Then, one or more blur filters are applied to the image. Among them, the most commonly used is Gaussian Blur. In specific implementation, the image is subjected to a two-dimensional convolution operation with a Gaussian kernel, and the parameters of the Gaussian kernel (such as kernel size and standard deviation σ) determine the degree and range of blur. The larger the standard deviation, the stronger the blur effect, and the more details of the image are lost. Gaussian blur can produce a smooth and natural blur effect, which is consistent with the human eye's visual perception of out-of-focus scenes. In addition to Gaussian blur, mean blur (Mean Blur / Box Blur) or median blur (Median Blur) can also be used. The former achieves blur by calculating the average value of neighborhood pixels, and the latter takes the median of neighborhood pixels, which has a better suppression effect on salt and pepper noise, but for background blur, Gaussian blur is more popular because of its smooth and natural characteristics. These filtering operations can be efficiently implemented through built-in functions in mature computer vision libraries (such as OpenCV, PIL / Pillow). Developers only need to specify the input image and blur parameters to obtain the blurred image.

[0086] Specifically, in step S700, the current frame Alpha mask image is used as a weight, the current frame image of the video conference is used as the foreground, and the blurred background image is used as the background, and Alpha blending is performed to obtain a blurred background image of the current frame of the video conference. It should be understood that Alpha blending uses the pixel-by-pixel weight information provided by the Alpha mask to determine to what extent each pixel in the final image should display the foreground content and to what extent the background content should be displayed, so as to generate a seamlessly integrated "current frame blurred background image of the video conference" with a sense of depth and visual focus (i.e., foreground characters). Specifically, the goal is to use the Alpha mask as a guide to keep the original clarity and details of the area identified as the foreground in the original video frame (mainly the participants), and to smoothly and naturally superimpose it on the previously generated blurred background image. The change in grayscale value in the alpha mask (usually from 0 for full background to 1 for full foreground) allows for a soft transition at the edge of the foreground and background, which is crucial for dealing with complex boundaries such as hair and translucent objects, thereby avoiding a harsh "cut and paste" feeling, making the final synthesized image look more realistic and credible, and improving the overall visual quality and professionalism.

[0087] More specifically, in a specific example of the present application, the "over" operation in the classical Porter-Duff synthesis equation or direct linear interpolation is adopted. First, the system needs to ensure that the "current frame Alpha mask image" (assuming its pixel value is α, normalized to the [0, 1] interval), the "current frame image of the video conference" (as the foreground image F), and the "blurred background image" (as the background image B) are spatially aligned, that is, they have the same size and pixel correspondence. Then, for each pixel point (x, y) in the output image (i.e., the "current frame background blurred image" O) of the video conference, its color value O(x, y) can be calculated by the following formula: O(x, y) = α(x, y) * F(x, y) + (1 - α(x, y)) * B(x, y), and this calculation is performed independently for each color channel of the image (such as the R, G, B channels). When α(x, y) is 1, the output pixel is completely taken from the foreground image F(x, y); when α(x, y) is 0, the output pixel is completely taken from the background image B(x, y); when α(x, y) is between 0 and 1, the output pixel is a weighted average of the foreground and background pixel color values, achieving a smooth transition. This operation can be efficiently completed by functions provided by standard image processing libraries (such as OpenCV, Pillow, etc.), and these libraries have optimized such pixel-level operations internally.

[0088] In summary, the deep learning-based video conference background blurring method according to the embodiments of the present application is elucidated. It first performs a preliminary positioning of the fast portrait area in the current video frame to meet the requirements of real-time processing. Subsequently, a unique edge refinement and stabilization mechanism is introduced: this mechanism does not process each frame independently, but intelligently uses the Alpha mask that has been refined in the previous frame as an important reference and combines it with the preliminary segmentation result of the current frame. Through the image edge refinement processing module that combines image semantic understanding ability and prior knowledge, the portrait mask of the current frame is deeply optimized, and finally a current frame Alpha mask with clearer edges, richer details, and more coherent in time is generated, and then this mask image is used to generate the current frame background blurred image. This design can effectively smooth and correct the portrait edge, reduce the flicker and jump of the background blurring effect, thereby improving the visual realism and user experience while ensuring performance.

[0089] Furthermore, a deep learning-based video conference background blurring system is also provided.

[0090] Figure 5 It is a block diagram of the deep learning-based video conference background blurring system according to the embodiments of the present application. As Figure 5As shown, the deep learning-based video conferencing background blurring system 500 according to an embodiment of the present application includes: a video conferencing current frame acquisition module 510, which always acquires raw video conferencing video stream data through a camera and extracts the video conferencing current frame image from the raw video conferencing video stream data; a human portrait semantic segmentation module 520, configured to perform human portrait semantic segmentation on the video conferencing current frame image to obtain a video conferencing current frame human portrait rough segmentation mask image; a previous frame image acquisition module 530 of the conference, configured to acquire the previous frame image of the video conferencing current frame image to obtain a previous frame image of the video conference; a previous frame Alpha mask module 540, configured to perform Alpha masking on the previous frame image of the video conference to obtain a previous frame Alpha mask image; an image edge refinement processing module 550, configured to perform image edge refinement processing on the video conferencing current frame human portrait rough segmentation mask image and the previous frame Alpha mask image through a prior information-based image edge refinement processing module to obtain a current frame Alpha mask image; a blurred background image generation module 560, configured to generate a blurred background image based on the video conferencing current frame image; a conference background blurring image generation module 570, which always uses the current frame Alpha mask image as a weight, uses the video conferencing current frame image as the foreground, and uses the blurred background image as the background, and performs Alpha blending to obtain a video conferencing current frame background blurring image.

[0091] As described above, the deep learning-based video conferencing background blurring system 500 according to an embodiment of the present application can be implemented in various wireless terminals, such as a server with a deep learning-based video conferencing background blurring algorithm. In a possible implementation manner, the deep learning-based video conferencing background blurring system 500 according to an embodiment of the present application can be integrated into a wireless terminal as a software module and / or a hardware module. For example, the deep learning-based video conferencing background blurring system 500 can be a software module in the operating system of the wireless terminal, or can be an application program developed for the wireless terminal; of course, the deep learning-based video conferencing background blurring system 500 can also be one of the many hardware modules of the wireless terminal.

[0092] The embodiments of the present disclosure have been described above. The above description is exemplary and not exhaustive, and is not limited to the disclosed embodiments. Many modifications and variations are obvious to those of ordinary skill in the art in the technical field without departing from the scope and spirit of the described embodiments. The selection of the terms used herein is intended to best explain the principles of the embodiments, practical applications, or improvements to the technologies in the market, or to enable other ordinary technical personnel in the technical field to understand the embodiments disclosed herein.

Claims

1. A method for virtual background blurring in video conferencing based on deep learning, characterized in that, Including: Collecting original video conference video stream data through a camera, and extracting the current frame image of the video conference from the original video conference video stream data; Performing human portrait semantic segmentation on the current frame image of the video conference to obtain a rough segmentation mask image of the current frame of the video conference; Obtaining the previous frame image of the current frame image of the video conference to obtain the previous frame image of the video conference; Performing an Alpha mask on the previous frame image of the video conference to obtain a previous frame Alpha mask image; Passing the rough segmentation mask image of the current frame of the video conference and the previous frame Alpha mask image through an image edge refinement processing module based on prior information to obtain a current frame Alpha mask image; Generating a blurred background image based on the current frame image of the video conference; Using the current frame Alpha mask image as a weight, using the current frame image of the video conference as the foreground and the blurred background image as the background, and performing Alpha blending to obtain a background blurred image of the current frame of the video conference.

2. The method for virtual background blurring in a video conference based on deep learning according to claim 1, wherein Performing human portrait semantic segmentation on the current frame image of the video conference to obtain a rough segmentation mask image of the current frame of the video conference, including: passing the current frame image of the video conference through a human portrait semantic segmentation module based on MobileNet to obtain the rough segmentation mask image of the current frame of the video conference.

3. The method for virtual background of video conferencing based on deep learning according to claim 2, wherein, Passing the rough segmentation mask image of the current frame of the video conference and the previous frame Alpha mask image through an image edge refinement processing module based on prior information to obtain a current frame Alpha mask image, including: Extracting video conference semantic features from the rough segmentation mask image of the current frame of the video conference to obtain a semantic feature map of the previous frame image of the video conference; Extracting previous frame Alpha mask image semantic features from the previous frame Alpha mask image to obtain a previous frame Alpha mask semantic feature vector; Passing the semantic feature map of the previous frame image of the video conference and the previous frame Alpha mask semantic feature vector through an image semantic joint encoder assisted by prior information optimization to obtain a semantic optimized representation vector of the current frame of the video conference; Generating the current frame Alpha mask image based on the semantic optimized representation vector of the current frame of the video conference.

4. The method for virtual background blurring in video conferencing based on deep learning according to claim 3, wherein Extracting video conference semantic features from the rough segmentation mask image of the current frame of the video conference to obtain a semantic feature map of the previous frame image of the video conference, including: passing the rough segmentation mask image of the current frame of the video conference through a video conference semantic feature extractor based on a depthwise separable convolutional neural network model to obtain the semantic feature map of the previous frame image of the video conference.

5. The method for virtual background blurring in a video conference based on deep learning according to claim 4, wherein Extracting previous frame Alpha mask image semantic features from the previous frame Alpha mask image to obtain a previous frame Alpha mask semantic feature vector, including: passing the previous frame Alpha mask image through a previous frame Alpha mask image semantic feature extractor based on a ViT model to obtain the previous frame Alpha mask semantic feature vector.

6. The method for virtual background blurring in a video conference based on deep learning according to claim 5, wherein, The semantic feature map of the previous video conference frame image and the semantic feature vector of the previous frame Alpha mask are input into an image semantic joint encoder assisted by prior information optimization to obtain the optimized portrait semantic representation vector of the current video conference frame, including: Performing feature decoupling on the semantic feature map of the previous video conference frame image to obtain a set of local semantic feature encoding vectors of the previous video conference frame image; Based on the feature class magnetic attraction effect between the semantic feature vector of the previous frame Alpha mask and each local semantic feature encoding vector in the set of local semantic feature encoding vectors of the previous video conference frame image, determining a fast interaction set of local semantic feature encoding vectors of the previous video conference frame image; Inputting the semantic feature vector of the previous frame Alpha mask and the fast interaction set of local semantic feature encoding vectors of the previous video conference frame image into a multi-modal multi-scale joint encoder based on an LSTM-Transformer hybrid architecture to obtain a local context semantic joint encoding vector of the current-previous video conference frame and a global context semantic joint encoding vector of the current-previous video conference frame; Fusing the local context semantic joint encoding vector of the current-previous video conference frame and the global context semantic joint encoding vector of the current-previous video conference frame to obtain the optimized portrait semantic representation vector of the current video conference frame.

7. The method for blurring the background of a video conference based on deep learning according to claim 6, wherein Based on the feature class magnetic attraction effect between the semantic feature vector of the previous frame Alpha mask and each local semantic feature encoding vector in the set of local semantic feature encoding vectors of the previous video conference frame image, determining a fast interaction set of local semantic feature encoding vectors of the previous video conference frame image, including: Calculating the class magnetic attraction characteristic parameters between the semantic feature vector of the previous frame Alpha mask and each local semantic feature encoding vector in the set of local semantic feature encoding vectors of the previous video conference frame image to obtain a set of class magnetic attraction characteristic parameters of the current-previous video conference frame; Based on the comparison between the set of class magnetic attraction characteristic parameters of the current-previous video conference frame and a preset threshold, extracting the fast interaction set of local semantic feature encoding vectors of the previous video conference frame image from the set of local semantic feature encoding vectors of the previous video conference frame image.

8. The method for virtual background blurring in a video conference based on deep learning according to claim 7, characterized in that, Inputting the semantic feature vector of the previous frame Alpha mask and the fast interaction set of local semantic feature encoding vectors of the previous video conference frame image into a multi-modal multi-scale joint encoder based on an LSTM-Transformer hybrid architecture to obtain a local context semantic joint encoding vector of the current-previous video conference frame and a global context semantic joint encoding vector of the current-previous video conference frame, including: Performing universal correlation correction on each local semantic feature encoding vector in the fast interaction set of local semantic feature encoding vectors of the previous video conference frame image to obtain a fast interaction set of corrected local semantic feature encoding vectors of the previous video conference frame image; Input the fast interaction set of the previous frame Alpha mask semantic feature vector and the corrected local semantic feature encoding vector of the video conference previous frame image into the multi-modal multi-scale joint encoder based on the LSTM-Transformer hybrid architecture to obtain the current frame-previous frame local context semantic joint encoding vector and the current frame-previous frame global context semantic joint encoding vector of the video conference.

9. The method for blurring the background of a video conference based on deep learning according to claim 8, wherein Generate the current frame Alpha mask image based on the current frame portrait semantic optimization representation vector of the video conference, including: passing the current frame portrait semantic optimization representation vector of the video conference through the Alpha mask image generator based on AIGC to obtain the current frame Alpha mask image.

10. A video conferencing background blurring system based on deep learning, characterized in that, Including: A video conference current frame acquisition module, configured to collect original video conference video stream data through a camera and extract the video conference current frame image from the original video conference video stream data; A portrait semantic segmentation module, configured to perform portrait semantic segmentation on the video conference current frame image to obtain a video conference current frame portrait rough segmentation mask image; A previous frame image acquisition module of the conference, configured to acquire the previous frame image of the video conference current frame image to obtain the video conference previous frame image; A previous frame Alpha mask module, configured to perform Alpha masking on the video conference previous frame image to obtain a previous frame Alpha mask image; An image edge refinement processing module, configured to pass the video conference current frame portrait rough segmentation mask image and the previous frame Alpha mask image through an image edge refinement processing module based on prior information to obtain the current frame Alpha mask image; A blurred background image generation module, configured to generate a blurred background image based on the video conference current frame image; A conference background blurring image generation module, configured to use the current frame Alpha mask image as a weight, use the video conference current frame image as the foreground, and use the blurred background image as the background, and perform Alpha blending to obtain the video conference current frame background blurring image.

Citation Information

Patent Citations

  • Portrait background blurring method and device

    CN113538270A

  • Method, device and equipment for tracking moving target in video monitoring

    CN118864537A

  • Image matting model training method and device, image matting method and device, equipment and medium

    CN119653032A

  • Screen splicable image processing device

    CN216700148U