Video communication with self-adaptive lighting background
By generating a self-adaptive lighting background image based on environmental illumination, the method addresses inconsistencies in video communication systems, improving immersion and realism while optimizing resource usage.
Patent Information
- Application Number
- PCT/US2024/060890
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-31
- Filing Date
- 2024-12-19
- Publication Date
- 2025-08-07
AI Technical Summary
Video communication systems often produce disharmonious background images due to inconsistencies in illumination between the user's foreground and virtual background, leading to a lack of immersion and realism.
A method to generate a self-adaptive lighting background image by calculating illumination information from a reference video frame and using it to adjust a virtual background image to match the user's environment, ensuring harmony with the foreground image.
Enhances immersion and realism in video communication by adapting the background image to the user's environment, reducing the likelihood of unreasonable outputs and minimizing computing resource usage.
Smart Images

Figure US2024060890_07082025_PF_FP_ABST
Abstract
Description
VIDEO COMMUNICATION WITH SELF-ADAPTIVE LIGHTING BACKGROUNDBACKGROUND
[0001] With the development of digital device, communication technology, video processing technology, etc., people may use terminal devices such as desktop computers, tablet computers, smart phones, etc., to conduct video communication with people located elsewhere for purposes such as chatting, work discussion, remote training, technical support, etc. Herein, video communication may broadly refer to a communication method based on Internet technology that can transmit voice and images of participants in real time. Video communication may include, e.g., video conference, video call, video live streaming, etc.SUMMARY
[0002] This Summary is provided to introduce a selection of concepts that are further described below in the Detailed Description. It is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to be used to limit the scope of the claimed subject matter.
[0003] Embodiments of the present disclosure propose a method, apparatus and computer program product for video communication. A first image may be received, the first image being an original background image. The first image may be encoded to generate a first image representation. A reference video frame in a video stream of a user may be identified, the reference video frame being a first-extracted video frame from the video stream, or a video frame with a brightness difference from a previous video frame. Illumination information of an environment where the user is located may be calculated using the reference video frame. The first image representation may be decoded according to the illumination information, to generate a second image, the second image being a self-adaptive lighting background image.
[0004] It should be noted that the above one or more aspects comprise the features hereinafter fully described and particularly pointed out in the claims. The following description and the drawings set forth in detail certain illustrative features of the one or more aspects. These features are only indicative of the various ways in which the principles of various aspects may be employed, and this disclosure is intended to include all such aspects and their equivalents.BRIEF DESCRIPTION OF THE DRAWINGS
[0005] The disclosed aspects will hereinafter be described in connection with the appended drawings that are provided to illustrate and not to limit the disclosed aspects.
[0006] FIG. 1 illustrates an exemplary process for video communication with self-adaptive lighting background according to an embodiment of the present disclosure.
[0007] FIG. 2 illustrates an exemplary process for identifying a reference video frame in avideo stream of a user according to an embodiment of the present disclosure.
[0008] FIG. 3A illustrates an exemplary process for preprocessing a reference video frame according to an embodiment of the present disclosure.
[0009] FIG. 3B illustrates another exemplary’ process for preprocessing a reference video frame according to an embodiment of the present disclosure.
[0010] FIG. 4A to FIG. 4C illustrate an example of preprocessing a reference video frame according to an embodiment of the present disclosure.
[0011] FIG. 5 illustrates an exemplary' process for generating a second image according to an embodiment of the present disclosure.
[0012] FIG. 6A to FIG. 6B illustrate an example of generating a second image according to an embodiment of the present disclosure.
[0013] FIG. 7 is a flowchart of an exemplary method for video communication according to an embodiment of the present disclosure.
[0014] FIG. 8 illustrates an exemplary apparatus for video communication according to an embodiment of the present disclosure.
[0015] FIG. 9 illustrates another exemplary’ apparatus for video communication according to an embodiment of the present disclosure.DETAILED DESCRIPTION
[0016] The present disclosure will now be discussed with reference to several example implementations. It is to be understood that these implementations are discussed only for enabling those skilled in the art to better understand and thus implement the embodiments of the present disclosure, rather than suggesting any limitations on the scope of the present disclosure.
[0017] Video communication may be implemented through communicating speeches and images among electronic devices of users participating in the video communication. During video communication, for purposes of protecting privacy, improving concentration or increasing personalization, a user may set a virtual background image as the background of his or her video frame. The virtual background image is usually a two-dimensional image, which may be a preset image provided by' a provider of the video communication service, or a custom image uploaded by the user. During the video communication of the user, a foreground image containing a user image of the user may be extracted from a video frame of the user captured in real-time, and the extracted foreground image may be combined with the virtual background image set by the user into an updated video frame. The updated video frame may then be presented at a device of the user and devices of other users participating in the video communication. Due to differences in illumination situations when capturing images or rendering approaches when producing images, the illumination situation presented by the foreground image may be inconsistent with theillumination situation presented by the virtual background image. Accordingly, the video frame produced by combining the foreground image and the virtual background image may be disharmonious.
[0018] Embodiments of the present disclosure propose video communication with self- adaptive lighting background. A first image may be received. The first image may be an original background image. The original background image may be a preset image provided by a provider of the video communication service, or a custom image uploaded by the user. Since this image is not the actual background image of the user, it may be referred to as a virtual background image. Illumination information of an environment where the user is located may be calculated using a video frame extracted from a video stream of the user. The video stream may be captured by a device used by the user to participate in the video communication. Herein, a video frame used to calculate illumination information of an environment where a user is located is referred to as a reference video frame. The reference video frame may include a foreground image containing a user image of the user and an actual background image reflecting an environment where the user is actually located. Illumination information of the environment where the user is located may be calculated based on the reference video frame or the background image obtained from the reference video frame through a control network containing a light extractor. The calculated illumination information may be taken as a control condition to guide an image generation model to generate a second image based on the first image. Preferably, the second image may explicitly include a light source, and an attribute of the light source may be the same or similar to an attribute of a light source in the reference video frame. The attribute of the light source may include, e.g., position and shape of the light source, intensity’, chromaticity, direction, and range of light emitted by the light source, etc. The technical effect of the above approach is that the second image whose illumination situation automatically adapts to the foreground image containing the user image of the user can be generated, thereby making the background image of the video frame harmonious with the foreground image. The illumination situation may include at least one of intensity, chromaticity, direction, and range of light. Accordingly, the second image may be referred to as a self-adaptive lighting background image. Presenting such a background image during video communication may enhance the immersion and realism of users participating in the video communication.
[0019] The first image may be a preset image provided by the provider of the video communication service, or a custom image uploaded by the user. There may be some images among these images that are not suitable for modifying illumination situations. For example, when the first image is an advertising poster of the organization of the user, it may be unreasonable and unnecessary to modify the illumination situations of the image. In view of this, the embodimentof the present disclosure propose that when it is detected that a virtual background image is used in video communication, that is, when the first image is detected, it is determined whether the first image is an image suitable for modifying illumination situations, and only when it is determined that the first image is an image suitable for modifying illumination situations, a second image, i.e., a self-adaptive lighting background image, is generated based on the first image. The technical effect of the above approach is to reduce the probability of outputting an unreasonable background image and reduce computing resource usage.
[0020] During the video communication, an illumination situation of the environment where the user is actually located may change. For example, when the user participates in the video communication in a conference room, lights in the conference room may be switched on or switched off as the outside brightness changes. Therefore, the illumination situation in the conference room will change. In view of this, the embodiments of the present disclosure propose to detect changes in the illumination situation of the environment where the user is located in real time, and regenerate the second image when a change in the illumination situation is detected. For example, a video frame may be extracted from the video stream of the video communication at a predetermined time interval. The video stream may be captured by the device used by the user to participate in the video communication. If the extracted video frame is a first-extracted video frame, the video frame may be taken as a reference video frame for calculating illumination information of the environment where the user is located, and the second image may be generated based on the calculated illumination information. If the extracted video frame being a non-first- extracted video frame, a brightness of the extracted video frame is detected. If it is determined that a brightness difference between the brightness and a brightness of a previous reference video frame exceeds a predetermined threshold, the video frame may be taken as a reference video frame, and the second image may be regenerated based on the illumination information that is calculated using the reference video frame; while if it is determined that the brightness difference between the brightness and a brightness of a previous reference video frame does not exceed the predetermined threshold, the current second image is maintained. The technical effect of the above approach is to enable the background image of the user to automatically adapt to illumination changes in the environment where the user is located, and maintain the harmonious between the background image and the foreground image of the video frame. In addition, the technical effect of determine whether the illumination situation of the environment where the user is located have changed through detecting whether there is a brightness difference between the extracted video frame and the previous reference video frame is to reduce computing resource usage. This is because the operation of detecting the brightness of the video frame requires lower computing resources compared to the operation of calculating the illumination information corresponding tothe video frame.
[0021] Various embodiments of the present disclosure will hereinafter be described in detail in connection with the appended drawings.
[0022] FIG. 1 illustrates an exemplary’ process 100 for video communication with self- adaptive lighting background according to an embodiment of the present disclosure.
[0023] At 102, a first image may be received. The first image may be an original background image. The first image may be set by a user when video communication starts, or preset by the video communication sendee. The video communication may include, e.g., video conference, video call, video live streaming, etc.
[0024] The first image may be a preset image provided by a provider of the video communication service, or a custom image uploaded by the user. There may be some images among these images that are not suitable for modifying illumination situations. For example, when the first image is an advertising poster of the organization of the user, it may be unreasonable and unnecessary to modify the illumination situations of the image. In view of this, the embodiments of the present disclosure propose that when it is detected that a virtual background image is used in video communication, that is, when the first image is detected, it is determined whether the first image is an image suitable for modifying illumination situations, and only when it is determined that the first image is an image suitable for modifying illumination situations, a second image, i.e., a self-adaptive lighting image, is generated based on the first image. At 104, it may be determined whether the first image is an image suitable for modifying illumination situations. As an example, an image suitable for modifying illumination situations may be a scenario-type image, e.g., photos or pictures depicting indoor arrangements, outdoor places, natural landscapes, etc. As another example, an image suitable for modifying illumination situations may be a science fiction image containing light effect elements. As yet another example, the image suitable for modifying illumination situations may be a blank image without objects, e.g., a black image, a white image, etc. It may be determined whether the first image is an image suitable for modifying illumination situations through a trained image classification model based on computer vision technology.
[0025] If it is determined that the first image is not an image suitable for modifying illumination situations at 104, the process 100 proceeds to a step 118. At 118, a foreground image may be combined with the first image into an updated video frame. The foreground image may include a user image of the user, which is extracted from a video stream of the video communication. The video stream is captured by a device used by the user to participate in the video communication. The updated video frame may then be presented at a device of the user and devices of other users participating in the video communication. That is, if it is determined that it is not suitable for performing self-adaptive lighting operation on the first image, the second imagewill not be generated based on the first image. The technical effect of the above approach is to reduce the probability of outputting an unreasonable background image and reduce computing resource usage.
[0026] If at 104, it is determined that the first image is an image suitable for modifying illumination situations, one or more second images may be generated based on the first image. Illumination information of an environment where the user is located may be obtained first, and a second image for video communication may be generated based on the obtained illumination information and the first image. The illumination information of the environment where the user is located may be calculated using a reference video frame in the video stream of the video communication. The process 100 may proceed to a step 106. At 106, a reference video frame in the video stream of the user may be identified. The video stream includes a number of video frames captured at a certain frame rate. The reference video frame may be a first-extracted video frame from the video stream, or a video frame with a brightness difference from a previous video frame. An exemplary process for identifying a reference video frame will be described later in conjunction with FIG. 2. It should be appreciated that the step 104 is optional. In the case where the step 104 is not performed, the process 100 may proceed directly to the step 106 after the step 102.
[0027] After the reference video frame is identified, at 108, a second image for video communication may be generated based on the reference video frame and the first image. One or more second images may be generated. The second image is a self-adaptive lighting background image. For example, illumination information of an environment where the user is located may be calculated using the reference video frame, and the second image may be generated based on the illumination information and the first image. An exemplary process for generating the second image will be described later in conjunction with FIG. 5.
[0028] At 110, the one or more second images generated at 108 may be presented to the user.
[0029] At 112, a selection of a second image by the user may be detected. If a selection of a second image by the user is detected at 112, the process 100 proceeds to a step 114. At 114, a foreground image may be combined with the second image selected by the user into an updated video frame. If no selection of a second image by the user is detected at 112, i.e., the user does not select any of the one or more generated second images, then the process 100 proceeds to a step 116. At 116. it may be determined whether there is a second image currently in use. If it is determined at 116 that there is no second image currently in use, the process 100 proceeds to a step 118. At 118, the foreground image may be combined with the first image into an updated video frame. If it is determined at 116 that there is a second image currently in use, the process 100 proceeds to a step 120. At 120. the foreground image may be combined with the second imagecurrently in use into an updated video frame. The updated video frame may then be presented at a device of the user or devices of other users participating in the video communication.
[0030] It should be appreciated that the step 110 to the step 112 are optional. In the case where the step 110 and the step 112, the process 100 may proceed directly to the step 114 after the step 110, and the step 116 and the step 120 may not be performed. For example, at 108, only one second image may be generated. The second image may be combined directly with the foreground image into an updated video frame without user confirmation.
[0031] It should be appreciated that the process for video communication with self-adaptive lighting background described above in conjunction with FIG. 1 is merely exemplary. Depending on actual application requirements, the steps in the process for video communication with self- adaptive lighting background may be replaced or modified in any manner, and the process may comprise more or fewer steps. In addition, the specific order or hierarchy of the steps in the process 100 is merely exemplary’, and the process for video communication with self-adaptive lighting background may be performed in an order different from the described order.
[0032] FIG. 2 illustrates an exemplary process 200 for identifying a reference video frame in a video stream of video communication according to an embodiment of the present disclosure. The process 200 may correspond to the step 106 in FIG. 1.
[0033] At 202. a video frame may be extracted from a video stream of video communication at a predetermined time interval. The video stream may be captured by a device used by a user to participate in the video communication. For example, one video frame may be extracted from the video stream every a few seconds.
[0034] At 204. it may be determined whether the extracted video frame is a first-extracted video frame.
[0035] If it is determined at 204 that the extracted video frame is a first-extracted video frame, then the process 200 proceeds to a step 210. At 210, the extracted video frame may be taken as a reference video frame. The reference video frame may be used to calculate illumination information of an environment where the user is located. The calculated illumination information may be further used to generate a second image, i.e., a self-adaptive lighting background image.
[0036] If it is determined at 204 that the extracted video frame is a non-first-extracted video frame, then the process 200 proceeds to a step 206. At 206, a brightness of the extracted video frame may be detected. For example, each pixel in the video frame may be represented by YUV color encoding. The value corresponding to the Y channel of each pixel may be obtained, i.e., the brightness (Luminance or Luma) value. The brightness of the video frame may be calculated through averaging the brightness values of all pixels in the video frame. It should be appreciated that the method for detecting the brightness of the extracted video frame described above is merelyexemplary, and the brightness of the extracted video frame may be detected through other methods.
[0037] At 208, it may be determined whether a brightness difference between the brightness and a brightness of a previous reference video frame exceeds a predetermined threshold.
[0038] If it is determined at 208 that the brightness difference between the brightness and the brightness of the previous reference video frame does not exceed the predetermined threshold, the process 200 returns to the step 202, that is, after waiting for a predetermined time, a next video frame is extracted from the video stream. In other words, the current extracted video frame will not be taken as a reference video frame. Accordingly, the second image will not be regenerated, i.e., the current second image will be maintained.
[0039] If it is determined at 208 that the brightness difference between the brightness and the brightness of the previous reference video frame exceeds the predetermined threshold, the process 200 proceeds to a step 210. At 210, the extracted video frame may be taken as a reference video frame. The reference video frame may be used to re-calculate the illumination information of the environment where the user is located. The calculated illumination information may be further used to regenerate the second image. After the step 210, the process 200 returns to the step 202, that is, after waiting for a predetermined time, a next video frame is extracted from the video stream.
[0040] During the video communication, an illumination situation of the environment where the user is actually located may change. For example, when the user participates in the video communication in a conference room, lights in the conference room may be switched on or switched off as the outside brightness changes. Therefore, the illumination situation in the conference room will change. The process 200 may detect, in real time, changes in illumination situation of the environment where the user is actually located. For example, if it is determined that the brightness difference between the brightness of the video frame extracted from the video stream and the brightness of the previous reference video frame exceeds the predetermined threshold, it may indicate that the illumination situation of the environment where the user is actually located have changed. In this case, the video frame may be taken as the reference video frame, and the second image may be regenerated based on the illumination information that is calculated using the reference video frame. If it is determined that the brightness difference between the brightness and the brightness of the previous reference video frame does not exceed the predetermined threshold, the current second image is maintained. The technical effect of the above approach is to enable the background image of the user to automatically adapt to illumination changes in the environment where the user is located, and maintain the harmonious between the background image and the foreground image of the video frame. In addition, thetechnical effect of determine whether the illumination situation of the environment where the user is located have changed through detecting whether there is a brightness difference between the extracted video frame and the previous reference video frame is to reduce computing resource usage. This is because the operation of detecting the brightness of the video frame requires lower computing resources compared to the operation of calculating the illumination information corresponding to the video frame.
[0041] It should be appreciated that the process for identifying the reference video frame in the video stream of video communication described above in conjunction with FIG. 2 is merely exemplary. Depending on actual application requirements, the steps in the process for identifying the reference video frame may be replaced or modified in any manner, and the process may comprise more or fewer steps.
[0042] As described above in conjunction with FIG. 1, the second image for video communication may be generated based on the reference video frame and the first image. For example, illumination information of an environment where the user is located may be calculated using the reference video frame, and the second image may be generated based on the illumination information and the first image. Preferably, when calculating the illumination information of the environment where the user is located, the reference video frame may be preprocessed, and the illumination information of the environment where the user is located may be calculated using the preprocessed video frame.
[0043] FIG. 3A illustrates an exemplary process 300a for preprocessing a reference video frame according to an embodiment of the present disclosure.
[0044] A reference video frame 302 is, e.g., a reference video frame extracted from a video stream of a user through the process 200 in FIG. 2. The video stream may be captured by a device used by the user to participate in video communication. FIG. 4A to FIG. 4C illustrate an example of preprocessing a reference video frame according to an embodiment of the present disclosure. A reference video frame 400a shown in FIG. 4A may be an example of the reference video frame 302. A switched-on ceiling lamp 402 located at the ceiling of the room is shown in the reference video frame 400a. Accordingly, in a foreground image 404 corresponding to a user image in the reference video frame 400a, the top of the user's head is brighter relative to other parts. Additionally, a door 406 behind the user is shown in the reference video frame 400a.
[0045] The reference video frame 302 may be preprocessed through a preprocessing module 310. The preprocessing module 310 may comprise a background image extracting module 320. An initial background image 322 may be extracted from the reference video frame 302 through the background image extracting module 320. Since the initial background image 322 corresponds to the actual background of the user, the initial background image 322 may be referred to as anactual background image of the user. For example, the initial background image 322 may be extracted from the reference video frame 302 through methods such as foreground and background segmentation, matting, etc. The initial background image 322 may be taken as a preprocessed video frame. The initial background image 322 is incomplete, missing parts that are occluded by the foreground image. An initial background image 400b shown in FIG. 4B may be an example of the initial background image 322. As may be seen from the initial background image 400b, the door 406 in the initial background image 400b is incomplete due to the occlusion by the foreground image 404.
[0046] FIG. 3B illustrates another exemplary process 300b for preprocessing a reference video frame according to an embodiment of the present disclosure.
[0047] In the process 300b, a preprocessing module 330 may comprise both a background image extracting module 320 and an image inpainting model 340.
[0048] After the initial background image 322 is extracted from the reference video frame 302 through the background image extracting module 320, the initial background image 322 may be inpainted through the image inpainting model 340, to obtain an inpainted background image 342. The image inpainting model 340 may be a model that can take advantage of the redundancy of an image itself and use information from known parts of the image to inpaint unknown parts. The inpainted background image 342 may be taken as a preprocessed video frame. An inpainted background image 400c shown in FIG. 4C may be an example of the inpainted background image 342. In the inpainted background image 400c, the door 406 has been inpainted.
[0049] Through the process 300a and the process 300b, the initial background image 322 directly extracted from the reference video frame 302 may be obtained, or the inpainted background image 342 obtained through performing image inpainting on the initial background image 322 may be obtained. The initial background image 322 and the inpainted background image 342 may be used to calculate illumination information of an environment w here a user is located. For example, the illumination information of the environment where the user is located may be calculated based on the initial background image 322 or the inpainted background image 342 through a control network containing a light extractor. Neither the initial background image 322 nor the inpainted background image 342 contains a foreground image. The technical effect of the above approach is that since training data used to train the control network is usually pictures that do not contain people and have complete backgrounds, therefore, using a preprocessed video frame, especially the inpainted background image 342, to calculate the illumination information of the environment where the user is located helps the control network to obtain more accurate results.
[0050] It should be appreciated that the processes for preprocessing the reference video framedescribed above in conjunction with FIG. 3A and FIG. 3B as well as FIG. 4A and FIG. 4C are merely exemplary. Depending on actual application requirements, the steps in the process for preprocessing the reference video frame may be replaced or modified in any manner, and the process may comprise more or fewer steps.
[0051] FIG. 5 illustrates an exemplary process 500 for generating a second image according to an embodiment of the present disclosure. The second image is a self-adaptive lighting background image. The process 500 may correspond to the step 108 in FIG. 1.
[0052] A reference video frame 502 is, e.g., a reference video frame extracted from a video stream of a user through the process 200 in FIG. 2. The video stream is captured by a device used by the user to participate in video communication. Illumination information 522 of an environment where the user is located may be calculated using the reference video frame 502. The environment where the user is located is usually a three-dimensional space.
[0053] In an implementation, the illumination information 522 may be calculated directly based on the reference video frame 502. The illumination information 522 of the environment where the user is located may be calculated based on the reference video frame 502 through a control network 520. The control network 520 may contain a light extractor 530. The light extractor 530 may be a machine learning model based on a convolutional neural network, which can calculate illumination information of the environment where the user is located based on the reference video frame 502. The initial illumination information may reflect brightness, light distribution, light chromaticity, shadow distribution, etc., of the three-dimensional space. The control network 520 may also contain an encoder 540 and a zero convolution layer 550. The encoder 540 and the zero convolution layer 550 may encode the initial illumination information, to produce illumination information 522 used to control an image generation model 560.
[0054] In another implementation, the reference video frame 502 may be preprocessed to obtain a preprocessed video frame 512, and the illumination information 522 may be calculated based on the preprocessed video frame 512. The reference video frame 502 may be preprocessed through a preprocessing module 510. The preprocessing module 510 may correspond to the preprocessing module 310 in FIG. 3A or the preprocessing module 330 in FIG. 3B. The preprocessed video frame 512 may correspond to the initial background image 322 in FIG. 3 A or the inpainted background image 342 in FIG. 3B. After the preprocessed video frame 512 is obtained, the illumination information 522 of the environment where the user is located may be calculated based on the preprocessed video frame 512 through the control network 520. The light extractor 530 may calculate illumination information of the environment where the user is located based on the preprocessed video frame 512. The initial illumination information may reflect brightness, light distribution, light chromaticity, shadow distribution, etc., of the three-dimensional space. The encoder 540 and the zero convolution layer 550 may encode the initial illumination information, to produce illumination information 522 used to control the image generation model 560. The preprocessed video frame does not contain a foreground image. The technical effect of the above approach is that since training data used to train the control network 520 is usually pictures that do not contain people and have complete backgrounds, therefore, using the preprocessed video frame 512, especially the inpainted background image, to estimate the illumination information of the environment where the user is located helps the control network 520 to obtain more accurate results.
[0055] A first image 504 may be an original background image. The first image 504 may be a preset image provided by a provider of the video communication service, or a custom image uploaded by the user.
[0056] The illumination information 522 may be taken as a control condition to guide the image generation model 560 to generate a second image 562 for video communication based on the first image 504. The image generation model 560 may be a model capable of generating a desired image based on an image and / or text, e.g., a StableDiffusion model, a DALL E 2 model, etc. The image generation model 560 may comprise an encoder 570 and a decoder 580. The encoder 570 may encode the first image 504 into a first image representation. The first image representation is in vector form. The decoder 580 may decode the first image representation of the first image 504 according to the illumination information 522, to generate one or more second images 562.
[0057] Preferably, in addition to providing the illumination information 522 to the decoder 580, a prompt 506 may also be provided to the decoder 580. The prompt 506 may be constructed based on a description of the first image 504 and / or an image generation requirement. The description of the first image 504 may be automatically generated through a model based on computer vision technology, or may be manually written. As an example, the description of the first image 504 may be "authentic photo of a comfortable room.” The image generation requirement may include a clarity requirement, a size requirement, etc., for the second image 562 to be generated. As an example, the image generation requirement could be "8K UHD". With the prompt 506 provided, the decoder 580 may decode the first image representation of the first image 504 according to both the illumination information 522 and the prompt 506, to generate one or more second images 562.
[0058] In the process 500, the illumination information of the environment where the user is located may be calculated using the reference video frame identified from the video stream captured by the device used by the user to participate in the video communication. The reference video frame may include a foreground image containing a user image of the user and a backgroundimage reflecting the environment where the user is located. The illumination information of the environment where the user is located may be calculated based on the reference video frame or the background image obtained from the reference video frame through the control network containing the light extractor. The calculated illumination information may be taken as a control condition to guide the image generation model to generate the second image based on the first image. Preferably, the second image may explicitly include a light source, and an attribute of the light source may be the same or similar to an attribute of a light source in the reference video frame. The attribute of the light source may include, e.g.. position and shape of the light source, intensity, chromaticity, direction, and range of light emitted by the light source, etc. The technical effect of the above approach is that the second image whose illumination situation automatically adapts to the foreground image containing the user image of the user can be generated, thereby making the background image of the video frame harmonious with the foreground image. The illumination situation may include at least one of intensity, chromaticity, direction, and range of light. Accordingly, the second image may be referred to as a self-adaptive lighting background image. Presenting such a background image during video communication may enhance the immersion and realism of users participating in the video communication.
[0059] It should be appreciated that the process for generating the second image described above in conjunction with FIG. 5 is merely exemplary. Depending on actual application requirements, the steps in the process for generating the second image may be replaced or modified in any manner, and the process may comprise more or fewer steps. For example, in the case where the control network 520 directly calculates the illumination information 522 of the environment where the user is located based on the reference video frame 502, the step of preprocessing the reference video frame 502 to obtain the preprocessed video frame 512 may be omitted. Additionally, the prompt 506 is optional. Furthermore, the specific order or hierarchy of the steps in the process 500 is merely exemplary, and the process for generating the second image may be performed in an order different from the described order.
[0060] FIG. 6A to FIG. 6B illustrate an example of generating a second image according to an embodiment of the present disclosure. The second image is a self-adaptive lighting background image.
[0061] FIG. 6A shows a first image 600a. The first image 600a is an original background image. The first image 600a includes a switched-on floor lamp 602 located on the left side of the screen. Assume that first image 600a will be combined with the foreground image 404 in the reference video frame 400a in FIG. 4A. If the first image 600a and the foreground image 404 are directly combined, the resulting updated video frame may be unreasonable and disharmonious. The reason is that in the foreground image 404, the top of the user's head is brighter relative toother parts; and the floor lamp 602 located on the left side of the first image 600a should make the side of the user close to the floor lamp brighter relative to other parts.
[0062] A self-adaptive lighting operation may be performed on the first image 600a to generate a second image. The second image is a self-adaptive lighting background image. FIG. 6B illustrate a second image 600b generated according to the embodiments of the present disclosure. The second image 600b may be generated, through the process 500 shown in FIG. 5, based on the first image 600a in FIG. 6A and any of the reference video frame 400a in FIG. 4A, the initial background image 400b in FIG. 4B, and the inpainted background image 400c in FIG. 4C. For example, any one of the reference video frame 400a, the initial background image 400b, and the inpainted background image 400c may be used to calculate the illumination information of the environment where the user is located. The calculated illumination information and / or a pre-constructed prompt may be used as a control condition to guide an image generation model to generate the second image 600b based on the first image 600a.
[0063] In the reference video frame 400a, the initial background image 400b, or the inpainted background image 400c, a light source is shown, i.e., a switched-on ceiling lamp 402 located at the ceiling of the room. Accordingly, in the second image 600b generated according to the embodiments of the present disclosure, a light source similar to that in the reference video frame 400a. the initial background image 400b, or the inpainted background image 400c is presented, i.e , a switched-on ceiling lamp 604 located at the ceiling of the room. The ceiling lamp 604 is not shown in the first image 600a. Some of the attributes of the ceiling lamp 604 are the same or similar to corresponding attributes of the ceiling lamp 402. For example, the position of the ceiling lamp 604 in the image is substantially the same as the position of the ceiling lamp 402 in the image, the direction of the light emitted by the ceiling lamp 604 is substantially the same as the direction of the light emitted by the ceiling lamp 402, etc. In addition, compared with the first image 600a, the floor lamp 602 located on the left side of the screen included in the first image 600a is deleted in the second image 600b.
[0064] The second image 600b may be combined with the foreground image 404 into an updated video frame. The updated video frame may then be presented at a device of the user and devices of other users participating in the video communication. Since the second image 600b includes a light source adapted to the foreground image 404, i.e., the switched-on ceiling lamp 604 located at the ceiling of the room, thus the illumination situation of the second image 600b may be adapted to the illumination situation of the foreground image 404. Accordingly, combining the second image 600b with the foreground image 404 can produce a harmonious video frame.
[0065] It should be appreciated that the light source, i.e., the ceiling lamp 604, is explicitly shown in the second image 600b, but the light source not mandatory. There may also be otherways to produce an illumination situation that is adapted to the foreground image 404 in the reference video frame 400a. As an example, if the first image is an outdoor scene on a cloudy day, the generated second image may be an image including a bright sky. As another example, if the first image is an undersea scene, the generated second image may be an image including light rays transmitted from above.
[0066] FIG. 7 is a flowchart of an exemplary method 700 for video communication according to an embodiment of the present disclosure.
[0067] At 710, a first image may be received, the first image being an original background image.
[0068] At 720, the first image may be encoded to generate a first image representation.
[0069] At 730, a reference video frame in a video stream of a user may be identified, the reference video frame being a first-extracted video frame from the video stream, or a video frame with a brightness difference from a previous video frame.
[0070] At 740, illumination information of an environment where the user is located may be calculated using the reference video frame.
[0071] At 750, the first image representation may be decoded according to the illumination information, to generate a second image, the second image being a self-adaptive lighting background image.
[0072] In an implementation, the method 700 may further comprise: determining that the first image is an image suitable for modifying illumination situations. The identifying a reference video frame in a video stream of a user may comprise: in response to determining that the first image is an image suitable for modifying illumination situations, identifying the reference video frame.
[0073] In an implementation, the identifying a reference video frame in a video stream of a user may comprise: extracting a video frame from the video stream at a predetermined time interval; in response to the extracted video frame being a first-extracted video frame, taking the extracted video frame as the reference video frame; in response to the extracted video frame being a non-first-extracted video frame, detecting a brightness of the extracted video frame; and in response to a brightness difference between the brightness and a brightness of a previous reference video frame exceeding a predetermined threshold, taking the extracted video frame as the reference video frame.
[0074] In an implementation, the calculating illumination information of an environment where the user is located may comprise: calculating the illumination information based on the reference video frame through a control network containing a light extractor.
[0075] In an implementation, the calculating illumination information of an environment where the user is located may comprise: preprocessing the reference video frame, to obtain apreprocessed video frame, the preprocessing comprising: extracting a background image from the reference video frame, or extracting a background image from the reference video frame, and inpainting the extracted background image through image inpainting; and calculating the illumination information based on the preprocessed video frame through a control network containing a light extractor.
[0076] In an implementation, the method 700 may further comprise: constructing a prompt based on a description of the first image and / or an image generation requirement. The generating a second image may comprise: decoding the first image representation according to the illumination information and the prompt, to generate a second image.
[0077] In an implementation, the reference video frame may include a foreground image. An illumination situation of the second image may be adapted to the foreground image. The illumination situation may include at least one of intensity7, chromaticity, direction, and range of light.
[0078] In an implementation, the second image may include a light source. An attribute of the light source may be the same or similar to an attribute of a light source in the reference video frame.
[0079] In an implementation, the method 700 may further comprise: detecting a selection of the second image by the user; and in response to detecting the selection of the second image by the user, combining a foreground image from the video stream with a selected second image into an updated video frame.
[0080] It should be appreciated that the method 700 may further comprise any other step / process for video communication according to the embodiments of the present disclosure as mentioned above.
[0081] FIG. 8 illustrates an exemplary' apparatus 800 for video communication according to an embodiment of the present disclosure.
[0082] The apparatus 800 may comprise: a first image receiving module 810, for receiving a first image, the first image being an original background image; a first image encoding module 820, for encoding the first image to generate a first image representation; a reference video frame identifying module 830, for identifying a reference video frame in a video stream of a user, the reference video frame being a first-extracted video frame from the video stream, or a video frame with a brightness difference from a previous video frame; an illumination information calculating module 840, for calculating illumination information of an environment where the user is located using the reference video frame; and a second image generating module 850, for decoding the first image representation according to the illumination information, to generate a second image, the second image being a self-adaptive lighting background image. Furthermore, the apparatus 800may further comprise any other modules configured for video communication according to the embodiments of the present disclosure as mentioned above.
[0083] FIG. 9 illustrates another exemplary' apparatus 900 for video communication according to an embodiment of the present disclosure.
[0084] The apparatus 900 may comprise a processor 910; and a memory 920 storing computer-executable instructions. The computer-executable instructions, when executed, may cause the processor 910 to: receive a first image, the first image being an original background image; encode the first image to generate a first image representation; identify a reference video frame in a video stream of a user, the reference video frame being a first-extracted video frame from the video stream, or a video frame with a brightness difference from a previous video frame; calculate illumination information of an environment where the user is located using the reference video frame; and decode the first image representation according to the illumination information, to generate a second image, the second image being a self-adaptive lighting background image.
[0085] In an implementation, the computer-executable instructions, when executed, may further cause the processor 910 to: determine that the first image is an image suitable for modifying illumination situations. The identifying a reference video frame in a video stream of a user may comprise: in response to determining that the first image is an image suitable for modifying illumination situations, identifying the reference video frame.
[0086] In an implementation, the identifying a reference video frame in a video stream of a user may comprise: extracting a video frame from the video stream at a predetermined time interval; in response to the extracted video frame being a first-extracted video frame, taking the extracted video frame as the reference video frame; in response to the extracted video frame being a non-first-extracted video frame, detecting a brightness of the extracted video frame; and in response to a brightness difference between the brightness and a brightness of a previous reference video frame exceeding a predetermined threshold, taking the extracted video frame as the reference video frame.
[0087] In an implementation, the calculating illumination information of an environment where the user is located may comprise: calculating the illumination information based on the reference video frame through a control network containing a light extractor.
[0088] In an implementation, the calculating illumination information of an environment where the user is located may comprise: preprocessing the reference video frame, to obtain a preprocessed video frame, the preprocessing comprising: extracting a background image from the reference video frame, or extracting a background image from the reference video frame, and inpainting the extracted background image through image inpainting; and calculating the illumination information based on the preprocessed video frame through a control networkcontaining a light extractor.
[0089] In an implementation, the computer-executable instructions, when executed, may further cause the processor 910 to: construct a prompt based on a description of the first image and / or an image generation requirement. The generating a second image may comprise: decoding the first image representation according to the illumination information and the prompt, to generate a second image.
[0090] In an implementation, the reference video frame may include a foreground image. An illumination situation of the second image may be adapted to the foreground image. The illumination situation may include at least one of intensity, chromaticity, direction, and range of light.
[0091] In an implementation, the second image may include a light source. An attribute of the light source may be the same or similar to an attribute of a light source in the reference video frame.
[0092] It should be appreciated that the processor 910 may further perform any other step / process of the method for video communication according to the embodiments of the present disclosure as mentioned above.
[0093] The embodiments of the present disclosure propose a computer program product for video communication, comprising a computer program that is executed by a processor for: receiving a first image, the first image being an original background image; encoding the first image to generate a first image representation; identifying a reference video frame in a video stream of a user, the reference video frame being a first-extracted video frame from the video stream, or a video frame with a brightness difference from a previous video frame; calculating illumination information of an environment where the user is located using the reference video frame; and decoding the first image representation according to the illumination information, to generate a second image, the second image being a self-adaptive lighting background image. Furthermore, the computer program may further be performed for implementing any other steps / processes of the method for video communication according to the embodiments of the present disclosure as mentioned above.
[0094] The embodiments of the present disclosure may be embodied in a computer-readable mediums for video communication. The computer-readable medium may comprise instructions, that when executed, may cause a processor to: receive a first image, the first image being an original background image; encode the first image to generate a first image representation; identify a reference video frame in a video stream of a user, the reference video frame being a first- extracted video frame from the video stream, or a video frame with a brightness difference from a previous video frame; calculate illumination information of an environment where the user islocated using the reference video frame; and decode the first image representation according to the illumination information, to generate a second image, the second image being a self-adaptive lighting background image. Furthermore, the instructions, when executed, may further cause the processor to perform any other steps / processes of the method for video communication according to the embodiments of the present disclosure as mentioned above.
[0095] It should be appreciated that all the operations in the methods described above are merely exemplary, and the present disclosure is not limited to any operations in the methods or sequence orders of these operations, and should cover all other equivalents under the same or similar concepts. In addition, the articles “a” and "an" as used in this specification and the appended claims should generally be construed to mean “one” or “one or more” unless specified otherwise or clear from the context to be directed to a singular form.
[0096] It should also be appreciated that all the modules in the apparatuses described above may be implemented in various approaches. These modules may be implemented as hardware, software, or a combination thereof. Moreover, any of these modules may be further functionally divided into sub-modules or combined together.
[0097] Processors have been described in connection with various apparatuses and methods. These processors may be implemented using electronic hardware, computer software, or any combination thereof. Whether such processors are implemented as hardware or software will depend upon the particular application and overall design constraints imposed on the system. By way of example, a processor, any portion of a processor, or any combination of processors presented in the present disclosure may be implemented with a microprocessor, microcontroller, digital signal processor (DSP), a field-programmable gate array (FPGA), a programmable logic device (PLD), a state machine, gated logic, discrete hardware circuits, and other suitable processing components configured for performing the various functions described throughout the present disclosure. The functionality of a processor, any portion of a processor, or any combination of processors presented in the present disclosure may be implemented with software being executed by a microprocessor, microcontroller, DSP, or other suitable platform.
[0098] Software shall be construed broadly to mean instructions, instruction sets, code, code segments, program code, programs, subprograms, software modules, applications, software applications, software packages, routines, subroutines, objects, threads of execution, procedures, functions, etc. The software may reside on a computer-readable medium. A computer-readable medium may include, by way of example, memory such as a magnetic storage device (e.g., hard disk, floppy disk, magnetic strip), an optical disk, a smart card, a flash memory device, random access memory' (RAM), read only memory (ROM), programmable ROM (PROM), erasable PROM (EPROM), electrically erasable PROM (EEPROM), a register, or a removable disk.Although memory is shown separate from the processors in the various aspects presented throughout the present disclosure, the memory may be internal to the processors, e.g., cache or register.
[0099] The previous description is provided to enable any person skilled in the art to practice the various aspects described herein. Various modifications to these aspects will be readily apparent to those skilled in the art, and the generic principles defined herein may be applied to other aspects. Thus, the claims are not intended to be limited to the aspects shown herein. All structural and functional equivalents to the elements of the various aspects described throughout the present disclosure that are known or later come to be known to those of ordinary skilled in the art are expressly incorporated herein and intended to be encompassed by the claims.
Claims
CLAIMS1 . A method for video communication, comprising: receiving a first image, the first image being an original background image; encoding the first image to generate a first image representation; identifying a reference video frame in a video stream of a user, the reference video frame being a first-extracted video frame from the video stream, or a video frame with a brightness difference from a previous video frame; calculating illumination information of an environment where the user is located using the reference video frame; and decoding the first image representation according to the illumination information, to generate a second image, the second image being a self-adaptive lighting background image.
2. The method of claim 1, further comprising: determining that the first image is an image suitable for modifying illumination situations, and wherein the identifying a reference video frame in a video stream of a user comprises: in response to determining that the first image is an image suitable for modifying illumination situations, identifying the reference video frame.
3. The method of claim 1, wherein the identifying a reference video frame in a video stream of a user comprises: extracting a video frame from the video stream at a predetermined time interval; in response to the extracted video frame being a first-extracted video frame, taking the extracted video frame as the reference video frame; in response to the extracted video frame being a non-first-extracted video frame, detecting a brightness of the extracted video frame; and in response to a brightness difference betw een the brightness and a brightness of a previous reference video frame exceeding a predetermined threshold, taking the extracted video frame as the reference video frame.
4. The method of claim 1, wherein the calculating illumination information of an environment where the user is located comprises: calculating the illumination information based on the reference video frame through a control network containing a light extractor.
5. The method of claim 1, wherein the calculating illumination information of an environment where the user is located comprises: preprocessing the reference video frame, to obtain a preprocessed video frame, the preprocessing comprising:extracting a background image from the reference video frame, or extracting a background image from the reference video frame, and inpainting the extracted background image through image inpainting; and calculating the illumination information based on the preprocessed video frame through a control network containing a light extractor.
6. The method of claim 1, further comprising: constructing a prompt based on a description of the first image and / or an image generation requirement, and wherein the generating a second image comprises: decoding the first image representation according to the illumination information and the prompt, to generate a second image.
7. The method of claim 1, wherein the reference video frame includes a foreground image, and an illumination situation of the second image is adapted to the foreground image, the illumination situation including at least one of intensity, chromaticity’, direction, and range of light.
8. The method of claim 1, wherein the second image includes a light source, an attribute of the light being the same or similar to an attribute of a light source in the reference video frame.
9. The method of claim 1, further comprising: detecting a selection of the second image by the user; and in response to detecting the selection of the second image by the user, combining a foreground image from the video stream with a selected second image into an updated video frame.
10. An apparatus for video communication, comprising: a processor; and a memory storing computer-executable instructions that, when executed, cause the processor to: receive a first image, the first image being an original background image, encode the first image to generate a first image representation, identify a reference video frame in a video stream of a user, the reference video frame being a first-extracted video frame from the video stream, or a video frame with a brightness difference from a previous video frame, calculate illumination information of an environment where the user is located using the reference video frame, and decode the first image representation according to the illumination information, to generate a second image, the second image being a self-adaptive lighting background image.
11. The apparatus of claim 10, wherein the computer-executable instructions, when executed, further cause the processor to: determine that the first image is an image suitable for modifying illumination situations, andwherein the identifying a reference video frame in a video stream of a user comprises: in response to determining that the first image is an image suitable for modifying illumination situations, identify ing the reference video frame.
12. The apparatus of claim 10, wherein the calculating illumination information of an environment where the user is located comprises: preprocessing the reference video frame, to obtain a preprocessed video frame, the preprocessing comprising: extracting a background image from the reference video frame, or extracting a background image from the reference video frame, and inpainting the extracted background image through image inpainting; and calculating the illumination information based on the preprocessed video frame through a control network containing a light extractor.
13. The apparatus of claim 10, wherein the reference video frame includes a foreground image, and an illumination situation of the second image is adapted to the foreground image, the illumination situation including at least one of intensify, chromaticity, direction, and range of light.
14. The apparatus of claim 10, wherein the second image includes a light source, an attribute of the light being the same or similar to an attribute of a light source in the reference video frame.
15. A computer program product for video communication, comprising a computer program that is executed by a processor for: receiving a first image, the first image being an original background image; encoding the first image to generate a first image representation; identifying a reference video frame in a video stream of a user, the reference video frame being a first-extracted video frame from the video stream, or a video frame with a brightness difference from a previous video frame; calculating illumination information of an environment where the user is located using the reference video frame; and decoding the first image representation according to the illumination information, to generate a second image, the second image being a self-adaptive lighting background image.
Citation Information
Patent Citations
Image Processing Method, Image Processing Apparatus and Electronic Device
US20200226729A1
Matching foreground and virtual background during a video communication session
US20220070389A1
Method for image processing, computer device, and storage medium
US20220284638A1
Virtual Background Adjustment Based On Conference Participant Lighting Levels
US20230231972A1
Image rendering method and apparatus
WO2022095757A1