Intelligent visual adaptation for image and / or video applications to enhance user experience

WO2025260104A3PCT designated stage Publication Date: 2026-04-09FUTUREWEI TECHNOLOGIES INC
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-08-27
Publication Date
2026-04-09

AI Technical Summary

Technical Problem

Existing image and video applications lack joint environment, user preference, and quality-aware visual adaptation, leading to inconsistent and visually unpleasing experiences, particularly in real-time virtual events.

Method used

Implementing intelligent visual adaptation using generative AI models for joint in-out-painting, view angle adaptation, and quality-aware enhancements, considering environmental conditions, user preferences, and visual quality assessment, with processing distributed across sender, cloud, and receiver devices.

Benefits of technology

Enhances user experience by providing aesthetically appealing and immersive visuals in real-time virtual events, ensuring visual coherence and adaptability across heterogeneous environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025043738_09042026_PF_FP_ABST
    Figure US2025043738_09042026_PF_FP_ABST
Patent Text Reader

Abstract

A computer-implemented method including receiving an input image; adapting, based on one or more visual adaptation criteria, one or more visual properties of the input image to generate a synthesized image, wherein the adapting includes generating one or more masked image patches, each comprising a respective masked portion of the input image; processing the one or more masked image patches using one or more computer vision models to generate respective ones of one or more output image patches; and generating the synthesized image based on the input image and the one or more output image patches, the synthesized image comprising at least a first portion extending a view of the input image and a second portion modifying content within the input image; and rendering, on a display device, the synthesized image.
Need to check novelty before this filing date? Find Prior Art

Description

Intelligent Visual Adaptation for Image and / or Video Applications to Enhance User ExperienceTECHNICAL FIELD

[0001] The present disclosure is generally related to image and / or video processing, and in particular, to intelligent visual adaptation for image and / or video applications to enhance user experience.BACKGROUND

[0002] Image and video applications are widely used in industries and daily life, changing how humans use technologies and interact. For instance, video conferencing allows real-time communication between participants in different locations to see, hear, and interact with each other using audio and video transmission over a network (e.g., the Internet). Video conferencing is widely used for various purposes, such as business meetings, remote work collaboration, online education, virtual social gatherings, etc. Additionally, recent advancements in computer video technologies have enabled many more virtual reality (VR) and / or augmented reality (AR) applications (e.g., for gaming, entertainment, medical procedures, etc.). Furthermore, the proliferation of mobile devices has led to the widespread and frequent use of mobile imaging and / or video applications (e.g., social media, image and / or video streaming, image and / or video editing, etc.).SUMMARY

[0003] This disclosure is related to intelligent visual adaptation for video conferencing, real-time virtual gathering events, and / or other image and / or video applications. The intelligent visual adaptation may include joint in-out-painting, view angle or perspective adaptation, quality-aware and / or content-aware enhancements (e.g., lighting, colors, brightness, contrast, etc.). In-painting refers to modifying (e.g., adding, removing, or replacing) a portion of the content within an image. Out-painting refers to extending beyond the view of an image. Joint in-out-painting can improve the overall user experience (e.g., providing more pleasant, realistic, aesthetically appealing, and / or immersive experience) for virtual real-time events. The visual adaptation may take into account environmental condition sensing information, visual quality assessment results, and user preferences. The visual adaptation may be based on image patch-based processing using various generative artificial intelligence (Al) models, followed by refinement, visual composition, andblending to ensure visual coherence across the visually adapted image. The intelligent visual adaptation can improve visual perception and enhance user experience for image and / or video applications.

[0004] A first aspect of the embodiments of the present disclosure relates to a computer- implemented method comprising receiving an input image; adapting, based on one or more visual adaptation criteria, one or more visual properties of the input image to generate a synthesized image, wherein the adapting comprises generating one or more masked image patches, each comprising a respective masked portion of the input image; processing the one or more masked image patches using one or more computer vision models to generate respective ones of one or more output image patches; and generating the synthesized image based on the input image and the one or more output image patches, the synthesized image comprising at least a first portion extending a view of the input image and a second portion modifying content within the input image; and rendering, on a display device, the synthesized image.

[0005] Optionally, in any of the preceding aspects, another implementation of the aspect provides that the one or more visual adaptation criteria comprises an indication of at least one of a background extension, a content extension size, a scene adaptation, an object adaptation, or a viewpoint adaptation.

[0006] Optionally, in any of the preceding aspects, another implementation of the aspect provides that the one or more visual adaptation criteria comprises a sample image.

[0007] Optionally, in any of the preceding aspects, another implementation of the aspect provides that the one or more visual adaptation criteria are based on at least one of a text prompt or a user preference.

[0008] Optionally, in any of the preceding aspects, another implementation of the aspect provides that the generating the one or more masked image patches comprises generating, based on the one or more visual adaptation criteria, one or more mask patches, each masking a respective portion of the input image for visual adaptation; and generating the one or more masked image patches based on respective ones of the one or more mask patches and the input image.

[0009] Optionally, in any of the preceding aspects, another implementation of the aspect provides that the one or more computer vision models used for generating the one or more output image patches comprises one or more generative artificial intelligence (Al) models.

[0010] Optionally, in any of the preceding aspects, another implementation of the aspect provides that the one or more computer vision models used for generating the one or more output image patches comprises a first generative artificial intelligence (Al) model trained for in-painting and a second generative Al model trained for out-painting.

[0011] Optionally, in any of the preceding aspects, another implementation of the aspect provides that the processing the one or more masked image patches comprises processing an individual one of the one or more masked image patches using one of the first generative Al model or the second generative Al model to generate an intermediate image patch; and processing the intermediate image patch using the other one of the first generative Al model or the second generative Al model to generate a corresponding one of the one or more output image patches.

[0012] Optionally, in any of the preceding aspects, another implementation of the aspect provides that the one or more computer vision models used for generating the one or more output image patches comprises a generative artificial intelligence (Al) model trained for joint in-out- painting.

[0013] Optionally, in any of the preceding aspects, another implementation of the aspect provides that the generating the synthesized image comprises refining at least one of a local coherence, a global coherence, or content detail of the one or more output image patches using a refinement network; and generating the synthesized image based on the refined one or more output image patches, the input image, and patch layout information associated with the one or more output image patches using a blending and composition network.

[0014] Optionally, in any of the preceding aspects, another implementation of the aspect provides that the processing the one or more masked image patches comprises processing a first masked image patch using a first computer vision model of the one or more computer vision models before processing a second masked image patch using a second computer vision model of the one or more computer vision models based on the first masked image patch being completely within the input image and the second masked image patch having a portion within the input image and a portion outside of the input image.

[0015] Optionally, in any of the preceding aspects, another implementation of the aspect provides that the input image comprises a plurality of images of participants of a real-time virtual event.

[0016] Optionally, in any of the preceding aspects, another implementation of the aspect provides that the adapting the one or more visual properties of the input image is further based on a visual quality assessment across the plurality of images.

[0017] Optionally, in any of the preceding aspects, another implementation of the aspect provides that the adapting the one or more visual properties of the input image is further based on information obtained from environmental condition sensing associated with capturing of one or more of the plurality of images.

[0018] Optionally, in any of the preceding aspects, another implementation of the aspect provides that the adapting comprises adapting at least one of a background, a viewpoint, a subject framing style, or a subject appearance associated with an individual participant of the plurality of participants.

[0019] Optionally, in any of the preceding aspects, another implementation of the aspect provides that the synthesized image comprises an image of two or more of the participants within the same virtual environment.

[0020] Optionally, in any of the preceding aspects, another implementation of the aspect provides that the adapting the one or more visual properties of the input image comprises offloading one or more operations associated with the adapting to a network server associated with the realtime virtual event.

[0021] Optionally, in any of the preceding aspects, another implementation of the aspect provides that the offloading the one or more operations associated with the adapting to the network server is based on at least one of a resource assessment or a real-time requirement.

[0022] Optionally, in any of the preceding aspects, another implementation of the aspect provides that one or more operations associated with the adapting the one or more visual properties of the input image is performed at a computing device associated with a receiver of the input image.

[0023] A second aspect of the embodiments of the present disclosure relates to a computing device comprising a memory storing instructions; and one or more processors coupled to the memory, the one or more processors configured to execute the instructions to cause the computing device to implement the method of any of the disclosed embodiments.

[0024] A third aspect of the embodiments of the present disclosure relates to a non-transitory computer readable medium comprising a computer program product for use by a computing device, the computer program product comprising computer executable instructions stored on the non-transitory computer readable medium that, when executed by one or more processors, cause the computing device to execute the method of any of the disclosed embodiments.

[0025] A fourth aspect of the embodiments of the present disclosure relates to a computing device comprising means for receiving an input image; means for adapting, based on one or more visual adaptation criteria, one or more visual properties of the input image to generate a synthesized image, wherein the adapting comprises generating one or more masked image patches, each comprising a respective masked portion of the input image; processing the one or more masked image patches using one or more computer vision models to generate respective ones of one or more output image patches; and generating the synthesized image based on the input image and the one or more output image patches, the synthesized image comprising at least a first portion extending a view of the input image and a second portion modifying content within the input image; and means for rendering, on a display device, the synthesized image.

[0026] For the purpose of clarity, any one of the foregoing embodiments may be combined with any one or more of the other foregoing embodiments to create a new embodiment within the scope of the present disclosure.

[0027] These and other features will be more clearly understood from the following detailed description taken in conjunction with the accompanying drawings and claimsBRIEF DESCRIPTION OF THE DRAWINGS

[0028] For a more complete understanding of this disclosure, reference is now made to the following brief description, taken in connection with the accompanying drawings and detailed description, wherein like reference numerals represent like parts.

[0029] FIG. 1 is a block diagram illustrating an example video conferencing system that implements environment, user preference, content, and / or quality aware visual adaptation according to an embodiment of the present disclosure.

[0030] FIG. 2 is a block diagram illustrating an example video conferencing scenario with environment, user preference, content, and / or quality aware visual adaptation according to an embodiment of the present disclosure.

[0031] FIG. 3 is a block diagram illustrating another video conferencing scenario with example environment, user preference, content, and / or quality aware visual adaptation according to an embodiment of the present disclosure.

[0032] FIG. 4 is a block diagram illustrating an example joint in-out-painting method according to an embodiment of the present disclosure.

[0033] FIGS. 5A-5C are block diagrams illustrating an example implementation of the joint in- out-painting method of FIG. 4 according to an embodiment of the present disclosure.

[0034] FIG. 6 is a block diagram illustrating another example joint in-out-painting method according to an embodiment of the present disclosure.

[0035] FIG. 7 is a block diagram illustrating an example quality aware visual adaptation method according to an embodiment of the present disclosure.

[0036] FIG. 8 is a block diagram illustrating an example video conferencing system architecture according to an embodiment of the present disclosure.

[0037] FIG. 9 is a block diagram illustrating another video conferencing system architecture according to an embodiment of the present disclosure.

[0038] FIG. 10 is a flowchart of a computer-implemented method according to an embodiment of the present disclosure.

[0039] FIG. 11 is a schematic diagram of a network apparatus according to an embodiment of the disclosure.DETAILED DESCRIPTION

[0040] It should be understood at the outset that although illustrative implementations of one or more embodiments are provided below, the disclosed systems and / or methods may be implemented using any number of techniques, whether currently known or in existence. The disclosure should in no way be limited to the illustrative implementations, drawings, and techniques illustrated below, including the exemplary designs and implementations illustrated and described herein, but may be modified within the scope of the appended claims along with their full scope of equivalents.

[0041] A video conferencing system (e.g., a computer device or system) may include a webcam, a microphone, and a speaker. The webcam may transmit live video, allowing participants to see each other. For instance, each participant may see a series of image frames, each including a view or image of each participant of the video conference on the same screen. The microphone and speaker may enable participants to hear and speak to each other in real-time, allowing for interactive conversations. To facilitate collaboration, some video conferencing platforms may also provide screen sharing, whiteboard sharing, etc. That is, a participant of the video conference may receive a series of video or image frames, each including a whiteboard or the screen that is being shared inaddition to images of the participants. These definitions should be considered as a supplement and should not be considered to limit any other definitions of descriptions provided for such terms herein.

[0042] As participants are at different locations with different environmental conditions (e.g., lighting, etc.) and may use different devices with different image sensing and / or processing capabilities, the image quality of the onscreen participant view for each individual may be different. One approach is to apply the same visual adjustment across the entire screen (e.g., the image frame including the images of all the participants and / or the shared whiteboard or screen). However, this may not result in a visually pleasing image. For instance, if the image of one of the participants is captured under a low light condition, simply adjusting the brightness at a full-screen level (e.g., across the entire image frame) may cause other parts of the screen (e.g., the images of other participants and / or the shared whiteboard or screen) to be overly bright. Additionally, different participants may have different preferences (e.g., in terms of viewing options and / or the appearances of the participants). As an example, a participant attending a virtual business meeting may desire to appear in business attire but without having to put on business attire. As another example, a participant of a virtual meeting may desire to appear facing the camera instead of looking away from the camera (e.g., due to the participant’s workspace set up) to create engagement with the meeting. Generally, there is a lack of joint environment, user preference, quality and content aware visual adaptation available for image and / or video applications. Furthermore, as real-time virtual events or gatherings become frequently used in today’s interconnected world, user experience is becoming an important factor for image and / or video applications.

[0043] Disclosed herein are techniques for providing intelligent visual adaptation for video conferencing, real-time virtual gathering events, and / or other image and / or video applications. More specifically, the intelligent visual adaptation may include joint in-out-painting, view angle or perspective adaptation, quality- aware and / or content-aware enhancements (e.g., lighting, colors, brightness, contrast, etc.). In-painting refers to modifying (e.g., adding, removing, or replacing) a portion of the content within an image. Out-painting refers to extending beyond the view of an image. Joint in-out-painting can improve the overall user experience (e.g., providing more pleasant, realistic, aesthetically appealing, and / or immersive experience) for virtual real-time events. The visual adaptation may take into account environmental condition sensing information, visual quality assessment results, and user preferences. In an embodiment, the visual adaptation may be based on an image patch-based processing approach using various generative artificial intelligence (Al)models, followed by refinement, visual composition, and blending to ensure visual coherence across the visually adapted image. Furthermore, the intelligent visual adaptation may be implemented in a heterogenous environment (e.g., across a sender device, a cloud device, an edge device, and / or a receiver device) based on an assessment of application requirements (e.g., real-time requirements) and / or computing and / or memory resource availability. The intelligent visual adaptation can improve visual perception and enhance user experience for image and / or video applications. While the disclosed embodiments are discussed in the context of video conferencing, the disclosed embodiments are applicable for other image and / or video related applications (e.g., image and / or video editing, augmented reality (AR) and / virtual reality (VR) applications, mobile imaging, photography, video processing and / or streaming, computer vision, medical imaging, etc.).

[0044] FIG. 1 is a block diagram illustrating an example video conferencing system 100 that implements environment, user preference, content, and / or quality aware visual adaptation according to an embodiment of the present disclosure. In the system 100, participants 102, 104, and 106 may participate in a video conference (e.g., a virtual meeting) via respective computing devices 120, 136, and 138. The participants 102, 104, and 106 may be located at different geographical locations and may communicate with each other in real-time over a network 110. The network 110 may include public network(s), private network(s), or a combination thereof. The network 110 may include wireless networks, wireline networks, or a combination thereof. In an example, the network 110 may include the Internet.

[0045] Each of the computing devices 120, 136, and 138 may be equipped with a camera (e.g., a webcam), a microphone, and a speaker, allowing the participants 102, 104, and 106 to see, hear, and speak to each other. Generally, the computing devices 120, 136, and 138 may be mobile phones, desktop computers, notebook computers, tablets, or any suitable devices that support video conferencing. For simplicity, FIG. 1 only shows an expanded view of the computing device 120 of the participant 102. However, the computing devices 136 and 138 may include substantially similar components as computing device 120. As shown in FIG. 1, the computing device 120 may include a video conferencing application 122, an environmental condition sensing engine 124, a quality measurement engine 126, a visual adaptation engine 128, a display device 130, an application and device resource adaptation component 131, one or more computer vision models 132, a user preference configuration 134, and user interface (UI) 135.

[0046] The video conferencing application 122 may include program instructions (e.g., software) stored in memory of the computing device 120 and executable by one or more processors of the computing device 120. The video conferencing application 122 may establish connections with the other participants 104 and 106 (or more specifically, with the video conferencing applications executing on the respective computing devices 136 and 138) for real-time audio and video exchange. In an example, a video conferencing server 112 (e.g., a network device) may be located in the network 110 and may facilitate the connection establishment. After establishing the connections, the video conferencing application 122 may receive audio and video of the participant 102 captured respectively by the camera and the speaker of the computing device 120. The video conferencing application 122 may transmit the audio and video of the participant 102 to the other participants 104 and 106 via the network 110. The video conferencing application 122 may also receive audio and video of each of the participants 104 and 106 via the network 110.

[0047] The video conferencing application 122 may display a combination or mix of the videos of the participants 102, 104, and 106 on the display device 130 (e.g., computer monitors, smartphone screens, etc.). A video may generally include a series of time-stamped image frames. In one example, each of the participants 102, 104, and 106 may transmit their respective videos and audios to the video conferencing server 112. The video conferencing server 112 may combine or mix the videos of the participants 102, 104, and 106 to generate a combined video. The video conferencing server 112 may also combine or mix the audios of the participants 102, 104, and 106 to generate a combined audio. The video conferencing server 112 may forward the combined video and the combined audio to each of the participants 102, 104, and 106. In another example, each of the participants 102, 104, and 106 may transmit their respective videos and audios to the video conferencing server 112, and the video conferencing server 112 may forward, to each one of the participants 102, 104, and 106, the videos and audios of the other ones of the participants 102, 104, 106, where the combining of the videos and / or audios may be performed at each receiving participant computing devices 120, 136, and 138. As an example, the video conferencing application 122 may receive videos and audios of the respective participants 104 and 106 separately. The video conferencing application 122 may combine the locally captured video of the participant 102 with the videos of the participants 104 and 106 received from the network 110. The video conferencing application 122 may also combine the locally captured audio of the participant 102 with the audios of the participants 104 and 106 received from the network 110. Generally, the combining or mixing of the videos and / or audios of theparticipants 102, 104, 106 can be performed at the video conferencing server 112 and / or at a receiving participant device (e.g., the computing devices 120, 136, and 138). In some examples, the video conferencing application 122 may also provide other functionalities, such as screen sharing, whiteboard sharing, etc. In such examples, the shared screen or shared whiteboard may be combined with the videos of the participants 102, 104, and 106 for display.

[0048] The top left portion of FIG. 1 illustrates an example of visual inputs of the respective participants 102, 104, and 106 at the sender side. For instance, the computing device 120 may capture an image 103 of the participant 102 and transmitted the image 103 over the network 110 to the other participants 104 and 106. Similarly, the computing device 136 may capture an image 105 of the participant 104 and transmit the image 105 over the network 110 to the other participants 102 and 106. The computing device 138 may capture an image 107 of the participant 106 and transmit the image 107 over the network 110 to the other participants 102 and 104. At the receiver end, the video conferencing application 122 may render a series of image frames (e.g., the combined video) on the display device 130. Each image frame may include an image (or view) of each of the participants 102, 104, and 106 and / or a shared whiteboard or a shared screen (if any). The top right portion of FIG. 1 illustrates an example receiver end display (e.g., on the display device 130). As shown, the receiver end display may include an image frame 140 including an image 142 corresponding to the image 103 of the participant 102, an image 144 corresponding to the image 105 of the participant 104, an image 146 corresponding to the image 107 of the participant 106, and a shared screen 148 (e.g., shared by one of the participants 102, 104, or 106).

[0049] According to embodiments of the present disclosure, the video conferencing application 122 may coordinate with the environmental sensing engine 124, the quality measurement engine 126, and the visual adaptation engine 128 to provide intelligent visual adaptation for the video conference (e.g., a series of image frames 140). In an embodiment, the environmental sensing engine 124 may include software and / or hardware components configured to sense the environmental condition (e.g., lighting, workspace set up, etc.) surrounding the participant 102. In an example, the environmental sensing engine 124 may include visual sensor(s), camera(s), etc. The environmental sensing engine 124 may provide the environmental sensing information to the quality measurement engine 126. The quality measurement engine 126 may include software and / or hardware components configured to determine whether the sensing information (e.g., lighting) satisfies certain criteria. In some examples, the quality measurement engine 126 may also determine whether the quality of acaptured image 103 (e.g., resolution, brightness, contrast, color, etc.) satisfies certain criteria. The quality measurement engine 126 may provide the quality measurement data to the visual adaptation engine 128. The visual adaptation engine 128 may include hardware and / or software components configured to adapt the visual properties of the image frame(s) 140 associated with the video conference based on the quality measurement data and / or various other visual adaptation criteria as will be discussed more fully below. The visual adaptation engine 128 may provide the visually adapted image frames 140 for display on the display device 130.

[0050] Because the participant 102 may act as a sender and a receiver in the video conference, some of the operations or components of the computing device 120 may be associated with sender end processing while other operations or components of the computing device 120 may be associated with receiver end processing. For instance, the environmental sensing engine 124 and the quality measurement engine 126 may perform operations associated with sender end processing (e.g., sensing the lighting condition and measuring the brightness of a visual input at the sender side) while the visual adaptation engine 128 may perform operations associated with receiver end processing (e.g., to adapt and enhance the visual input at the receiver end for display) and / or sender end processing (e.g., to adapt and enhance a captured image of the participant 102 before transmitting over to the network 110). As will be discussed more fully below with reference to FIGS. 8 and 9, in some instances, the visual adaptation operations may be implemented in a heterogenous environment, distributing between a sender device (e.g., the computing device 120 at a transmitting end), a receiving device (e g., the computing device 120 at a receiving end), and / or a network device (e.g., the video conferencing server 112).

[0051] In an embodiment, the visual adaptation engine 128 may perform various visual adaptations or enhancements, for example, including, but not limited to, human perception adaptive quality-aware adaptation, user preference-aware adaptation, environment sensing-aware adaptation, viewpoint adaptation, out-painting, in-painting, and / or joint in-out-painting for enhanced visual experience, flexible participant appearance adaptation, flexible environmental condition adaptation, intelligent static and dynamic content adaptation. The human perception adaptive quality-aware adaptation may include adaptations of hue, saturation, contrast, and / or brightness of individual pixels or regions in an image (e.g., the image frame 140 and / or the images 103, 105, and / or 107) to provide a visually pleasing output image.

[0052] The user preference-aware adaptation may be based on the user preference configuration 134. The user preference configuration 134 may be input by a user (e.g., the participant 102) and stored in the memory of the computing device 120. For example, the user preference configuration 134 may be part of a user profile associated with the video conferencing application 122. The user preference configuration 134 may include various settings, for example, including, but not limited to, image properties (e.g., hue, saturation, contrast, and / or brightness), participant appearance (e.g., clothing, hairstyle, makeup, accessories, etc.), background of the participant (e.g., a virtual environment, such as an office, virtual paintings on a wall, virtual plants, etc.), dimension for out- painting (e.g., dimension of an extended view), and / or view angle (e.g., for view point adaptation). The viewpoint adaptation may include adjusting the view angle or perspective of the participant 102. For instance, the participant 102 may be facing the side of the camera, and the viewpoint adaptation may adjust an image of the participant 102 so that the participant 102 may appear to face the direction of the camera. Additionally or alternatively, the user preference-aware adaptation may be based on instructions input by a user and received via the UI 135. The instructions may be in text or any other suitable forms, such as audio or video. In an example, a text prompt received from the UI 135 may include instruction on how the user would like to adjust the appearance of an individual participant 102, 104, and / or 106 and / or the overall application display window (e.g., the receiver end display).

[0053] The in-painting adaptation may include modifying content within an original image (e.g., the images 103, 105, and / or 107). For instance, in-painting may generally be used for restoring or modifying regions of an image or a video, such as filling in missing details, removing unwanted objects, and / or changing existing elements. The out-painting adaptation may include extending the view of an image (e.g., the images 103, 105, and / or 107) beyond the original view of the image. For instance, out-painting may extend the canvas of an image or a video frame by generating new content beyond its original boundaries, extrapolates what might be beyond the original view of the content, adding new visual elements, and seamlessly blending them with the existing content. The joint in- out-painting combines the processing of both in-painting and out-painting to leverage the strengths and capabilities of both modalities to achieve better results (e.g., more pleasant, realistic, aesthetically appealing, and / or immersive experience) for image and / or video applications. Various examples of joint in-out-paining are shown in FIGS. 1-3 as will be discussed below.

[0054] The environmental condition adaptation may include adjusting one or more visual properties of an image 103 of the participant based on surroundings (e.g., lighting or workspace setup) of the participant 102. The intelligent static and dynamic adaptation may include adapting static properties and dynamic properties of image frames 140. For example, the static visual adaptation may include adjusting the image size based on the screen size of the display device 130, the resolution based on the display device 130, the layout and / or the virtual background (e.g., based on the user preference configuration 134). The dynamic visual adaptation may include dynamically adjusting the resolution based on changes in network conditions, device processing loadings, and / or memory availability, correcting color and / or lighting based on changes in the surroundings of the participant 102, adjusting appearance of the participant 102 based on movements of the participant 102 (e.g., hand gestures, hair movements, facial movements, etc.).

[0055] In an embodiment, the visual adaptation engine 128 may adapt the visual properties of an image frame (e.g., the image frame 140 and / or individual participant images 103, 105, and / or 107) based on one or more visual adaptation criteria. In an embodiment, the one or more visual adaptation criteria may include, for example, but are not limited to, an indication of at least one of a background extension, a content extension size, a scene adaptation, an object adaptation, or a viewpoint adaptation. A background extension may refer to extending a background view of a participant 102 beyond the original background view (e.g., similar to a zoom out operation). For instance, the original background view of the participant 102 may only show a portion of a room and a portion of a painting on a wall, and the background extension may extend the background view to show an entire room with a full view of the painting on the wall. A content extension size may be in terms of dimensions or a particular view, such as a half body view or a full body view (e.g., for an original headshot view). An example of a scene adaptation may include adapting a virtual office environment by adding a desk, painting(s) on walls, and / or plant(s) behind a participant 102. In some instances, the one or more visual adaptation criteria may specify certain portion(s) of a scene (e.g., top, bottom, left, and / or right portion(s), etc.) for scene adaptation. An example of an object adaptation may include modifying the appearance of a participant 102 (e.g., with a different hair style, a different attire, etc.). In some instances, the one or more visual adaptation criteria may indicate certain portion(s) and / or certain feature(s) of an object (e.g., top, bottom, left, and / or right portion(s), or hairstyle, attire in the case of an individual) for object adaptation. An example of a viewpoint adaptation may include modifying a participant view from a side profile view to a frontal view. In some instances, the one or more visual adaptation criteria may specify a certain perspective angle or a range of perspective angles for viewpoint adaptation.

[0056] In an embodiment, the one or more visual adaptation criteria may include sample image(s) (e.g., indicating an example of how a participant image may be presented in terms of appearance or a background of the participant image, etc.). In an embodiment, the indication for the one or more visual adaptation criteria may be in the form of a text prompt(s) (e.g., via the UI 135) and / or user preference(s) stored in the user preference configuration 134. Additionally or alternatively, the visual adaptation engine 128 may adapt the visual properties of an image frame (e.g., the image frame 140 and / or individual participant images 103, 105, and / or 107) based on environmental sensing associated with capturing of the image 103, 105, and / or 107. Additionally or alternatively, the visual adaptation engine 128 may adapt the visual properties of an image frame (e.g., the image frame 140) based on a visual quality assessment across images (e.g., the individual images 142, 144, and 146 and the shared screen 148) in the image frame.

[0057] In the illustrated example of FIG. 1, the visual adaptation engine 128 may adjust the brightness of the image 103 of the participant 102 and perform joint in-out painting to provide the image 142 for the receiver end display. For example, the environmental sensing engine 124 may sense a light condition surrounding the participant 102 (e.g., as shown by the image 103). The quality measurement engine 126 may determine that the sensing information associated with the light condition fails to satisfy a certain threshold (e.g., a brightness threshold). In an embodiment, the visual adaptation engine 128 may increase the brightness of the participant 102’s video or images (e.g., the image 103) before transmitting the video or images to the other participants 104 and 106. The visual adaptation engine 128 may also increase the brightness of the participant 102’s video or images (e.g., the image 103) for local display on the display device 130 (e.g., as shown by the image 142 at the receiver end display). Depending on the video mixing mode, the visual adaptation engine 128 may increase the brightness of the image 103 before mixing or after mixing. For example, when the video mixing is performed at the receiver end, the visual adaptation engine 128 may adjust the brightness of the image 103 before combining the image 103 with the images 105 and 107 of the other participants 104 and 106 received from the network 110. Alternatively, when the video mixing is performed at the video conferencing server 112, the visual adaptation engine 128 may receive a combined image frame 140 and may adjust the brightness of the portion of the combined image frame 140 corresponding to the image 103 (e.g., as shown by the image 142) before displaying the image frame 140 on the display device 130. In some embodiments, a transmitting participant device (e.g., the computing devices 120, 136, and / or 138) may not have visual adaptation capabilities. Insuch embodiments, instead of adjusting the visual properties of a locally captured image (e.g., the images 103, 105, and / or 107) before transmitting the image over to the network 110, the transmitting participant device may transmit the quality measurement information along with the image to a receiving participant’s device (e.g., the computing devices 120, 136, and / or 138).

[0058] Additionally, the visual adaptation engine 128 may perform joint in-out painting (e.g., based on the user preference configuration 134) to provide flexible background and user appearance adaptation. For example, the participant 102 may configure an appearance to be in business attire with a half body view and at a desk in an office environment. Generally, the user preference configuration 134 may include an indication of at least one of a background extension (e.g., a virtual office environment), a content extension size (e.g., in terms of dimensions or a particular view, such as a half body view), a scene adaptation (e.g., a virtual office environment with a desk, painting(s) on a wall, and / or plant(s) behind the participant 102), or an object adaptation (e.g., appearance adaptation). In some instances, the user preference configuration 134 may also include an example image (e.g., an example background image with a desk in an office environment, an example image of a person in a half body view, an example image of business attire, etc.). In some examples, the video conference application 122 may provide a template of images of different backgrounds, different clothing styles, different poses, etc. and the user preference configuration 134 may select one or more of those images from the template as examples for the visual adaptation.

[0059] As can be seen from the image 103 (at the sender end), the actual participant 102 is in casual attire with her hair tied up and in front of a blank wall while attending the video conference and the image 103 may only capture a headshot of the participant 102. As shown by the onscreen view or image 142 of the participant 102 (on the display device 130), the participant 102 appears in business attire with her hair down and at a desk in an office environment. For instance, the visual adaptation engine 128 may perform in-painting to modify the hairstyle and the attire of the participant 102, and out-painting to extend the view of the captured image 103 of the participant 102 (the headshot view shown on the left) to provide a half body view behind a desk. Stated differently, the visual adaptation engine 128 may generate a synthesized image 142 of the participant 102 based on a received captured image 103 of the participant 102, where the synthesized image 142 may include at least a first portion extending a view of the received image 103 and a second portion modifying the content within the received image 103. The visual adaptation engine 128 may performin-painting and out-painting jointly to satisfy the various visual adaptation criteria and provide a visually cohesive image (e.g., with aesthetically harmonized visual appearances).

[0060] In some embodiments, the visual adaptation engine 128 may also perform joint in-out- painting on the image 105 to provide the image 144. In some embodiments, the visual adaptation engine 128 may also perform joint in-out-painting on the image 107 to provide the image 146. As shown in FIG. 1, the images 105 and 107 (captured at the sender sides) are headshots of respective participants 104 and 106, and the adapted images 144 and 146 include half body views of the respective participants 104 and 106. Generally, the visual adaptation engine 128 may adapt a subject framing style (e.g., a headshot) in an image 103, 105, and / or 107 to any suitable framing style (e.g., a half-body view, a full-body view, etc.).

[0061] In an embodiment, the visual adaptation engine 128 may utilize one or more computer vision models 132 to perform joint in-out-painting and may consider global semantic or contextual information and low-level pixel or spatial information jointly to provide joint in-out-painting as will be discussed more fully below with reference to FIGS. 4, 5A-5C, and 6. The one or more computer vision models 132 may include program instructions (e.g., software) stored in the memory of the computing device 120 and executable by the one or more processors of the computing device 120 and / or model parameters and related data stored in the memory of the computing device 120. In an embodiment, the one or more computer vision models 132 may include generative Al models trained to perform in-painting, out-painting, joint in-out-painting, refinement (e.g., local coherence, consistency across in-painted portion(s), out-painted portion(s), original portion(s), etc ), visual composition, and / or blending as will be discussed more fully below with reference to FIGS. 4, 5A- 5C, and 6-7. In general, the one or more computer vision models 132 may have any suitable model architectures (e.g., convolutional neural networks (CNNs), vision transformers (ViTs), generative Al vision models, recurrent neural networks (RNNs), attention models, u-net models, encode- decoder networks, diffusion models, etc.).

[0062] In some embodiments, because different participant devices (e.g., the computing devices 120, 136, and 138) may have different resources (e.g., computing, memory, display, etc.), the application and device resource adaptation component 131 may determine parameters for adapting the visual adaptation algorithms used by the visual adaptation engine 128 to the device capabilities. For instance, if the computing device 120 has high processing power but small memory, the application and device resource adaptation component 131 may adapt a visual adaptation algorithmto use a smaller amount of memory. Generally, the application and device resource adaptation component 131 may adapt visual adaptation algorithm(s) and / or distribute or offload some computations of visual adaptation to other device(s) (e.g., to a network device as will be discussed more fully below with reference to FIGS. 8 and 9).

[0063] FIG. 1 is merely an example of components of a video conferencing system 100, and variations are contemplated to be within the scope of the present disclosure. In some embodiments, the video conferencing system may include other components not illustrated in FIG. 1. In some embodiments, the video conferencing system may not include every component illustrated in FIG. 1. In some embodiments, the components and connections may be implemented with different connections than those illustrated in FIG. 1. Such and other embodiments are contemplated to be within the scope of the present disclosure.

[0064] FIGS. 2 and 3 illustrate additional examples of environment, user preference, content, and / or quality aware visual adaptation that may be implemented by the video conferencing system 100. FIG. 2 is a block diagram illustrating an example environment, user preference, and / or quality aware visual adaptation scenario 200 according to an embodiment of the present disclosure. For simplicity, FIG. 2 may use the same reference numerals as in FIG. 1 to refer to the same elements and / or components. The scenario 200 may be substantially similar to the visual adaptation example shown in FIG. 1 but may illustrate a visual adaptation including viewpoint or view angle adaptation and quality enhancement in addition to brightness adjustment and joint in-out-painting. As shown in FIG. 2, the participant 102 may be at a location with a low light condition and may be facing a different direction than the camera, and thus an image 203 of the participant 102 (captured at the sender side) shows a side profile view of a headshot of the participant 102. In the scenario 200, the visual adaptation engine 128 at the computing device 120 of the participant 102 may adapt a view angle of the participant 102 in addition to joint in-out-painting. As shown in FIG. 2, an image frame 240 displayed on the display device 130 includes an image 242 of the participant 102 in a frontal view and with the headshot extended to a half-body view behind a desk in an office environment. In an example, the view angle adaptation may be configured at the user preference configuration 134 and / or a user text prompt input. In other examples, the view angle adaptation may be based on sensing and / or image analysis. Additionally, the visual adaptation engine 128 may adjust the brightness of the image 203 without causing other portions of the image frame 240 to be overly bright. As shown, the image 242 appears to be brighter than the original input image 203.

[0065] FIG. 3 is a block diagram illustrating another example with environment, user preference, content, and / or quality aware visual adaptation scenario 300 according to an embodiment of the present disclosure. For simplicity, FIG. 3 may use the same reference numerals as in FIG. 1 to refer to the same elements and / or components. The scenario 300 may be substantially similar to the video conferencing scenario shown in FIG. 1 but may further illustrate the application of joint in-out- painting to provide immersive video conferencing experience. For simplicity, FIG. 3 only includes the participants 102 and 104 in the video conference. To provide immersive experience, the visual adaptation engine 128 at the computing device 120 of the participant 102 may generate an image including both the participants 102 and 104 in the same virtual environment (e.g., seated seamlessly in the same scene) with joint in-out-painting and scene composition and / or fusion as shown by the image frame 340 displayed on the display device 130. In an example, the immersive setting may be part of the user preference configuration 134 and / or a user text prompt input.

[0066] FIG. 4 is a block diagram illustrating an example joint in-out-painting method 400 according to an embodiment of the present disclosure. In an embodiment, the method 400 may be performed by a visual adaptation engine (e.g., the visual adaptation engine 128) to enhance user visual experience. The method 400 may use a patch-based approach for joint in-out-painting. The method 400 may perform joint in-out-paining 404 using an image mask generation component 410, an image patch generation component 420, an output refinement component 430, and a visual composition and blending component 440.

[0067] As shown in FIG. 4, an image mask generation component 410 may receive an input image 402, represented by X. In an example, the input image 402 may be an image of a conference participant (e.g., similar to the images 103, 105, 107, 203, and / or 303). The image mask generation component 410 may generate an image mask 412. The image masks 412 may be represented by M. The image mask 412 may mask a portion of the input image 402. For example, the image mask 412 may include a value of 1 for each pixel of the input image 402 in the region of interest (e.g., to be adapted for in-painting and / or out-painting) and a value of 0 for each pixel outside the region of interest.

[0068] To perform joint in-out-painting using a patch-based processing approach, the image patch generation component 420 may generate a plurality of output image patches 422 based on the input image 402, the image masks 412, and one or more visual adaptation criteria 408. In an embodiment, the one or more visual adaptation criteria 408 may include an indication of at least oneof a background extension, a content extension size, a scene adaptation, an object adaptation (e.g., adaptation of appearances, such as clothing, hairstyles, accessories, etc.) or a viewpoint adaptation. In an embodiment, the one or more visual adaptation criteria 408 may include a sample image (e.g., of a background, participant appearances, etc.) to be used as an example for the joint in-out-painting. In an embodiment, the one or more visual adaptation criteria 408 may be based on at least one of a text prompt or a user preference (e.g., the user preference configuration 134). The output image patches 422 may be represented by x , where i may vary from 1 to N and N may be any suitable integer (e.g., 2, 3, 4, 5 or more). The output image patches 422 may include one or more in-painted image patches, one or more out-painted image patches, and / or one or more joint in-out-painted image patches. An in-painted images patch may include a portion modifying the content within the input image 402. An out-painted images patch may include a new portion extending a view of the original input image 402. A joint in-out-painted image patch may include at least a first portion extending a view of the original input image 402 and a second portion modifying the content within the input image 402.

[0069] In an embodiment, the image patch generation component 420 may use a generative network (e.g., one or more computer vision models 132) to generate the output image patches 422. For instance, the generative network may include a deep feature extraction network that extracts a rich, hierarchical set of feature maps from the input image 402. The feature maps may capture low- level details (e.g., at a pixel level) and high-level semantic information from the known parts of the input image 402. The generative network may further include a context encoder that may generate context embeddings from features in the unmasked regions of the input image 402. The context embeddings may capture an understanding of the overall scene in the input image 402 so that a coherent generation of both in-painting areas and out-painting areas can be achieved. The generative network may further include a mask-aware feature processing component that may make the generative network explicitly aware of the masked regions and learn to propagate information from the known regions of the input image 402 into the masked feature space. The generative network may further perform visual feature fusion to fuse deep features representing the global context and the mask-aware local features using deep learning operations (e.g., by concatenating the feature maps along the channel dimension and using attention to weight the importance of different feature maps during the fusion process).

[0070] Next, the output refinement component 430 may refine the content of the initially generated output image patches 422 to generate refined image patches 432. For instance, the output refinement component 430 may refine or improve the local coherence within each generated output image patch 422, add finer details to the generated output image patches 422, and refine the global consistency across the in-painted image portions (within the in-painted image patches 422), the out- painted image portions (within the out-painted image patches 422 and / or joint in-out-painted image patches 422), and / or the original image portions within the output image patches 422. The refined image patches 432 may be represented by x , where t may vary from 0 to N-l.

[0071] Next, the visual composition and blending component 440 may take the input image 402 and the refined image patches 432 as input and generate a synthesized joint in-out-painted image 406, represented by A*, based on the input image 402 and the refined image patches 432. The visual composition and blending component 440 may determine the spatial relationship between the refined image patches 432 and / or the elements in the refined image patches 432 and predict the blended result. The visual composition and blending component 440 may ensure a smooth and visually natural transition, textures, and / or lighting, for example, across the in-out-painting boundaries.

[0072] FIGS. 5A-5C are block diagrams illustrating an example implementation 500 of the joint in-out-painting method of FIG. 4 according to an embodiment of the present disclosure. In an embodiment, the implementation 500 may be implemented by a visual adaptation engine (e.g., the visual adaptation engine 128). For simplicity, FIGS. 5A-5C may use the same reference numerals as in FIG. 4 to refer to the same elements and / or components.

[0073] Turning now to FIG. 5 A, the image mask generation component 410 may receive an input image 502 (e.g., the images 103, 105, 107, 203, 303, and / or 402) and may generate an image mask 504 (e.g., the image mask 412). As similarly discussed above, the image mask 504 may include a value of 1 for a pixel corresponding to a pixel in a region of interest of the input image 502 (e.g., depicted by the grey color in FIG. 5 A) and a value of 0 for a pixel corresponding to a pixel outside the region of interest (e.g., depicted by the white color in FIG. 5A). FIG. 5A may provide a more detailed view of internal components and / or operations of the image patch generation component 420. For example, the image patch generation component 420 may include a patch generation and layout encoding component 510, a mask patch generation component 520, and a masked image patch generation component 524.

[0074] The patch generation and layout encoding component 510 may generate patch layout encoded information 514 from the input image 502. The patch generation and layout encoding may include determining a patch layout including the locations of image patches to be processed for inpainting and / or out-painting and the spatial relationships between those image patches. In some instances, the determination of those patch locations may be based on one or more vision adaptation criteria (e.g., the vision adaptation criteria 408 discussed above with reference to FIG. 4). To determine those patch locations, the patch generation and layout encoding component 510 may extract features (e.g., high-level semantic information and / or low-level detail) from the input image 502 to determine which portion(s) of the input image 502 are to be processed for in-painting, which portion(s) of the input image 502 are to be processed for out-painting, and which portion(s) of the input image 502 are to be processed for in-painting and out-painting.

[0075] Generally, the patch generation and layout encoding component 510 may determine a plurality of patch locations (e.g., N patch locations) to be processed for visual adaptation. For simplicity, FIG. 5A only illustrates five image patch locations shown by the bounded boxes 512 (individually shown as 512a, 512b, 512c, 512d, and 512e). The bounded boxes 512a-512e may generally include non-overlapping regions and / or partially overlapping regions of the input image 502. In some instances, the patch generation and layout encoding component 510 may also define regions that are beyond the original size or canvas of the original input image 502 (e.g., adding blank placeholders that are to be filled to grow the image scene outward, e.g., for out-painting). The patch layout encoded information 514 may provide layout information, such as the shapes and sizes of the patches (the bounded boxes 512), the spatial relationship between the patches, and the locations or layout (e.g., pixel coordinate information) of the patches with respect to the input image 502. In some instances, the patches can have the same shape and / or the same size. In other instances, the patches can have different shapes and / or different sizes. In an embodiment, the patch generation and layout encoding component 510 may determine the patch layout based on one or more visual adaptation criteria (e.g., the visual adaptation criteria 408)

[0076] The mask patch generation component 520 may generate, based on the patch layout encoded information 514, a plurality of mask patches 522, represented by m1m2, ... mN. That is, the mask patches 522 may correspond to portions of the input image 402 to be processed for inpainting and / or out-painting. For simplicity, FIG. 5A only shows the expanded view for the mask patches 522a and 522b. Next, the masked image patch generation component 524 may generatemasked image patches 526 based on the input image 502 and the mask patches 522. For instance, the masked image patch generation component 524 may apply each mask patch 522a, 522b, 522c, 522d, and 522e to the input image 502 (e.g., based on a tensor product computation) to generate a respective masked image patch 526, represented by xlzx2, ... xN. As an example, for each pixel, p, in the input image 502, if the pixel is unmasked, the pixel value for the respective pixel in the masked image patch 526 may be set to a value of 0. If, however, the pixel is masked, the pixel value for the respective pixel in the masked image patch 526 may be set to the original pixel value of the pixel p in the input image 402.

[0077] Turning to FIG. 5B, the image patch generation component 420 may further apply generative Al network(s) 530 (e.g., computer vision model(s) 132) to each masked image patches 526 to generate a respective output image patch 542 (e.g., corresponding to the output image patches 422). In an example, a generative Al network 530 may include an encoding network 532 and a decoding network 536 trained to perform in-painting (e.g., remove, add, or replace element(s) in an image) and / or out-painting (e.g., add new portions to extend a view of an image). The encoding network 532 may encode each input masked image patch 526, x into respective encoded embeddings 534, represented by x;= En(xi), where t may vary from 1 to N and £'n( ) may represent the embedding encoder function. The decoding network 536 may decode the encoded embeddings 534, xt, to generate respective output image patches 542, represented by x = De(xt), where i may vary from 1 to N and £>e( ) may represent the decoding function. In an example, the encoding network 532 may extract meaningful features from a masked image patch 526 and represent the extracted features in a vector space as encoded embeddings 534 (e.g., a context vector or latent representations), and the decoding network 536 may generate a respective output image patch 542 from the respective encoded embeddings 534.

[0078] The image patch generation component 420 may generally process the masked image patches 526 in any suitable order. In an embodiment, the image patch generation component 420 may use a grow and complete approach. For instance, the image patch generation component 420 may process (e.g., in-painting) the masked image patch(es) 526 that are completely within the original input image 502, followed by processing (e.g., in-painting) the masked image patch(es) 526 that include at least a portion of the originally input image 502, then the masked image patch(es) 526 that partially cover those already processed masked image patch(es) 526, and so on. The resulting output image patches 542 may include in-painted portion(s) and / or out-painted portion(s). The growand complete approach may generally provide more realistic and context-aware in-painting and / or out-painting.

[0079] Turning now to FIG. 5C, the output refinement component 430 may apply a refinement network 550 (e.g., computer vision model(s) 132) to the output image patches 542 to generate respective refined image patches 552. The refinement network 550 may be trained to improve the local coherence within an output image patch 542, add finer details to an output image patch 542, and ensure the global consistency across the in-painted image portions, the out-painted image portions, and / or the original image portions in the image patches 542. In an example, the refinement network 550 may have a substantially similar encoder-decoder network architecture as the generative Al network(s) 530 of FIG. 5B. Next, the visual composition and blending component 440 may apply a visual composition and blending network 560 (e.g., computer vision model(s) 132) to the refined image patches 552, the input image 502, and the patch layout encoded information 514 to generate a synthesized joint in-out-painted image 506. The visual composition and blending network 560 may be trained to combine the refined image patches 552 with the original portions of the input image 502 and blend the transition boundaries (e.g., in terms of scenes, objects, textures, colors, and / or lighting) between the different image portions (e.g., the in-painted portions, out-painted portions, and / or the original image portions) to ensure a smooth and visually natural transition in the synthesized joint in-out-painted image 506. In some instances, the patch layout encoded information 514 may be propagated along the visual adaptation processing chain to provide information on how to compose the synthesized image 506 from the refined image patches 552. In an example, the visual composition and blending network 560 may have a substantially similar encoder-decoder network architecture as the generative Al network 530 of FIG. 5B. As shown in FIG. 5C, the synthesized image 506 may include a portion extending the view of the input image 502 (e.g., shown by the ovals and triangles) and a portion modifying the content within the input image 502 (e.g., represented by a different color than the input image 502).

[0080] FIG. 6 is a block diagram illustrating another example joint in-out-painting method 600 according to an embodiment of the present disclosure. In an embodiment, the method 600 may be performed by a visual adaptation engine (e.g., the visual adaptation engine 128) to enhance user visual experience. The method 600 may be substantially similar to the method 400. For simplicity, FIG. 6 may use the same reference numerals as in FIG. 4 to refer to the same elements and / or components. As shown in FIG. 6, the image patch generation component 420 may include aconcatenation of an in-painting generative network 610 and an out-painting generative network 620 to generate in-out-painted image patches 622. For instance, the in-painting generative network 610 may process (e.g., in-paint) the image mask and the input image 402 to generate intermediate image patches 612 and the out-painting generative network 620 may process (e.g., out-paint) the intermediate image patches 612 to generate in-out-painted image patches 622. In an embodiment, the image patch generation component 420 may enhance or refine the intermediate image patches 612 (e.g., in terms of local coherency, visual details, and / or global consistency). Generally, the image patch generation component 420 may apply the in-painting generative network 610 and the out- painting generative network 620 in any order and / or iterate multiple iterations of in-painting and out- painting. In an example, the in-painting generative network 610 and / or the out-painting generative network 620 may have a substantially similar encoder-decoder network architecture as the generative Al network(s) 530 of FIG. 5B. After performing in-painting and out-painting, the output refinement network 550 may refine the local coherency, global consistency, and / or details of the in-out-painted image patches 622. Subsequently, the refined image patches 632 may be processed by the visual composition component 640 and the blending component 650 as discussed above with reference to the visual composition and blending component 440 to generate a synthesized joint in-out-painted image 606.

[0081] In some embodiments, instead of using a concatenation of an in-painting generative network 610 and an out-painting generative network 620, the image patch generation component 420 may use a joint in-out-painting generative network (e g., a computer vision model 132) to generate joint in-out-painted image patches in a single step. The joint in-out painting generative Al network may be trained to perform point in-out-painting. In some embodiments, the joint in-out painting generative network may have an encoded-decoder network architecture as the generative Al network 530 of FIG. 5B. Generally, the joint in-out-painting generative Al network may be a large network with long-sequence cross-attention.

[0082] FIG. 7 is a block diagram illustrating an example quality-aware adaptation method 700 according to an embodiment of the present disclosure. In an embodiment, the method 700 may be performed by a visual adaptation engine (e.g., the visual adaptation engine 128), an environmental sensing engine (e.g., the environmental sensing engine 124), and / or a quality measurement engine (e.g., the quality measurement engine 126). The method 700 may begin with receiving an input image 702 (e.g., the image frames 140, 240, and / or 340, the images 103, 105, 107, 203, 303, 402,and / or 502). At operation 710, color space conversion may be performed on the input image 702, for example, by the visual adaptation engine. The color space conversion may include transforming an image or pixel data from one color representation (color space) to another.

[0083] At operation 712, color and tone adjustment may be further applied to the color space converted image 711. The color and tone adjustment may include enhancing, correcting, and / or modifying the colors and contrast and / or brightness (the tone) of the image 711. The color and tone adjustment may be based on environmental condition sensing information 722, human perception aware quality assessment result 732, and / or user preferences and / or profile information 742. For instance, at operation 720, environmental condition sensing (e.g., lighting, etc.) associated with the capturing of the input image 702 may be performed, for example, by the environmental condition sensing engine. At operation 730, human perception aware quality assessment of the input image 702 may be performed, for example, by the quality measurement engine. The human perception aware quality assessment may include an assessment of color, hues, saturation, contrast, brightness, sharpness, and / or any visual properties that a human may perceive as visually appealing, natural, non-appealing, and / or unnatural. The user preferences and / or profile information 742 may be obtained from a user preference and profile configuration 740. The user preference and / or profile information 742 may include an indication of color, hue, saturation, contrast, brightness, sharpness preferences, etc.

[0084] Next, at operation 714, sharpening and noise reduction may be performed on the color and tone adjusted image 713, for example, by the visual adaptation engine, to generate an output image 706. The sharpening and noise reduction may be based on the human perception aware quality assessment result 732 and / or the user preference and / or profile information 742. In some embodiments, the operations at operations 710, 712, and / or 714 may be performed using techniques, such as histogram equalization, contrast stretching, and / or gamma value adjustment. Histogram equalization may refer to the process where the histogram of pixel intensities is flattened or spread- out across the full value range (e.g., 0-255 for 8-bit pixel values). Contrast stretching may refer to the process of expanding the dynamic range of pixel values to use the full intensity scale. Gamma value adjustment may refer to the process of adjusting the brightness of an image using a nonlinear curve.

[0085] FIG. 8 is a block diagram illustrating an example video conferencing system architecture 800 according to an embodiment of the present disclosure. The video conferencing systemarchitecture 800 may use joint sender-receiver device processing to provide environment, user preference, content, and / or quality aware visual adaptation. In an embodiment, the environment, user preference, content, and / or quality aware visual adaptation discussed herein may be implemented using the video conferencing system architecture 800.

[0086] The video conferencing system architecture 800 may include a sender device 810 and a receiver device 830 connected by a network 820. The network 820 may be substantially similar to the network 110. The sender device 810 and the receiver device 830 may be devices (e.g., similar to the computing devices 120, 136, and / or 138) used by participants (e.g., the participants 102, 104, and / or 106) of a video conference. For instance, a video conferencing application (e.g., the video conferencing application 122) may be executed on each of the sender device 810 and the receiver device 830 for the video conference. As shown, the sender device 810 may include an environmental sensing engine 812 (e.g., the environmental sensing engine 124), an application and device resource adaptation component 814 (e.g., the application and device resource adaptation component 131), and a quality measurement engine 816 (e.g., the quality measurement engine 126). The receiver device 830 may include a visual adaptation engine 832 (e.g., the visual adaptation engine 128), a user preference processing engine 834, and a user preference and profile configuration 836 (e.g., the user preference configuration 134).

[0087] During the video conference, the sender device 810 may capture an input image 802 of a participant of the sender device 810 and may transmit the input image 802 to the receiver device 830 via the network 820. At the sender side, the environmental sensing engine 124 may sense the surroundings (e.g., light condition) of the sending participant. The application and device resource adaptation component 814 may determine parameters for adjusting the behavior and / or operations of the video conferencing application executing on the sender device 810 and / or the visual adaptation algorithm(s) to optimize performance and / or user experience. The determination of the adjustment parameters may be based on the network conditions (e.g., bandwidth, latency, packet error rate, packet loss, etc.), device resources (e.g., processing and / or memory resources), and / or display capabilities (e.g., screen size, resolution, etc.). The quality measurement engine 816 may determine whether the environmental sensing information sensed by the environmental sensing engine 812 satisfies certain criteria (e.g., visual quality criteria). The quality measurement engine 816 may transmit the quality measurement information and / or the adjustment parameters to the receiver device 830 over the network 820.

[0088] At the receiver side, the receiver device 830 may receive the input image 802, the quality measurement information, and / or the adjustment parameters. The visual adaptation engine 832 may adapt the visual properties of the received image 802 to generate an output image 806 for display on a display device (e.g., the display device 130) of the receiver device 830. The visual adaptation may be based on the received quality measurement information, adjustment parameters, and / or user preference and / or profile information provided by the user preference processing engine 834. For instance, the user preference and / or profile information of the receiving participant of the receiver device 830 may be stored at the user preference and profile configuration 836. The visual adaptation may include human perception adaptive quality-aware adaptation, user preference-aware adaptation, environment sensing-aware adaptation, viewpoint adaptation, out-painting, in-painting, and / or joint in-out-painting for enhanced visual experience, flexible participant appearance adaptation, flexible environmental condition adaptation, and / or intelligent static and dynamic content adaptation as discussed above with reference to FIGS. 1-4, 5A-5C, 6-7.

[0089] FIG. 9 is a block diagram illustrating another example video conferencing system architecture 900 according to an embodiment of the present disclosure. In an embodiment, the environment, user preference, content, and / or quality aware visual adaptation discussed herein may be implemented using the video conferencing system architecture 900. For simplicity, FIG. 9 may use the same reference numerals as in FIG. 8 to refer to the same elements and / or components. The video conferencing system architecture 900 may be substantially similar to the video conferencing system architecture 800. However, the video conferencing system architecture 900 may use joint cloud-device processing to provide environment, user preference, content, and / or quality aware visual adaptation. More specifically, the visual adaptation may be distributed between a visual adaptation engine 912 at a network device 910 (e.g., a cloud server or an edge server) in the network 820 and the receiver device 830 as shown in FIG. 9. Stated differently, the visual adaptation processing performed by the visual adaptation engine 832 may be split between the network device 910 and the lightweight visual adaptation engine 920 at the receiver device 830.

[0090] Generally, the visual adaptation processing discussed herein may be distributed across a sender device 810, a network device 910 (e.g., a cloud server or an edge server), and / or a receiver device 830 in any suitable way to provide optimize network and / or device resource usages, video conferencing application performance, and user experience. The visual adaptation processing may be further distributed between a cloud server and an edge server for joint cloud-edge-deviceprocessing. For instance, operations that are related to static content adaptation may be performed at a cloud device and operations that are related to dynamic content adaptation may be performed at an edge device. In some instances, the distribution (or offloading) may be based on the availability of network and / or device processing and / or memory resources, network conditions, and / or application real-time requirements.

[0091] FIG. 10 is a flowchart of a method 1000 implemented by a computing device according to an embodiment of the present disclosure. In an embodiment, the method 1000 may be implemented by or on a personal computer (PC), a smart phone, a smart tablet, or some other computing device used to participate in a video conference, a virtual real-time event, and / or other image / video related applications. In an embodiment, the method 1000 may be implemented by a computing device (e.g., the computing device 120) including an environmental sensing engine (e.g., the environmental sensing engines 124 and / or 812), a quality measurement engine (e g., the quality measurement engines 126 and / or 816), and a visual adaptation engine (e.g., the visual adaptation engines 128, 832, 912, and / or 920). In an embodiment, the method 1000 may be implemented using a computing device with components as shown in FIG. 11. The method 1000 may use similar mechanisms as discussed above with reference to FIGS. 1-4, 5A-5C, 6-9. As illustrated, FIG. 10 includes a number of enumerated operations, but embodiments of the operations in FIG. 10 may include additional operations before, after, and in between the enumerated operations. In some embodiments, one or more of the enumerated operations may be omitted or performed in a different order.

[0092] At operation 1002, the computing device receives an input image (e.g., an image frame 140 and / or an image 103, 105, 107, 203, 303, 402, 502, 702 or 802). At operation 1004, the computing device adapts, based on one or more visual adaptation criteria (e.g., the visual adaptation criteria 408), one or more visual properties of the input image to generate a synthesized image (e.g., an image frame 140 and / or an image 142, 144, 146, 242, 342, 406, 506, 606, 706, or 806). In an embodiment, the one or more visual adaptation criteria includes an indication of at least one of a background extension, a content extension size, a scene adaptation, an object adaptation, or a viewpoint adaptation. In an embodiment, the one or more visual adaptation criteria includes a sample image. In an embodiment, the one or more visual adaptation criteria are based on at least one of a text prompt or a user preference (e.g., the user preference configuration 134 and / or the userpreference and profile configurations 740 and / or 836). As part of adapting the one or more visual properties, the computer device performs operations 1006, 1008, and 1010.

[0093] At operation 1006, the computing device generates one or more masked image patches (e.g., the masked image patches 526), each including a respective masked portion of the input image. In an embodiment, as part of the generating the one or more masked image patches, the computer device generates, based on the one or more visual adaptation criteria, one or more mask patches (e.g., the mask patches 522), each masking a respective portion of the input image for visual adaptation. The computing device generates the one or more masked image patches based on respective ones of the one or more mask patches and the input image.

[0094] At operation 1008, the computer device processes the one or more masked image patches using one or more computer vision models (e.g., the computer vision models 132 and / or the generative Al network(s) 530, to generate respective ones of one or more output image patches (e.g., the output image patches 542). In an embodiment, the one or more computer vision models used for generating the one or more output image patches comprises one or more generative Al models (e.g., the generative Al network(s) 530, the in-painting generative Al network 610, and / or the out-painting generative Al network 620).

[0095] In an embodiment, the one or more computer vision models used for generating the one or more output image patches comprises a generative Al model trained for joint in-out-painting. In an embodiment, the one or more computer vision models used for generating the one or more output image patches comprises a first generative Al model (e.g., the in-painting generative Al network 610) trained for in-painting and a second generative Al model (e.g., the out-painting generative Al network 620) trained for out-painting. In a further embodiment, as part of processing the one or more masked image patches at operation 1008, the computing device processes an individual one of the one or more masked image patches using one of the first generative Al model or the second generative Al model to generate an intermediate image patch and process the intermediate image patch using the other one of the first generative Al model or the second generative Al model to generate a corresponding one of the one or more output image patches.

[0096] In an embodiment, as part of processing the one or more masked image patches at operation 1008, the computing device processes a first masked image patch using a first computer vision model of the one or more computer vision models before processing a second masked image patch using a second computer vision model of the one or more computer vision models based onthe first masked image patch being completely within the input image and the second masked image patch having a portion within the input image and a portion outside of the input image (e.g., a grow and complete approach).

[0097] At operation 1010, the computer device generates the synthesized image based on the input image and the one or more output image patches. The synthesized image includes at least a first portion extending a view of the input image and a second portion modifying content within the input image. In an embodiment, as part of generating the synthesized image, the computing device refines at least one of a local coherence, a global coherence, or content detail of the one or more output image patches using a refinement network (e.g., the refinement network 550). Further, the computing device generates the synthesized image based on the refined one or more output image patches, the input image, and patch layout information associated with the one or more output image patches using a blending and composition network (e.g., the visual composition and blending network 560, the visual composition component 640, and / or the blending component 650).

[0098] At operation 1012, the computing device renders the synthesized image on a display device (e.g., the display device 130).

[0099] In an embodiment, the input image received at operation 1002 includes a plurality of images of participants of a real-time virtual event (e.g., a video conference, etc.). In an embodiment, the adapting the one or more visual properties of the input image at operation 1004 is further based on a visual quality assessment (e g., including threshold comparisons, etc.) across the plurality of images. In an embodiment, the adapting the one or more visual properties of the input image at operation 1004 is further based on information obtained from environmental condition sensing (e.g., of lighting condition) associated with capturing of one or more of the plurality of images. In an embodiment, the one or more visual properties of the input image being adapted at operation 1004 are associated with at least one of a background (e.g., from a blank wall to an office environment, etc.), a viewpoint (e.g., from a side profile view to a frontal view, etc.), a subject framing style (e.g., from a headshot view to a half body view, etc.), or a subject appearance (e.g., in terms of clothing, hairstyle, etc.) associated with an individual participant of the plurality of participants. In an embodiment, the synthesized image includes an image of two or more of the participants within the same virtual environment (e.g., for immersive experience as shown in the image 342 discussed above with reference to FIG. 3).

[0100] In an embodiment, the adapting the one or more visual properties of the input image at operation 1004 includes offloading one or more operations associated with the adapting to a network server (e.g., the network device 910) associated with the real-time virtual event. In an embodiment, the offloading the one or more operations associated with the adapting to the network server is based on at least one of a resource assessment or a real-time requirement (e.g., associated with a video conferencing application 122).

[0101] FIG. 11 is a schematic diagram of an apparatus 1100 (e.g., a network apparatus, a network node, a network router, a router, etc.). The apparatus 1100 is suitable for implementing the disclosed embodiments as described herein. The apparatus 1100 comprises ingress ports / ingress means 1110 (a.k.a., upstream ports) and receiver units (Rx) / receiving means 1120 for receiving data; a processor, logic unit, or central processing unit (CPU) / processing means 1130 to process the data; transmitter units (Tx) / transmitting means 1140 and egress ports / egress means 1150 (a.k.a., downstream ports) for transmitting the data; and a memory / memory means 1160 for storing the data. In an embodiment, the receiver units (Rx) / receiving means 1120 comprise a discrete circuit, integrated circuit, chip set, package, hardware module, electronic device, or other structure capable of receiving signals. In an embodiment, the transmitter units (Tx) / transmitting means 1140 comprise a discrete circuit, integrated circuit, chip set, package, hardware module, electronic device, or other structure capable of transmitting signals. The apparatus 1100 may also comprise optical-to-electrical (OE) components and electrical-to-optical (EO) components coupled to the ingress ports / ingress means 11 10, the receiver units / receiving means 1 120, the transmitter units / transmitting means 1 140, and the egress ports / egress means 1150 for egress or ingress of optical or electrical signals.

[0102] The processor / processing means 1130 is implemented by hardware and software. The processor / processing means 1130 may be implemented as one or more CPU chips, cores (e.g., as a multi-core processor), field-programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), and digital signal processors (DSPs). The processor / processing means 1130 is in communication with the ingress ports / ingress means 1110, receiver units / receiving means 1120, transmitter units / transmitting means 1140, egress ports / egress means 1150, and memory / memory means 1160. The processor / processing means 1130 comprises a visual adaptation module 1170. The visual adaptation module 1170 may be formed by software including program instructions stored at the memory / memory means 1160 and executed by the processor / processor means 1130, hardware (e.g., discrete logic, ASIC, or FPGA), or a combination of software and hardware and is configuredto implement the methods disclosed herein. The inclusion of the visual adaptation module 1170 therefore provides a substantial improvement to the functionality of the apparatus 1100 and effects a transformation of the apparatus 1100 to a different state. Alternatively, the visual adaptation module 1170 is implemented as instructions stored in the memory / memory means 1160 and executed by the processor / processing means 1130.

[0103] The apparatus 1100 may also include input and / or output (I / O) devices or I / O means 1180 for communicating data to and from a user. The I / O devices or I / O means 1180 may include output devices such as a display for displaying video data, speakers for outputting audio data, etc. The I / O devices or I / O means 1180 may also include input devices, such as a keyboard, mouse, trackball, etc., and / or corresponding interfaces for interacting with such output devices.

[0104] The memory / memory means 1160 comprises one or more disks, tape drives, and solid- state drives and may be used as an over-flow data storage device, to store programs when such programs are selected for execution, and to store instructions and data that are read during program execution. The memory / memory means 1160 may be volatile and / or non-volatile and may be readonly memory (ROM), random access memory (RAM), ternary content-addressable memory (TCAM), and / or static random-access memory (SRAM).

[0105] While several embodiments have been provided in the present disclosure, it should be understood that the disclosed systems and methods might be embodied in many other specific forms without departing from the spirit or scope of the present disclosure. The present examples are to be considered as illustrative and not restrictive, and the intention is not to be limited to the details given herein. For example, the various elements or components may be combined or integrated in another system or certain features may be omitted, or not implemented.

[0106] In addition, techniques, systems, subsystems, and methods described and illustrated in the various embodiments as discrete or separate may be combined or integrated with other systems, modules, techniques, or methods without departing from the scope of the present disclosure. Other items shown or discussed as coupled or directly coupled or communicating with each other may be indirectly coupled or communicating through some interface, device, or intermediate component whether electrically, mechanically, or otherwise. Other examples of changes, substitutions, and alterations are ascertainable by one skilled in the art and could be made without departing from the spirit and scope disclosed herein.

Claims

CLAIMSWhat is claimed is:

1. A computer-implemented method, comprising: receiving an input image; adapting, based on one or more visual adaptation criteria, one or more visual properties of the input image to generate a synthesized image, wherein the adapting comprises: generating one or more masked image patches, each comprising a respective masked portion of the input image; processing the one or more masked image patches using one or more computer vision models to generate respective ones of one or more output image patches; and generating the synthesized image based on the input image and the one or more output image patches, the synthesized image comprising at least a first portion extending a view of the input image and a second portion modifying content within the input image; and rendering, on a display device, the synthesized image.

2. The method of claim 1, wherein the one or more visual adaptation criteria comprises an indication of at least one of a background extension, a content extension size, a scene adaptation, an object adaptation, or a viewpoint adaptation.

3. The method of any of claims 1-2, wherein the one or more visual adaptation criteria comprises a sample image.

4. The method of any of claims 1-3, wherein the one or more visual adaptation criteria are based on at least one of a text prompt or a user preference.

5. The method of any of claims 1-4, wherein the generating the one or more masked image patches comprises: generating, based on the one or more visual adaptation criteria, one or more mask patches, each masking a respective portion of the input image for visual adaptation; andgenerating the one or more masked image patches based on respective ones of the one or more mask patches and the input image.

6. The method of any of claims 1-5, wherein the one or more computer vision models used for generating the one or more output image patches comprises one or more generative artificial intelligence (Al) models.

7. The method of any of claims 1-6, wherein the one or more computer vision models used for generating the one or more output image patches comprises a first generative artificial intelligence (Al) model trained for in-painting and a second generative Al model trained for out-painting.

8. The method of claim 7, wherein the processing the one or more masked image patches comprises: processing an individual one of the one or more masked image patches using one of the first generative Al model or the second generative Al model to generate an intermediate image patch; and processing the intermediate image patch using the other one of the first generative Al model or the second generative Al model to generate a corresponding one of the one or more output image patches.

9. The method of any of claims 1-6, wherein the one or more computer vision models used for generating the one or more output image patches comprises a generative artificial intelligence (Al) model trained for joint in-out-painting.

10. The method of any of claims 1-9, wherein the generating the synthesized image comprises: refining at least one of a local coherence, a global coherence, or content detail of the one or more output image patches using a refinement network; and generating the synthesized image based on the refined one or more output image patches, the input image, and patch layout information associated with the one or more output image patches using a blending and composition network.

11. The method of any of claims 1 -9, wherein the processing the one or more masked image patches comprises: processing a first masked image patch using a first computer vision model of the one or more computer vision models before processing a second masked image patch using a second computer vision model of the one or more computer vision models based on the first masked image patch being completely within the input image and the second masked image patch having a portion within the input image and a portion outside of the input image.

12. The method of any of claims 1-11, wherein the input image comprises a plurality of images of participants of a real-time virtual event.

13. The method of any of claims 1-12, wherein the adapting the one or more visual properties of the input image is further based on a visual quality assessment across the plurality of images.

14. The method of any of claims 1-13, wherein the adapting the one or more visual properties of the input image is further based on information obtained from environmental condition sensing associated with capturing of one or more of the plurality of images.

15. The method of any of claims 1-14, wherein the adapting comprises adapting at least one of a background, a viewpoint, a subject framing style, or a subject appearance associated with an individual participant of the plurality of participants.

16. The method of any of claims 1-15, wherein the synthesized image comprises an image of two or more of the participants within the same virtual environment.

17. The method of any of claims 1-16, wherein the adapting the one or more visual properties of the input image comprises offloading one or more operations associated with the adapting to a network server associated with the real-time virtual event.

18. The method of any of claims 1 -17, wherein the offloading the one or more operations associated with the adapting to the network server is based on at least one of a resource assessment or a real-time requirement.

19. The method of any of claims 1-18, wherein one or more operations associated with the adapting the one or more visual properties of the input image is performed at a computing device associated with a receiver of the input image.

20. A computing device, comprising: a memory storing instructions; and one or more processors coupled to the memory, the one or more processors configured to execute the instructions to cause the computing device to implement the method in any of claims 1 - 19.

21. A non-transitory computer readable medium comprising a computer program product for use by a computing device, the computer program product comprising computer executable instructions stored on the non-transitory computer readable medium that, when executed by one or more processors, cause the computing device to execute the method in any of claims 1-19.

22. A computing device, comprising: means for receiving an input image; means for adapting, based on one or more visual adaptation criteria, one or more visual properties of the input image to generate a synthesized image, wherein the adapting comprises: generating one or more masked image patches, each comprising a respective masked portion of the input image; processing the one or more masked image patches using one or more computer vision models to generate respective ones of one or more output image patches; and generating the synthesized image based on the input image and the one or more output image patches, the synthesized image comprising at least a first portion extending a view of the input image and a second portion modifying content within the input image; and means for rendering, on a display device, the synthesized image.

Citation Information

Patent Citations

  • Information processing apparatus, information processing system, information processing method, and program

    US20190014288A1

  • Conditional image generation

    US20250200827A1

  • Adaptive teleconferencing experiences using generative image models

    US20250200843A1

  • Distributed generation of virtual content

    WO2024040054A1