System and method for image correction in a camera system using adaptive image deformation
Adaptive image deformation technology corrects the image distortion of the camera system and dynamically adjusts the view and perspective, solving the image distortion and insufficient field of view problems caused by wide-angle lenses, and improving the user experience and space utilization efficiency of video meetings.
Patent Information
- Application Number
- CN202510270759.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-03-08
- Filing Date
- 2025-03-07
- Publication Date
- 2025-09-09
AI Technical Summary
Existing camera systems are prone to image distortion when shooting with wide-angle or fisheye lenses, and are unable to dynamically adjust the view and perspective angle to adapt to different meeting room layouts and participant positions, resulting in some participants in the video meeting being invisible or distorted.
Adaptive image warping technology is used to process the video stream obtained by the camera through the image warping unit, identify the area of interest, adjust the perspective to generate a natural and attractive video stream, and dynamically adjust the camera view according to the participant's position and attention.
Effectively correct image distortion to ensure all participants are visible on the screen, improve inclusivity and participation in video meetings, enhance user experience, and optimize space utilization and equipment costs.
Smart Images

Figure FT_1 
Figure FT_2 
Figure FT_3
Abstract
Description
Technical Field
[0001] The present disclosure relates generally to camera systems and, more particularly, to systems and methods for correcting image distortion in camera systems in real time using adaptive image warping.
[0002] In some conference room settings, wide-angle or fisheye lenses may be used to capture a larger field of view. Such lenses may distort the image or video stream captured and displayed by the lens, causing objects (e.g., meeting participants, whiteboards, furniture) to appear stretched or distorted, and straight lines to appear curved. Therefore, there is a need for a camera system or method that can correct for such distortion, so that the video feed in, for example, a video conference appears more natural and visually appealing.
[0003] Furthermore, conference rooms come in all sizes and shapes, and not all participants in a meeting or video conference can be seated directly in front of the camera. Therefore, a need exists for a camera system or method that can adjust its view and perspective to ensure that most or all participants are visible in the video stream, display, or screen. Such a system and method can improve inclusivity and ensure that every participant's contribution is equally visible during the meeting.
[0004] Additionally, conference rooms may be used for different purposes. For example, a conference room may be used for a presentation, and the same conference room may later be used for a training session. Therefore, a need exists for a camera system or method that can include adjustable camera settings to ensure that the content being discussed is clear and visible to remote participants in a video conferencing scenario.
[0005] Some conference rooms may have limited space for camera placement, so cameras placed in those conference rooms may not have the best field of view for the conference room. Therefore, there is a need for a camera system or method that can capture a wider field of view and present it in standard screen aspect ratios without unnecessary cropping or distortion. Summary of the Invention
[0006] The disclosed embodiments can address one or more of these challenges. The disclosed cameras and camera systems can include intelligent cameras or multi-camera systems that understand the dynamics of conference room participants (e.g., using artificial intelligence (AI), such as a trained network) and provide an engaging experience to far-end or remote participants based on, for example, the number of people in the room, who is speaking, who is listening, and where attendees are focusing their attention. Examples of conference rooms or meeting environments can include, but are not limited to, meeting rooms, boardrooms, classrooms, lecture halls, conference spaces, and the like.
[0007] Consistent with the disclosed embodiments, a video conferencing system for adjusting perspective views using adaptive image warping includes an image warping unit comprising at least one processor programmed to: receive an overview video stream from a camera in the video conferencing system; determine at least one region of interest represented within the at least one test frame based on analysis of at least one test frame from the overview video stream; determine one or more indicators of an actual camera perspective relative to the at least one region of interest; determine a target camera perspective relative to the at least one region of interest, wherein the target camera perspective is different from the actual camera perspective; determine at least one image transformation based on the difference between the actual camera perspective and the target camera perspective; apply the at least one image transformation to one or more subframe regions of a plurality of image frames of the overview stream to generate at least one image warped primary video stream; and cause the at least one image warped primary video stream to be displayed on a display. BRIEF DESCRIPTION OF THE DRAWINGS
[0008] The accompanying drawings, which are incorporated into and constitute a part of this specification, illustrate the disclosed embodiments and, together with the description, serve to explain the disclosed embodiments. The details shown are by way of example and for the purpose of illustrative discussion of the embodiments of the present disclosure. The description taken in conjunction with the accompanying drawings makes it apparent to those skilled in the art how the embodiments of the present disclosure may be practiced.
[0009] Figure 1 is an illustration of an example of a multi-camera system consistent with some embodiments of the present disclosure.
[0010] Figure 2 is an illustration of a camera including a video processing unit, consistent with some embodiments of the present disclosure.
[0011] Figure 3 is an example illustration of perspective correction of an image or video stream, consistent with some embodiments of the present disclosure.
[0012] Figure 4 1 shows a processing flow in an image deformation system consistent with some embodiments of the present disclosure.
[0013] Figure 5 is a flow chart illustrating an example method of automatic framing consistent with some embodiments of the present disclosure.
[0014] Figure 6 is a flow chart illustrating an example method of constructing a perspective frame fit, consistent with some embodiments of the present disclosure.
[0015] Figure 7 is a flow chart illustrating an example dewarping technique according to an exemplary disclosed embodiment.
[0016] Figure 8 is a flow chart illustrating an example method of mesh generation consistent with some embodiments of the present disclosure.
[0017] Figure 9 Examples of warped image sub-frames of a primary output video stream are shown, consistent with some embodiments of the present disclosure. DETAILED DESCRIPTION
[0018] The present disclosure provides a video conferencing system and a camera system for video conferencing. Therefore, when a camera system is mentioned herein, it should be understood that it may be referred to alternatively as a video conferencing system, a video conferencing camera system or a camera system for video conferencing. As used herein, the term "video conferencing system" refers to a system that can be used for video conferencing (such as a video conferencing camera) and may be referred to alternatively as a system for video conferencing. The video conferencing system does not need to be able to provide video conferencing capabilities by itself, but may interface with other devices or systems (such as a laptop, PC or other networked device) to provide video conferencing capabilities.
[0019] The video conferencing system / camera system according to the present disclosure may include at least one camera and a video processor for processing video output generated by the at least one camera. The video processor may include one or more video processing units.
[0020] According to an embodiment of the present disclosure, the video conferencing camera may include at least one video processing unit. The at least one video processing unit may be configured to process the video output generated by the video conferencing camera. As used herein, the video processing unit may include any electronic circuit designed to read, manipulate and / or change a computer-readable memory to create, generate or process video images and video frames intended to be output (e.g., in a video output or video delivery) to a display device. The video processing unit may include one or more microprocessors or other logic-based devices configured to receive a digital signal representing the acquired image. The disclosed video processing unit may include an application-specific integrated circuit (ASIC), a microprocessor unit or any other suitable structure for analyzing the acquired image, selectively framing an object (subject) based on the analysis of the acquired image, generating an output video stream, etc.
[0021] In some cases, at least one video processing unit may be located within a single camera. In other words, the video conferencing camera may include a video processing unit. In other embodiments, the at least one video processing unit may be located remotely from the camera, or may be distributed across multiple cameras and / or devices. For example, the at least one video processing unit may include more than one (or more) video processing units distributed across a group of electronic devices including one or more cameras (e.g., a multi-camera system), a personal computer, a mobile device (e.g., a tablet, a phone, etc.), and / or one or more cloud-based servers. As another example, the at least one video processing unit may be located on a central video computing unit connected to the camera system or on a cloud computing platform, both of which are remotely located from the at least one camera. Thus, disclosed herein is a video conferencing system comprising at least one camera and at least one video processing unit, as described herein. The at least one video processing unit may or may not be implemented as part of the at least one camera. The at least one video processing unit may be configured to receive video output generated by one or more video conferencing cameras. The at least one video processing unit may decode digital signals to display video and / or may store image data in a memory device. In some embodiments, the video processing unit may include a graphics processing unit. It should be understood that although video processing units are referred to in the singular herein, more than one video processing unit is also contemplated. The various video processing steps described herein can be performed by at least one video processing unit, and the at least one video processing unit can therefore be configured to perform the method (e.g., video processing method) as described herein or any step of such a method. Where determination of a parameter, value, or amount relating to such a method is disclosed herein, it should be understood that the at least one video processing unit can perform such determination and can therefore be configured to perform such determination.
[0022] Single cameras as well as multi-camera systems are described herein. Although some features may be described with reference to a single camera and other features may be described with reference to a multi-camera system, it should be understood that any and all of the features, embodiments, and elements herein may relate to, or be implemented in, both single-camera and multi-camera systems. For example, some features, embodiments, and elements may be described as relating to a single camera system. It should be understood that those features, embodiments, and elements may relate to and / or be implemented in a multi-camera system. In addition, other features, embodiments, and elements may be described as relating to a multi-camera system. It should also be understood that these features, embodiments, and elements may relate to and / or be implemented in a single camera system.
[0023] Embodiments of the present disclosure include a multi-camera system. As used herein, a multi-camera system may include two or more cameras deployed in an environment, such as a conference environment, and may simultaneously record or broadcast one or more representations of the environment. The disclosed cameras may include any device including one or more light-sensitive sensors configured to capture a stream of image frames. Examples of cameras may include, but are not limited to, or S1 camera, Camera, digital camera, smartphone camera, compact camera, digital single-lens reflex (DSLR) video camera, mirrorless camera, action (adventure) camera, 360-degree camera, medium format camera, webcam, or any other device for recording visual images and generating corresponding video signals.
[0024] refer to Figure 1, provides an illustration of an example of a multi-camera system 100 consistent with some embodiments of the present disclosure. Multi-camera system 100 may include a primary camera 110, one or more peripheral cameras 120, one or more sensors 130, and a host computer 140. In some embodiments, primary camera 110 and one or more peripheral cameras 120 may be the same camera type, such as, but not limited to, the examples of cameras discussed above. Furthermore, in some embodiments, primary camera 110 and one or more peripheral cameras 120 may be interchangeable, such that primary camera 110 and one or more peripheral cameras 120 may be located together in a meeting environment, and any one of the cameras may be selected to function as the primary camera. Such selection may be based on various factors, such as, but not limited to, the location of speakers, the layout of the meeting environment, the location of auxiliary items (e.g., whiteboards, presentation screens, televisions), and the like. In some cases, the primary camera and peripheral cameras may operate in a "supervisor-worker" arrangement. For example, the primary camera may include most or all of the components for video processing associated with the multiple outputs of the various cameras included in the multi-camera system. In other cases, the system may include a more distributed arrangement in which the video processing components (and tasks) are more evenly distributed across the various cameras of the multi-camera system. Additionally, in some embodiments, the video processing components may be remotely located relative to the various cameras of the multi-camera system, such as on an adapter, computer, or server / network.
[0025] like Figure 1As shown, the main camera 110 and one or more peripheral cameras 120 can each include an image sensor 111, 121. Furthermore, the main camera 110 and one or more peripheral cameras 120 can include directional audio (DOA / audio) units 112, 122. The DOA / audio units 112, 122 can detect and / or record audio signals and determine the direction from which one or more audio signals originate. In some embodiments, the DOA / audio units 112, 122 can determine, or be used to determine, the direction of a speaker in a conference environment. For example, the DOA / audio units 112, 122 can include a microphone array that can detect audio signals from different positions relative to the main camera 110 and / or one or more peripheral cameras 120. The DOA / audio units 112, 122 can use audio signals from different microphones and determine the angle and / or position from which the audio signal (e.g., speech) originates. Additionally or alternatively, in some embodiments, the DOA / audio units 112, 122 can distinguish between situations in the conference environment where a conference participant is speaking and other situations in the conference environment where there is silence. In some embodiments, determining the direction from which one or more audio signals originate and / or distinguishing between different situations in a conference environment may be determined by units other than the DOA / audio units 112 , 122 , such as one or more sensors 130 .
[0026] The main camera 110 and the one or more peripheral cameras 120 may include vision processing units 113, 123. The vision processing units 113, 123 may include one or more hardware-accelerated programmable convolutional neural networks with pre-trained weights that can detect different properties from video and / or audio. For example, in some embodiments, the vision processing units 113, 123 may use a vision pipeline model (e.g., a machine learning model) to determine the positions of meeting participants in a meeting environment based on the representations of the meeting participants in the overview stream. As used herein, the overview stream may include a video recording of the meeting environment at a standard zoom and perspective of the camera used to capture the recording, or at the camera's most zoomed-out perspective. In other words, the overview shot or stream may include the camera's maximum field of view. Alternatively, the overview shot may be a zoomed or cropped portion of the camera's full video output, but may still capture an overview shot of the meeting environment. Typically, an overview screen or overview video stream can capture an overview of the conference environment and can be framed to feature representations of all or substantially all conference participants, for example, within the field of view of a camera, or present in the conference environment and detected or identified by the system (e.g., by a video processing unit based on analysis of the camera output). A primary or focus stream can include a focused, enhanced, or magnified recording of the conference environment. In some embodiments, the mainstream or focus stream can be a substream of the overview stream. As used herein, a substream can involve a video recording that captures a portion or subframe of the overview stream. In addition, in some embodiments, the visual processing units 113, 123 can be trained to be non-biased with respect to various parameters including, but not limited to, gender, age, race, scene, light, and size, thereby allowing for a robust conference or video conferencing experience.
[0027] like Figure 1As shown, the main camera 110 and one or more peripheral cameras 120 may include virtual director units 114, 124. In some embodiments, the virtual director units 114, 124 can control the main video stream that can be consumed by the connected host computer 140. In some embodiments, the host computer 140 may include one or more of a television, a laptop, a mobile device, a projector, or any other computing system. The virtual director units 114, 124 may include software components that use input from the visual processing units 113, 123 and determine the video output stream, as well as from which camera (e.g., from the main camera 110 and one or more peripheral cameras 120) to stream to the host computer 140. The virtual director units 114, 124 can create an automated experience that may be similar to a television talk show production or an interactive video experience. In some embodiments, the virtual director units 114, 124 can frame the representation of each conference participant in the conference environment. For example, the virtual director units 114, 124 can determine that a camera (e.g., a main camera 110 and / or one or more peripheral cameras 120) can provide an ideal frame or picture of the conference participants in the conference environment. The ideal frame or picture can be determined by various factors, including but not limited to the angle of each camera relative to the conference participants, the position of the conference participants, the level of participation of the conference participants, or other attributes associated with the conference participants. More non-limiting examples of attributes associated with conference participants that can be used to determine the ideal frame or picture of the conference participants can include: whether the conference participant is speaking, the duration of time the conference participant has been speaking, the direction of the conference participant's gaze, the percentage of the conference participant that is visible in the frame, the reactions and body language of the conference participant, or other conference participants that may be visible in the frame.
[0028] The multi-camera system 100 may include one or more sensors 130. The sensors 130 may include one or more smart sensors. As used herein, a smart sensor may include a device that receives input from the physical environment and, upon detecting a specific input, uses built-in or associated computing resources to perform a predefined function and processes the data before sending it to another unit. In some embodiments, the one or more sensors 130 may send data to the main camera 110 and / or one or more peripheral cameras 120, or to at least one video processing unit. Non-limiting examples of sensors may include level sensors, current sensors, humidity sensors, pressure sensors, temperature sensors, proximity sensors, thermal sensors, flow sensors, fluid velocity sensors, and infrared sensors. Furthermore, non-limiting examples of smart sensors may include touchpads, microphones, smartphones, GPS trackers, echolocation sensors, thermometers, humidity sensors, and biometric sensors. Furthermore, in some embodiments, one or more sensors 130 may be placed throughout the meeting environment. Additionally or alternatively, the sensors in the one or more sensors 130 may be of the same type or different types. In other cases, the sensor 130 may generate (one or more) raw signal outputs and send them to one or more processing units, which may be located on the main camera 110 or distributed among two or more cameras included in the multi-camera system. The processing unit(s) may receive the (one or more) raw signal outputs, process the received signals, and use the processed signals to provide various features of the multi-camera system (these features will be discussed in more detail below).
[0029] like Figure 1 As shown, one or more sensors 130 may include an application programming interface (API) 132. Figure 1 As shown, the main camera 110 and the one or more peripheral cameras 120 may include APIs 116, 126. As used herein, an API may refer to a defined set of rules that may enable different applications, computer programs, or units to communicate with each other. For example, the API 132 of the one or more sensors 130, the API 116 of the main camera 110, and the API 126 of the one or more peripheral cameras 120 may be connected to each other (e.g., Figure 110 ), and allows the one or more sensors 130, the primary camera 110, and the one or more peripheral cameras 120 to communicate with each other. It is contemplated that the APIs 116, 126, 132 can be connected in any suitable manner, such as, but not limited to, via Ethernet, a local area network (LAN), a wired or wireless network. It is further contemplated that each of the one or more sensors 130 and each of the one or more peripheral cameras 120 can include an API. In some embodiments, the host computer 140 can be connected to the primary camera 110 via the API 116, and the API 116 can allow communication between the host computer 140 and the primary camera 110.
[0030] The main camera 110 and the one or more peripheral cameras 120 may include stream selectors 115, 125. The stream selectors 115, 125 may receive the overview stream and the focus stream of the main camera 110 and / or the one or more peripheral cameras 120 and provide an updated focus stream (e.g., based on the overview stream or the focus stream) to the host computer 140. The selection of the stream to be displayed to the host computer 140 may be performed by the virtual director unit 114, 124. In some embodiments, the selection of the stream to be displayed to the host computer 140 may be performed by the host computer 140. In other embodiments, the selection of the stream to be displayed to the host computer 140 may be determined by user input received via the host computer 140, where the user may be a conference participant.
[0031] In some embodiments, an autonomous video conferencing (AVC) system is provided. The AVC system may include any or all of the features described above with respect to the multi-camera system 100 in any combination. In addition, in some embodiments, one or more peripheral cameras and smart sensors of the AVC system may be placed in a separate video conferencing space (or conference environment) as an auxiliary space for video conferencing (or conference). These peripheral cameras and smart sensors may be networked with the main camera and may be adapted to provide image and non-image inputs from the auxiliary space to the main camera. In some embodiments, the AVC system may be adapted to generate an automatic television studio production for the combined video conferencing space based on inputs from cameras and smart sensors in the two spaces.
[0032] In some embodiments, the AVC system can include smart cameras adapted to have different field of view angles. For example, in a small video meeting (or conference) space with fewer smart cameras, the smart cameras can have a wide field of view (e.g., approximately 150 degrees). As another example, in a large video meeting (or conference) space with more smart cameras, the smart cameras can have a narrow field of view (e.g., approximately 90 degrees). In some embodiments, the AVC system can be equipped with smart cameras with various field of view angles, thereby allowing for optimal coverage of the video meeting space.
[0033] In addition, in some embodiments, at least one image sensor of the AVC system can be adapted to zoom to 10X, thereby enabling close-up images of objects at the far end of the video meeting space. Additionally or alternatively, in some embodiments, at least one smart camera in the AVC system can be adapted to capture content on or around an object, which can be a non-human item within the video meeting space (or conference environment). Non-limiting examples of non-human items include a whiteboard, a television (TV) display, a poster, or a presentation stand. The camera adapted to capture content on or around an object can be smaller and positioned differently from other smart cameras in the AVC system, and can be mounted to, for example, a ceiling to provide effective coverage of the target content.
[0034] At least one audio device (e.g., a DOA audio device) in a smart camera of an AVC system may include a microphone array adapted to output audio signals representing sounds originating from different locations and / or directions around the smart camera. Signals from the different microphones may allow the smart camera to determine the direction of audio (DOA) associated with the audio signal and to discern, for example, whether there is silence in a particular location or direction. This information may be available to the vision pipeline and virtual director included in the AVC system. Thus, in some embodiments, the machine learning model disclosed herein may include an audio model that provides both direction of audio (DOA) and voice activity detection (VAD) associated with audio signals received from, for example, the microphone array to provide information about when someone is speaking. In some embodiments, a computing device with high computing power may be connected to the AVC system via an Ethernet switch. The computing device may be adapted to provide additional computing power to the AVC system. In some embodiments, the computing device may include one or more high-performance CPUs and GPUs and may run portions of the vision pipeline for the main camera and any designated peripheral cameras.
[0035] In some embodiments, multiple system cameras can create a varied, flexible, and engaging experience by placing multiple wide-field-of-view single-lens cameras that collaborate to frame conference participants in the conference environment as they join and participate in the conversation from different camera angles and zoom levels. This can give far-end participants (e.g., participants located further away from the camera, participants participating remotely or via video conferencing) a natural sense of what is happening in the conference environment.
[0036] The disclosed embodiments may include a multi-camera system including multiple cameras. Each camera may be configured to generate a video output stream representing a conference environment. Each video output stream may feature one or more conference participants present in the conference environment. In this context, "featured" means that the video output stream includes representations of one or more conference participants or features representations of one or more conference participants. For example, a first representation of a conference participant may be included in a first video output stream from a first camera included in the multiple cameras, and a second representation of the conference participant may be included in a second video output stream from a second camera included in the multiple cameras. As used herein, a conference environment may refer to any space in which there is a gathering of people interacting with each other. Non-limiting examples of conference environments may include a boardroom, a classroom, a lecture hall, a video conferencing space, or an office space. As used herein, a representation of a conference participant may refer to an image, video, or other visual rendering of the conference participant that can be captured, recorded, and / or displayed, for example, to a display unit. The video output stream or video stream may refer to a media component (which may include visual and / or audio renderings) that can be delivered to, for example, a display unit via a wired or wireless connection and played back in real time. Non-limiting examples of a display unit may include a computer, tablet, television, mobile device, projector, projector screen, or any other device that can display or show an image, video, or other rendering of a conference environment.
[0037] Figure 2 is a diagram of a camera 200 including a video processing unit 210. Figure 2 As shown, the video processing unit 210 (which may include one or more trained neural networks (e.g., convolutional neural networks CNN)) can process video data from the sensor 220. In some examples, the video processing unit 210 may include the same as described above with respect to Figure 1 Similar features to those of the visual processing units 113, 123 described above and providing Figure 1 The visual processing units 113, 123 described have similar functions.
[0038] Video processing unit 210 can receive overview video stream 230 and, based on analysis of the overview video stream, can cause primary video stream 232 to be generated. In some examples, the primary video stream can include a cropped and scaled video stream based on a portion of a frame included in the overview video stream. In other examples, as discussed in the following sections, primary video stream 232 can include multiple warped subframes of the overview video stream, to which one or more image transformations have been applied, such that the primary video stream appears to be captured from a different camera perspective than the camera perspective associated with the overview video stream. Using specialized hardware and software, camera 200 can use a wide-angle lens (not shown) and / or a high-resolution sensor (such as sensor 220) to detect the location of conference participants. Furthermore, in some embodiments, camera 200 can determine who is speaking based on the head orientation of the conference participant(s), detect facial expressions, and determine where attention is focused based on the head orientation(s). This information can be sent to virtual director 240, which can determine appropriate video setting selections for the video stream(s) output.
[0039] In a video conferencing setting, one or more cameras can be positioned relative to a meeting space (e.g., a home office, a meeting room, a boardroom, a classroom, or any space from which images representing meeting participants can be captured). In limited circumstances, a video conferencing camera can be positioned in the meeting space so that the meeting participants appear at the center of the camera's field of view. In such circumstances, the primary meeting participant's line of sight relative to the camera can indicate to other meeting participants (or other recipients of an image stream including one or more images) that the primary meeting participant is looking directly into the lens of the camera at a relatively level line of sight. In other words, the central optical axis of the capture camera can be substantially perpendicular to a center point associated with the primary meeting participant (e.g., perpendicular to a center point of the participant's face, neck, torso, etc.).
[0040] However, more commonly, the camera of a video conferencing system will be positioned in the conference space such that the central optical axis of the camera is not perpendicular to the center point of the conference participants. For example, the camera is typically positioned at a certain height in the conference space (e.g., at the top or bottom of a display, computer monitor, on a desktop, etc.) at which the camera can be tilted up or down at a non-zero tilt angle in order to capture images of the conference participants. Similarly, the camera can be positioned in the conference space relative to the conference participants such that the center of the camera's field of view is offset from the center point associated with the conference participant by a positive or negative pan angle.
[0041] As a result of this off-axis camera perspective, at least some amount of distortion in the image representation of the conference participants within the acquired image frame (e.g., in the overview video stream) will typically be present in all camera perspectives except for the straight-on example (where the optical axis of the camera is substantially perpendicular to the center point of the conference participants). For example, Figure 3 A first example is shown in which a captured image 310 is taken from an elevated camera position relative to a conference participant. Due to this elevated camera perspective, distortion may occur such that features closer to the camera (e.g., a participant's head and face) appear proportionally larger than features of the conference participant further from the camera (e.g., a participant's feet).
[0042] In contrast, Figure 3 Image 320 shown in does not include the same distortion as exhibited in image 310. Instead, image 320 appears to be captured from a camera perspective, such as an on-axis camera perspective in which the camera's central optical axis is substantially aligned with a normal to a central location associated with the conference participant. In one example, image 320 can be generated by placing a video conferencing camera at an on-axis perspective position relative to the conference participant without the same distortion as exhibited in image 310.
[0043] However, in other cases, which will be described in more detail in the following sections, image 320 can be generated by image warping an image acquired from an off-axis camera perspective relative to a conference participant. Image warping can result in an image warped image with less or no distortion associated with the original off-axis camera perspective. In such an example, an acquired image (such as image 310) that exhibits distortion associated with an off-axis camera perspective (e.g., where the central optical axis of the camera is tilted downward toward a central position of a conference participant (such as the center of the participant's face, the center of the neck, the center of the torso, etc.)) can be image warped to be presented as image 320. This image warping can be achieved by determining the difference between the actual camera perspective (e.g., the downward tilt angle of the camera in the example of image 310) and the target camera perspective (e.g., an on-axis camera perspective). Upon determining this difference, an image transformation can be determined that, when applied to the acquired image (e.g., image 310), causes the transformed, image-warped image (e.g., image 320) to appear as if it were acquired at the target image perspective rather than from the actual camera perspective.
[0044] Other video conferencing systems already include the ability to highlight specific conference participants represented in a wide-angle camera view. For example, some systems may provide the ability to zoom in on an individual (e.g., a speaker) to highlight that individual in the video stream shown on the display. However, such automatic framing models typically only provide simple cropping and scaling capabilities. As a result, the cropped and scaled image will include distortions that are the same or similar to those that already existed in the original wide-angle image frame. For example, one disadvantage of this approach is that cropping and scaling, while conceptually simple, will leave behind unnecessary straight line distortions (such as straight lines appearing curved) or perspective distortions (such as volumetric deformation) that were present in the main image. This is a known trade-off in the field of panoramic photography. If the field of view of the image is large, the rendering will have to balance these distortions.
[0045] Early video conferencing camera models also lack the ability to perform changes in the acquired image to change the perceived camera perspective relative to the acquired image. That is, such systems do not provide the ability to perform image warping on a segment of the captured image to change the effective camera perspective relative to a characterized segment of the captured image (e.g., performing image warping on image 310 to provide image 320 having a perceived camera perspective (e.g., a target camera perspective) that is different from the actual camera perspective associated with acquired sensor image 310).
[0046] These features in a video conferencing system can provide specific benefits. For example, the features described herein are intended to enhance the overall meeting experience and can also help improve the efficiency of communication and collaboration. The currently disclosed embodiments can help reduce distortion. In some conference room settings, a wide-angle or fisheye lens can be used to capture a larger field of view. Any such lens will distort the image in some way, such as making people appear stretched or straight lines appear curved. Image warping can adaptively correct for this distortion, making the video delivery more natural and visually appealing, which can be especially important for professional presentations and client meetings.
[0047] Perspective can be improved or optimized. Meeting rooms come in all sizes and shapes, and not all participants can sit directly in front of the camera. Image warping allows the camera to adjust its view and perspective to ensure all participants are visible on screen, even if they're not in the camera's direct line of sight. This improves inclusivity and helps ensure everyone is visible during the meeting.
[0048] Additionally, image deformation according to the disclosed technology can dynamically adapt to the layout of the conference room. For example, if the conference room is being used for a specific purpose (such as a presentation or training session), the camera settings (including the effective camera perspective) can be adjusted (automatically or based on user input) to focus on the presenter or whiteboard, thereby ensuring that the content being discussed is clear and visible to remote participants.
[0049] Image warping according to the disclosed embodiments can also allow for more efficient use of available meeting space. Some conference rooms have limited space for camera placement. Image warping can increase or maximize the use of that space by capturing a wider field of view and presenting it in a standard screen aspect ratio without unnecessary cropping or distortion.
[0050] The disclosed image warping techniques can enhance the user experience by, for example, ensuring that remote conference participants have a clear, undistorted view of the room and its occupants. When participants can clearly see each other and interact effectively, this enhances engagement, understanding, and collaboration. Adaptive perspective (e.g., an effective camera perspective that changes over time relative to one or more conference participants) can produce more variation in the video than a static perspective. This can help reduce meeting fatigue. The disclosed techniques can also provide visually appealing and well-composed video delivery that can enhance an organization's professionalism and brand image; for example, demonstrating a commitment to the quality of communication and collaboration.
[0051] The disclosed system can simplify setup complexity. For example, image morphing can eliminate the need to physically move or adjust the camera each time the room layout changes or a new participant enters the meeting. This saves time and effort, making meetings more efficient. Furthermore, cost savings can be achieved. Instead of investing in multiple cameras or expensive motorized camera systems for different room configurations, image morphing can even achieve similar or better results with a single camera, potentially saving equipment costs.
[0052] Returning to the implementation, the disclosed embodiments can use various techniques to accomplish the described image warping of sub-frame regions of image frames summarizing a video stream. In one example, the disclosed embodiments can use a projection model to create a dynamic image warping mesh, thereby giving a new synthetic view and projection that is different from the projection of the optical lens used to capture the image. By relying on camera calibration methods that allow for mapping pixels to spatial directions, it is possible to use the model to create image warping that simulates the projection of nearly any camera geometry, as long as it is within the field of view of the original image. In other words, it enables the use of a single fixed camera to change camera angles and types.
[0053] In addition to reducing unintended distortion, choosing the right projection adapted to the scene content can improve the experience by balancing the unavoidable distortion tradeoffs and placing the distortion where it is less likely to be noticed (distortion masking). Furthermore, the concept allows transitions between projections, for example to give the appearance of physical camera motion.
[0054] Projection models can be broadly classified as non-parametric or parameterized. Several types of parameterized models can be used. These models can relate to the physical surface that the model simulates and is projected onto, such as a plane or cylindrical model. In each parameterized model, parameters can control various aspects (such as plane orientation or cylinder curvature). It is also possible to combine two parameterized models and to mix between them. For example, a model can be plane in the central field of view and cylindrical in the peripheral field of view.
[0055] On the other hand, non-parametric models do not need to have physically interpretable parameters, but can include multiple similar model variables. Examples include splines or Bezier curves that describe the pixel displacement between the input image and the output image. This is a more direct way to construct deformable meshes with high flexibility.
[0056] The adaptation of the projection can be controlled by a projection control model. This model can be designed or generated through machine learning. The projection control model takes as input features from a computer vision model and auxiliary sensor signals. Machine vision features can detect objects in the scene, estimate distances to objects, infer facial and body pose keypoints, determine facial embeddings, or estimate the geometry of a room's furniture layout, among a variety of other possible tasks.
[0057] Various auxiliary sensor signal inputs may be used to determine camera orientation. In some cases, these sensors may include accelerometers, inertial measurement units (IMUs), directional microphones, and / or lidar systems that provide range information relative to various locations associated with meeting participants, various locations associated with the meeting space (e.g., room corners, etc.), and / or objects in the meeting environment. Other examples of these sensors include, but are not limited to, sensors such as those regarding Figure 1 Sensor 130 is shown and described.
[0058] Figure 4A conceptual representation of the use of auxiliary sensor inputs in the disclosed adaptive image projection techniques is provided. For example, as shown, a camera image sensor 410 in a video conferencing system 410 can generate an initial input (e.g., an overview video stream comprising a plurality of acquired image frames). A computer vision module 420 (e.g., comprising one or more trained neural networks executed by a video processing unit (such as processing unit 113 or 123)) can receive the overview video stream from the image sensor 410. Based on an analysis of at least one test frame, the computer vision model 420 can determine, for example, regions of the test frame that include representations of meeting participants or other objects of interest. These regions of interest can be extracted from frames of the overview video stream and characterized in one or more primary video streams, which can be modified using the disclosed image deformation techniques.
[0059] System 410 may also include a projection control module 430, for example, executed by processing unit 113 or 123. The projection control module may receive input from auxiliary sensors 432 (e.g., accelerometers, directional microphones, lidar, etc.) indicating the orientation of the camera including image sensor 410. In some cases, projection control unit 430 may determine the orientation and / or actual perspective of the camera used to acquire the overview video stream based on the auxiliary sensor input. In other cases, projection control unit 430 may include one or more trained models configured to receive a test image frame or subframe as input and output one or more parameter values indicating the actual perspective / orientation of the camera.
[0060] Using the actual camera orientation / perspective information, projection model 440 (e.g., also executed by processing unit 113 or 123) can be used to determine one or more image transformations to convert an image frame or image sub-frame generated using image sensor 410 from the actual perspective to a target perspective (e.g., an on-axis perspective as described above). In some examples, the one or more image transformations can include a deformation grid that provides pixel-by-pixel transformation information to deform the acquired image frame or sub-frame so that the image frame or sub-frame appears to be acquired from the target camera perspective rather than the actual camera perspective.
[0061] The image transformation(s) generated by projection model 440 may be implemented by image warping unit 450 (e.g., performed by processing unit 113 or 123). Figure 4 As shown, the image warping unit 450 may operate directly on frames or sub-frames of the overview video stream, in particular as indicated or guided by the computer vision model 420 .
[0062] Figure 5 Another embodiment of an adaptive image warping system 500 is shown, in which the automatic framing model is extended to provide a perspective frame fitting stage. Figure 4 The automatic framing model generates framing based on input from the object detector model (e.g., in addition to other machine vision models that provide features such as facial landmarks or embedded features). The framing model may also receive additional input based on other modalities such as audio signal analysis from microphone input. The object detector or machine vision model receives input in the form of an overview image, which may be an image sensor image processed by a suitable image signal processing (ISP) device. This may be referred to as a preview image.
[0063] In this example embodiment, the output from auto-framing is a rectangular frame in some input coordinate space (e.g., image pixel coordinates of a preview image or calibrated normalized camera coordinates) 510. Here and in other disclosed embodiments, the disclosed auto-framing model is configured to include / generate perspective changes relative to the actual perspective associated with the acquired image frame.
[0064] Thus, rather than simply cropping and scaling the acquired image sensor image or the sensor image processed by an image signal processor (ISP), which, as described above, retains primarily straight line distortions (e.g., straight lines appear curved) or perspective distortions (e.g., volumetric deformation) present in the image, the disclosed embodiments apply image warping techniques to change the perceived perspective of the warped image relative to the image acquired by the image sensor.
[0065] The disclosed embodiments can fit a perspective transform to a frame in a separate step, taking a rectangular frame as input. The perspective frame boundaries can be adapted to approximate the field of view of the rectangular frame, with the perspective center close to the center of the frame or the center of the object of interest. By adapting the perspective in this way, distortion associated with cropping wide-angle images can be eliminated. The output from the perspective frame fitting step can be used in an image warping step, taking one or more overview video image frames as input.
[0066] Figure 6
[00136] An example of perspective frame fitting is shown. The rotation is first determined based on a rectangular input frame and optionally based on where the perspective center should be located, which sensibly defaults to the center of the rectangular input frame. To determine the rotation, it may be useful to first convert the coordinates of the input rectangle into normalized camera coordinates, unless they are already in that form. The vertices of the input rectangle can then be transformed to the new perspective using a rotation matrix. The vertices of the input rectangle can be, for example, corners and midpoints, or more densely sampled points. The new rectangle can then be fitted to the transformed vertices, and finally, a data structure (e.g., a target perspective frame) can be constructed with the rotated and fitted rectangles.
[0067] The translation and tilt rotations determined in perspective frame fitting can also be based on the measured camera orientation received from the auxiliary sensor. For example, the tilt axis can be determined from the measured tilt angle so that the resulting image after image deformation will be horizontal, while vertical features appear parallel.
[0068] The perspective frame can also be used to guide image scaling and cropping during the main ISP step, or before the main image signal processor step, to provide an image warping step with a desired resolution level. Figure 7 One possible implementation of the image warping step is shown, which also includes the use of such guidance with respect to image scaling.
[0069] Figure 8 A possible implementation of the mesh generation step is shown. When including perspective in the framing transformation, more variation and naturalness of motion can be created compared to using rectangular frames.
[0070] Example Processing Flow
[0071] In some embodiments, the disclosed system can be configured to perform a series of steps to convert an input image frame or subframe (e.g., from an overview video) into a deformed image frame that simulates capture from a target camera perspective that is different from the actual camera perspective used to capture the overview video stream. In some examples, the method performed by the disclosed embodiments may include: receiving one or more video frames from an image sensor by an image signal processor; determining one or more preview video frames based on the received images by the image signal processor; storing the determined one or more preview video frames to a memory; reading the stored one or more preview video frames from the memory by a neural network processing unit; determining one or more computer vision features representing a scene depicted in the one or more preview video frames (e.g., a region of interest in the captured image frame) using the neural network processing unit; determining a polygonal region of interest in one or more of the preview video frames based on the one or more computer vision features representing the scene; reading camera model parameters for a camera model from the memory; determining, by the camera model, normalized camera coordinates of vertices of the polygonal region of interest from the polygonal region of interest; determining a camera orientation (e.g., based on a vector graph of the camera model); ... input from one or more auxiliary sensors); determining one or more geometric projection parameters based on the polygonal region of interest in the normalized camera coordinates and the camera orientation; determining a polygonal region of interest that is transformed into a new geometric projection through the camera model and the one or more geometric projection parameters; determining a rectangular region of interest that approximates the polygonal region of interest based on the transformed polygonal region of interest; determining a main video frame based on the received sensor video frame by the image signal processor; storing the determined main video frame in a memory; reading the stored main video frame from the memory by an image deformation unit; generating a warped grid that correlates pixel coordinates in the output video frame with pixel coordinates in the main video frame by the camera model based on the geometric projection parameters and the rectangular region of interest; applying the generated warped grid in the image deformation unit to generate pixels of the output video frame based on pixels of the main video frame; and storing the generated main video frame in the memory.
[0072] In the example process steps described above, the camera model parameters read from the memory may be sufficient to convert pixel coordinates in one or more preview video frames into normalized camera coordinates according to the camera model. The camera model parameters may include a camera matrix and a geometric distortion model coefficient vector sufficient to convert pixel coordinates in one or more preview video frames into normalized camera coordinates according to the camera model. Furthermore, in the described steps, the normalized camera coordinates may be in a form that can be converted into azimuth-elevation angles relative to the camera body. The camera orientation received from the auxiliary sensor may include Euler angles for tilt and roll relative to a horizontal orientation. The geometric projection parameters may include Euler angles describing translation, tilt, and roll, which are 3D rotations about a camera projection center. The Euler tilt angle is determined based on the camera orientation so that the geometric projection, when used to generate a deformation mesh and applied in an image warping unit, flattens the image. The Euler tilt angle may be further defined as a maximum absolute value that will produce a high-quality output image in the image warping unit. If the determined angle is below a minimum absolute value defined by the expected accuracy of the measured camera orientation, the Euler angle may be clamped to zero. The Euler roll angle can be determined based on the camera orientation so that the geometric projection will flatten the image when used to generate the deformation mesh and applied in the image warping unit. The Euler roll angle can be further constrained to a maximum absolute value that will produce a high-quality output image in the image warping unit. If the determined angle is below a minimum absolute value defined by the expected accuracy of the measured camera orientation, the Euler roll angle can be clamped to zero. The Euler translation angle can be determined based on the center point of the polygonal region of interest.
[0073] In the disclosed embodiments, the image warping unit may include a configurable resampling hardware accelerator in a system on a chip (SoC), a programmable graphics processing unit (GPU), or a software component running on a general-purpose processor (CPU).
[0074] In the disclosed embodiments, the perspective frame fitting stage need not be limited to rectangular frames as input. For example, it may be beneficial to share other information from the automatic framing step (e.g., the locations of salient features that should not be cropped, and features outside the current region of interest that should not be included) to guide the perspective frame fitting. The perspective frame fitting stage need not be limited to perspective transforms and frames as its outputs, but can be generalized to include other projection model parameters (such as cylindrical projection curvature).
[0075] In other variations of the disclosed embodiments, the perspective frame can be directly adapted to the machine vision model output, thereby eliminating the intermediate rectangular frame output from the automatic framing. Such a unitary projection control model can, for example, include a neural network trained by machine learning. However, one advantage of retaining the intermediate step of having a rectangular frame output can include modularity and simpler integration with existing automatic framing implementations.
[0076] As mentioned above, the disclosed embodiments are compatible with multi-camera setups. In some cases, the director component may be responsible for selecting between framings proposed by the framing components running the various cameras.
[0077] Figure 9 Image subframes extracted from an overview video stream are shown, and the images are warped to change the perceived camera perspective from the actual perspective to the target perspective. In this example, the subjects are seated across from each other at a wide conference room table. The camera is located in an elevated position in the conference room, with its central optical axis extending parallel to the longitudinal axis of the conference room table. This configuration results in an actual camera perspective with respect to subject 910 that includes a tilt angle of -20 degrees (downward) and a pan angle of -10 degrees (leftward). Similarly, this configuration results in an actual camera perspective with respect to subject 920 that includes a tilt angle of -20 degrees (downward) and a pan angle of +10 degrees (rightward). With this configuration, the representation of the corresponding object in the resulting image subframe includes distortion caused by the off-axis camera perspective (and camera optics). However, using the disclosed image warping techniques, the image subframes can be warped according to the target camera perspective, reducing or eliminating some or all of the distortion in the originally captured subframes.
[0078] Additional Example Embodiments
[0079] The disclosed embodiments may include a video conferencing system for adjusting a perspective view using adaptive image warping. The video conferencing system may include an image warping unit comprising at least one processor programmed to: receive an overview video stream from a camera in the video conferencing system; determine at least one region of interest represented within the at least one test frame based on an analysis of at least one test frame from the overview video stream; determine one or more indicators of an actual camera perspective relative to the at least one region of interest; determine a target camera perspective relative to the at least one region of interest, wherein the target camera perspective is different from the actual camera perspective; determine at least one image transformation based on the difference between the actual camera perspective and the target camera perspective; apply the at least one image transformation to one or more subframe regions of a plurality of image frames of the overview stream to generate at least one image-warped primary video stream; and cause the at least one image-warped primary video stream to be displayed on a display.
[0080] Various types of image segments in a test frame of the overview video stream can be identified (e.g., by a trained neural network) as comprising representations of regions of interest in the conference environment. In addition, in some cases, multiple discrete regions of interest can be identified based on analysis of a single test frame. The determined one or more regions of interest can be used to determine how sub-frames can be extracted from one or more complete frames of the overview video stream for generating a main video stream (e.g., focusing on image representations of conference participants represented in sub-frame regions of the overview video frames). For example, once regions of interest in the conference environment have been identified based on analysis of one or more test frames, image segments of the overview video that include representations of the regions of interest can be used to generate the main video stream.
[0081] In some cases, as noted, the regions of interest of the conference environment identified based on the analysis of the test frames may, for example, include video conference participants (e.g., areas where conference participants sit or stand). The at least one identified region of interest may also include a first region of interest for a first video conference participant and at least a second region of interest that includes a second video conference participant. In some cases, at least one region of interest includes two or more video conference participants (e.g., where two participants are close to each other such that a primary video stream is generated to include representations of more than one participant within a single frame of the primary video stream). Additionally, at least one region of interest may be identified based on its inclusion of one or more objects (e.g., a podium, a lectern, a presentation screen, etc.).
[0082] The identified regions of interest can be defined by any suitable boundary indicator. In some cases, at least one region of interest is delineated by a rectangular boundary within at least one test frame. At least one region of interest can also be delineated by a polygonal boundary within at least one test frame. The polygonal boundary can be a quadrilateral. In other cases, at least one region of interest is delineated by a boundary in at least one test frame that traces the outline of a perimeter associated with the representation of at least one video conference participant.
[0083] Any suitable convention may be used to express the one or more indicators of the actual camera perspective relative to the at least one region of interest. For example, in some cases, the one or more indicators of the actual camera perspective relative to the at least one region of interest include coordinates in the camera's reference frame of at least one point associated with a determined boundary that delineates the representation of the at least one region of interest in the at least one test frame. The one or more indicators may also include coordinates in the camera's reference frame of at least one point associated with an object or video conference participant located in the at least one region of interest. Similarly, the one or more indicators may include coordinates in the camera's reference frame of each of a plurality of pixels included in the representation of the at least one region of interest.
[0084] Various techniques can be used to determine one or more indicators of the actual camera perspective relative to at least one region of interest. For example, the one or more indicators can be determined based at least in part on the output of a sensor (e.g., an auxiliary sensor) that is integrated with or separate from the camera. In some cases, the sensor includes an accelerometer or a directional microphone. In some examples, the one or more indicators of the actual camera perspective relative to at least one region of interest are determined based at least in part on a predetermined camera model, wherein the predetermined camera model indicates at least one of a field of view angle, a pitch value, a tilt value, a roll value, a yaw value, or a translation value associated with the camera.
[0085] A target camera perspective that is different from the actual camera perspective used to obtain the overview video stream can be determined in various ways. In some examples, the target camera perspective is determined to have an opposite pan and / or tilt angle relative to the actual camera perspective. In some examples, the target camera perspective can be determined as a line of sight from a camera origin to a center point associated with an object or video meeting participant represented in at least one test frame.
[0086] In other cases, a plurality of different target camera perspectives may be determined, and the different target perspectives may be used as a basis for simulating a changed camera perspective relative to a conference participant or region of interest. In such cases, the target camera perspective would include a plurality of different target camera perspectives, each associated with one or more corresponding image transformations. To provide the changed perspective effect in the generated primary video stream, the one or more corresponding image transformations may be applied to one or more sub-frame regions of at least one of the plurality of image frames of the overview video stream to generate at least one image-warped primary video stream representing the changed camera perspective.
[0087] The simulated change in camera perspective can occur at a variety of different rates. For example, each of the one or more corresponding image transforms can be applied to the same number of frames in the multiple image frames of the overview video stream. Such application can result in the perspective appearing to change at a constant rate (e.g., linearly). In other cases, a nonlinear perspective change effect can be achieved by varying how the image transforms are applied. For example, each of the one or more corresponding image transforms can be applied to a linearly varying number of frames in the multiple image frames of the overview video stream. In this case, the perspective change effect may appear to accelerate, but at a constant acceleration rate. In other cases, each of the one or more corresponding image transforms can be applied to a nonlinearly varying number of frames in the multiple image frames of the overview video stream, so that the effect achieved is that the change in camera perspective appears to start slowly (or quickly), but then accelerate (or decelerate) toward the final target camera perspective.
[0088] When the target camera perspective(s) are determined, at least one image transformation can be determined based on the difference between the actual camera perspective and the target camera perspective. The image transformation can, for example, be suitable for deforming original image segments acquired at the actual camera perspective into corrected image segments that appear to be acquired from the target camera perspective. In some cases, the image transformation can indicate one or more image adjustments on a pixel-by-pixel basis that depend on the difference between the translation angle of a first camera associated with the actual camera perspective and the translation angle of a second camera associated with the target camera perspective. Additionally, the image transformation can indicate one or more image adjustments on a pixel-by-pixel basis that depend on the difference between the tilt angle of a first camera associated with the actual camera perspective and the tilt angle of a second camera associated with the target camera perspective.
[0089] To obtain a simulated acquisition of a warped image from the perspective of the target camera, at least one image transformation can be applied to one or more sub-frame regions of a plurality of image frames of the overview stream to generate at least one warped primary video stream. The application of the at least one image transformation can be accomplished using an image warping grid (as described above) that indicates, for a plurality of pixel coordinates of the image warping grid, the one or more transformations to be applied relative to pixel coordinates of the overview video stream.
[0090] Where one or more warped primary video streams are generated, the primary video streams can be displayed on a display. The display can include, among other examples, a boardroom display, a desktop computer display, a laptop or mobile device display, etc. Alternatively or additionally, the warped primary video stream(s) can be provided to the video conferencing platform for display as part of the receiving platform's video feed.
[0091] The various analysis techniques described above can be accomplished using algorithmic image analysis. However, in some examples, one or more trained neural networks k can be configured to perform one or more of the described tasks. For example, in some cases, at least one processor (e.g., included in video processing units 113, 123) can also be programmed to provide at least one trained neural network configured to receive as input at least one test frame and an indicator of a target camera perspective, and in response, output at least one image transformation for generating an image-warped primary video stream. Furthermore, the at least one trained neural network can be configured to receive as input at least one test frame, determine at least one region of interest represented in the test frame, and output one or more indicators of an actual camera perspective. The at least one trained neural network can be configured to receive as input at least one test frame and an indicator of a target camera perspective, determine at least one region of interest represented in the test frame, determine one or more indicators of an actual camera perspective, determine at least one image transformation, and output at least one image-warped primary video stream. Alternatively, at least one trained neural network may be configured to receive as input at least one test frame, determine at least one region of interest represented in the test frame, determine one or more indicators of an actual camera perspective, determine a target camera perspective, determine at least one image transformation, and output at least one image-warped primary video stream.
[0092] It should also be noted that the guidance of one or more of the above tasks can be based on input received from the user of the video conferencing system. For example, the target camera perspective and / or the expected perspective change effect (e.g., constant rate change, non-linear rate change, etc.) can be determined based on the input received from the user of the video conferencing system.
[0093] The configuration of the disclosed video conferencing system may also vary. For example, in some cases, the image warping unit (e.g., included as part of the video processing unit 113, 123) may be located on the camera or at a location (or in a device) remote from the camera. Furthermore, as noted, the disclosed video conferencing system may include a single camera or may include multiple cameras, with one of the multiple cameras being used to obtain the described output video stream.
[0094] Embodiments of the present disclosure may provide a multi-camera video conferencing system or a non-transitory computer-readable medium containing instructions for real-time image correction using adaptive image deformation. Some embodiments may involve a machine language vision / audio pipeline that can detect people, objects, voices, movements, gestures, canvas enhancements, documents, and depth in the video conferencing space. In some embodiments, a virtual director unit (or component) may use machine language vision / audio and previous events in the video conferencing to determine the specific portion of the image or video output (from one or more cameras) to be placed in the composite video stream. The virtual director unit (or component) may determine the specific layout of the composite video stream.
[0095] Although many of the disclosed embodiments are described in the context of camera systems, video conferencing systems, etc., it should be understood that the present disclosure specifically contemplates corresponding methods associated with all disclosed embodiments. More specifically, methods corresponding to actions, steps, or operations performed by a video processing unit, as described herein, are disclosed. Therefore, the present disclosure discloses a video processing method performed by at least one video processing unit, which includes any or all steps or operations performed by (one or more) video processing units as disclosed herein. In addition, at least one (or one or more) video processing units are disclosed herein. Therefore, it is specifically contemplated that protection may be claimed for at least one video processing unit in any configuration as disclosed herein. (One or more) video processing units can be defined separately and independently of other hardware components of (one or more) cameras or video conferencing systems. Also disclosed herein are one or more computer-readable media storing instructions that, when executed by one or more video processing units, cause the one or more video processing units to perform a method according to the present disclosure (e.g., any or all steps or operations performed by the video processing units, as described herein).
[0096] Other embodiments will be apparent from consideration of the specification and practice of the embodiments disclosed herein. It is intended that the specification and examples be considered as exemplary only, with a true scope and spirit of the disclosed embodiments being indicated by the following claims.
Claims
1. A video conferencing system for adjusting a perspective view using adaptive image deformation, the system comprising: An image deformation unit comprising at least one processor programmed to: receiving an overview video stream from a camera in the video conferencing system; determining, based on an analysis of at least one test frame from the overview video stream, at least one region of interest represented within the at least one test frame; determining one or more indicators of an actual camera perspective relative to the at least one region of interest; determining a target camera perspective relative to the at least one region of interest, wherein the target camera perspective is different from the actual camera perspective; determining at least one image transformation based on a difference between the actual camera perspective and the target camera perspective; applying the at least one image transformation to one or more sub-frame regions of a plurality of image frames of the overview stream to generate at least one image-warped primary video stream; and The at least one warped primary video stream is caused to be displayed on a display.
2. The video conferencing system according to claim 1, wherein: The at least one region of interest includes a video conference participant.
3. The video conferencing system according to claim 1, wherein: The at least one region of interest includes: a first region of interest including a first video conference participant and at least a second region of interest including a second video conference participant.
4. The video conferencing system according to claim 1, wherein: The at least one region of interest includes two or more video conference participants.
5. The video conferencing system according to claim 1, wherein: The at least one region of interest includes one or more objects.
6. The video conferencing system according to claim 1, wherein: The at least one region of interest represented by the at least one test frame is determined by a trained neural network.
7. The video conferencing system according to claim 1, wherein: The at least one region of interest is delineated by a rectangular boundary within the at least one test frame.
8. The video conferencing system according to claim 1, wherein: The at least one region of interest is delineated by a polygonal boundary within the at least one test frame.
9. The video conferencing system according to claim 8, wherein: The polygonal boundary is a quadrilateral.
10. The video conferencing system according to claim 1, wherein: The at least one region of interest is delineated by a boundary in the at least one test frame that traces an outline of a perimeter associated with a representation of at least one video meeting participant.
11. The video conferencing system according to claim 1, wherein: The one or more indicators of the actual camera perspective relative to the at least one region of interest include coordinates in the reference system of the camera of at least one point associated with a determined boundary, wherein the determined boundary delineates the representation of the at least one region of interest in the at least one test frame.
12. The video conferencing system according to claim 1, wherein: The one or more indicators of the actual camera perspective relative to the at least one region of interest include coordinates in the camera's reference frame of at least one point associated with an object or video conference participant located in the at least one region of interest.
13. The video conferencing system according to claim 1, wherein: The one or more indicators relative to the actual camera perspective of the at least one region of interest include coordinates of each pixel of a plurality of pixels included in the representation of the at least one region of interest in a reference system of the camera.
14. The video conferencing system according to claim 1, wherein: The one or more indicators of an actual camera perspective relative to the at least one region of interest are determined based at least in part on output of a sensor separate from the camera.
15. The video conferencing system according to claim 14, wherein: The sensor includes an accelerometer.
16. The video conferencing system according to claim 14, wherein: The sensor includes a directional microphone.
17. The video conferencing system according to claim 1, wherein: The one or more indicators of an actual camera perspective relative to the at least one region of interest are determined at least in part based on a predetermined camera model.
18. The video conferencing system according to claim 17, wherein: The predetermined camera model indicates at least one of a field of view angle, a pitch value, a tilt value, a roll value, a yaw value, or a translation value associated with the camera.
19. The video conferencing system according to claim 1, wherein: The image transformation indicates one or more image adjustments on a pixel-by-pixel basis that depend on a difference between a translation angle of a first camera associated with the actual camera perspective and a translation angle of a second camera associated with the target camera perspective.
20. The video conferencing system according to claim 1, wherein: The image transformation refers to one or more image adjustments on a pixel-by-pixel basis that depend on a difference between a tilt angle of a first camera associated with the actual camera perspective and a tilt angle of a second camera associated with the target camera perspective.
21. The video conferencing system according to claim 1, wherein: The target camera perspective includes a plurality of different target camera perspectives, each target camera perspective being associated with one or more corresponding image transformations.
22. The video conferencing system according to claim 21, wherein: Respective ones of the one or more respective image transforms are applied to the one or more sub-frame regions of at least one image frame of the plurality of image frames of the overview video stream to generate at least one image-warped primary video stream representing the changed camera perspective.
23. The video conferencing system according to claim 22, wherein: Each respective one of the one or more respective image transforms is applied to a same number of frames of the plurality of image frames of the overview video stream.
24. The video conferencing system according to claim 22, wherein: Each respective one of the one or more respective image transforms is applied to a linearly varying number of frames of the plurality of image frames of the overview video stream.
25. The video conferencing system according to claim 22, wherein: Each respective one of the one or more respective image transforms is applied to a non-linearly varying number of frames in the plurality of image frames of the overview video stream.
26. The video conferencing system according to claim 1, wherein: The at least one processor is further programmed to provide at least one trained neural network configured to receive as input the at least one test frame and the indicator of the target camera perspective, and to output the at least one image transformation.
27. The video conferencing system according to claim 1, wherein: The at least one processor is further programmed to provide at least one trained neural network configured to receive the at least one test frame as input, determine the at least one region of interest represented in the test frame, and output the one or more indicators of the actual camera perspective.
28. The video conferencing system according to claim 1, wherein: The at least one processor is further programmed to provide at least one trained neural network configured to receive as input the at least one test frame and the indicator of the target camera perspective, determine the at least one region of interest represented in the test frame, determine the one or more indicators of actual camera perspective, determine the at least one image transformation, and output the at least one image-warped primary video stream.
29. The video conferencing system according to claim 1, wherein: The at least one processor is further programmed to provide at least one trained neural network configured to receive the at least one test frame as input, determine the at least one region of interest represented in the test frame, determine the one or more indicators of an actual camera perspective, determine the target camera perspective, determine the at least one image transformation, and output the at least one image-warped primary video stream.
30. The video conferencing system according to claim 1, wherein: The target camera perspective is along a line substantially perpendicular to a center point associated with an object or video meeting participant represented in the at least one test frame.
31. The video conferencing system according to claim 1, wherein: The target camera perspective is determined based on input received from a user of the video conferencing system.
32. The video conferencing system according to claim 1, wherein: Applying the at least one image transformation is done using an image deformation grid that indicates, for a plurality of pixel coordinates of the image deformation grid, one or more transformations to be applied relative to pixel coordinates of the overview video stream.
33. The video conferencing system according to claim 1, wherein: The image deformation unit is positioned on the camera.
34. The video conferencing system according to claim 1, wherein: The image deformation unit is remotely located relative to the camera.
35. The video conferencing system according to claim 1, wherein: The video conferencing system is a multi-camera video conferencing system including a plurality of cameras, and the camera is included in the plurality of cameras.
Citation Information
Cited By
Screen fault detection method and device, electronic equipment and readable storage medium
CN121354446A