Restoration of image fov for stereoscopic rendering

By selecting and correcting frames in the camera array and inserting patches into blank areas, the FOV loss problem caused by parallel camera arrays is solved, and FOV restoration and quality improvement of stereoscopic rendered images are achieved.

CN115997379BActive Publication Date: 2025-12-16SAMSUNG ELECTRONICS CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202180052089.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2021-08-20
Filing Date
2021-08-25
Publication Date
2025-12-16
Estimated Expiration
2041-08-25

AI Technical Summary

Technical Problem

Images or videos captured using parallel camera arrays result in a significant loss of the available field of view (FOV) in 3D autostereoscopic displays, leading to suboptimal content.

Method used

By selecting the first and second frames, correcting and aligning to the reference frame, and transforming the first frame to have near-optimal overlap with the second frame, a patch from the transformed first frame is inserted into the blank area of ​​the second frame to restore the usable field of view (FOV).

Benefits of technology

It restores the FOV loss caused by the parallel camera array setup, improving the image quality and usable field of view in stereo rendering.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115997379B_ABST
    Figure CN115997379B_ABST
Patent Text Reader

Abstract

An apparatus includes a memory and a processor. The memory receives a plurality of frames of a scene captured from a camera array. The processor selects a first frame and a second frame from the plurality of frames. The processor also rectifies and aligns the first frame and the second frame to a reference frame, where a blank area of the second frame has a greater area than a blank area of the first frame. The processor also transforms the first frame to have near-optimal overlap with the second frame. The processor inserts a patch in the transformed first frame into the blank area of the second frame.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure relates generally to image processing devices and processes. More specifically, the present disclosure relates to methods and apparatus for recovering image fields of view (FOV) captured via a multi-view camera rig setup for stereoscopic rendering. BACKGROUND

[0002] One-dimensional (ID) or two-dimensional (2D) parallel camera arrays are a common way of capturing multi-view and light field video. The captured frames need to be transformed so that they are visible in a three-dimensional (3D) autostereoscopic display. However, the transformation of the images or video can result in a significant loss of available FOV for the multi-view video or light field content. Such loss of FOV results in suboptimal content. The techniques described in the present disclosure aim to recover the available FOV for images or video captured using a parallel camera array setup. SUMMARY

[0003] The present disclosure provides methods and apparatus for recovering image FOV captured via a parallel camera setup for stereoscopic rendering. BRIEF DESCRIPTION OF DRAWINGS

[0004] For a more complete understanding of the present disclosure and its advantages, reference is now made to the following description taken in conjunction with the accompanying drawings in which like parts are marked with like numerals throughout the drawings:

[0005] Figure 1 An example communication system is shown in accordance with embodiments of the present disclosure;

[0006] Figure 2 And Figure 3 An example electronic device is shown in accordance with embodiments of the present disclosure;

[0007] Figure 4 An example end-to-end pipeline for a stereoscopic rendering system using a camera array and a display is shown in accordance with the present disclosure;

[0008] Figure 5A And Figure 5B Example available FOVs from a first camera and a second camera are shown in accordance with the present disclosure.

[0009] Figures 6A to 6F Example FOV recovery for an image array is shown in accordance with the present disclosure;

[0010] Figures 7A to 7D Example hierarchical FOV recovery is shown in accordance with the present disclosure;

[0011] Figure 8A And Figure 8B An example method for hierarchical FOV recovery is shown in accordance with the present disclosure; and

[0012] Figure 9An example method for rectifying FOV of images captured via a parallel camera setup for stereo rendering is shown in accordance with the present disclosure. DETAILED DESCRIPTION

[0013] The present disclosure provides methods and apparatus for rectifying FOV of images captured via a parallel camera setup for stereo rendering.

[0014] In a first embodiment, an apparatus includes at least one memory and at least one processor operatively coupled to the memory. The at least one memory is configured to receive a plurality of frames of a scene captured from an array of cameras. The at least one processor is configured to select a first frame and a second frame from the plurality of frames. The at least one processor is further configured to rectify and align the first frame and the second frame to a reference frame, wherein a blank region of the second frame has a larger area than a blank region of the first frame. The at least one processor is further configured to transform the first frame to have near-optimal overlap with the second frame in an overlapping region of the FOV. In perfect overlap, every point (feature point at any depth) from the first frame would have the same pixel coordinates as the corresponding point in the second frame. However, since the two frames belong to physically separated cameras, we can not find a 2D to 2D transformation of the first frame such that the transformed first frame can perfectly overlap on the second frame. Thus, we find a near-optimal overlap between the two frames such that all feature points originating from a plane at a certain depth in the scene (most commonly, a depth plane corresponding to the convergence plane) overlap with the corresponding feature points in the second frame. Further, the at least one processor is configured to insert a patch from the transformed first frame into the blank region of the second frame.

[0015] In a second embodiment, a method includes receiving a plurality of frames of a scene captured from an array of cameras; and selecting a first frame and a second frame from the plurality of frames. The method further includes rectifying and aligning the first frame and the second frame to a reference frame, wherein a blank region of the second frame has a larger area than a blank region of the first frame. The method further includes transforming the first frame to have near-optimal overlap with the second frame. Additionally, the method includes inserting a patch from the transformed first frame into the blank region of the second frame.

[0016] In a third embodiment, a non-transitory machine-readable medium stores instructions that, when executed, cause a processor to receive a plurality of frames of a scene captured from an array of cameras; and select a first frame and a second frame from the plurality of frames. The instructions, when executed, further cause the processor to rectify and align the first frame and the second frame to a reference frame, wherein a blank region of the second frame has a larger area than a blank region of the first frame. The instructions, when executed, further cause the processor to transform the first frame to have near-optimal overlap with the second frame. Additionally, the instructions, when executed, cause the processor to insert a patch from the transformed first frame into the blank region of the second frame.

[0017] Other technical features can be readily apparent to one skilled in the art from the following figures, descriptions, and claims.

[0018] Before undertaking the detailed description below, it can be advantageous to set forth definitions of certain terms and phrases used throughout this patent document. The term "couple" and its derivatives refer to any direct or indirect communication between two or more elements, regardless of the nature of the The terms "transmit," "receive," and "communicate," and derivatives thereof, encompass both direct and indirect communication. The terms "include" and "comprise," as well as derivatives thereof, mean inclusion without limitation. The term "or" is inclusive, meaning and / or. The phrase "associated with," as well as derivatives thereof, means to include, be included within, interconnect with, contain, be contained within, connect to or with, couple to or with, be communicable with, cooperate with, interleave, juxtapose, be proximate to, be bound to or with, have a property of, have relations with, have agreements, or other related terms, and the like. The term "controller" means any device, system or part thereof that controls at least one operation. Such a controller can be implemented in hardware or a combination of hardware and software and / or firmware. The functionality associated with any particular controller can be centralized or distributed, whether locally or remotely. The phrase "at least one of," when used with a list of items, means that different combinations of one or more of the listed items can be used and only one item from the list can be needed. For example, "at least one of A, B, and C" includes: A alone, B alone, C alone, A and B together, A and C together, B and C together, and A and B and C together.

[0019] Furthermore, various functions described below can be implemented or supported by one or more computer programs, each of which is formed from computer readable program code and embodied in a computer readable medium. The terms "application" and "program" refer to one or more computer programs, software components, sets of instructions, procedures, functions, objects, classes, instances, related data, or a portion thereof that perform particular tasks or implement particular abstract data types. The phrase "computer readable medium" includes any mechanism that stores, communicates, propagates, or transports a program for use by or in connection with an information processing system, such as a computer readable storage medium. The computer readable storage medium can be, for example, but is not limited to, an electronic storage device, a magnetic storage device, an optical storage device, a tape, a floppy disk, a CD-ROM, a DVD, a memory stick, or a hard disk. The computer readable storage medium can be non-transitory. The phrase "non-transitory" includes a tangible medium that can store data for a period of time (e.g., a hard disk). The phrase "non-transitory" excludes wired, wireless, optical, or other communication links that transport a program for use by or in connection with an information processing system. The computer readable storage medium includes a tangible medium that is non-transitory, and a medium that can store data for a period of time (e.g., a memory stick). The computer readable medium can be non-transitory.

[0020] Definitions for other certain words and phrases are provided throughout this patent document. Those of ordinary skill in the art will understand that such definitions apply not only to the respective words and phrases that they follow, but to instances thereof throughout this document.

[0021] The description that follows describes examples only. Figures 1 to 9 The principles of the disclosure described herein can be employed in any type of suitably arranged device or system.

[0022] Figure 1 An example communication system 100 according to embodiments of the disclosure is shown. Figure 1 The illustrated embodiment of the communication system 100 is for illustration only. Other embodiments of the communication system 100 can be used without departing from the scope of the disclosure.

[0023] The communication system 100 includes a network 102 that facilitates communication between various components in the communication system 100. For example, the network 102 can communicate IP packets, frame relay frames, Asynchronous Transfer Mode (ATM) cells, or other information between network addresses. The network 102 includes one or more local area networks (LANs), metropolitan area networks (MANs), wide area networks (WANs), all or a portion of a global network such as the Internet, or one or more other communication systems at one or more locations.

[0024] In this example, the network 102 facilitates communication between a server 104 and various client devices 106-116. The client devices 106-116 may, for example, be smartphones, tablet computers, laptop computers, personal computers, wearable devices, or HMDs. The server 104 can represent one or more servers. Each server 104 includes any suitable computing or processing device that can provide computing services to one or more client devices, such as the client devices 106-116. Each server 104 may, for example, include one or more processing devices, one or more memories that store instructions and data, and one or more network interfaces that facilitate communication over the network 102. As described in greater detail below, the server 104 can send to one or more display devices, such as the client devices 106-116, a compressed bitstream that includes one or more FOV restored frames captured from a linear camera array. In certain embodiments, each server 104 can include an encoder.

[0025] Each client device 106-116 represents any suitable computing or processing device that interacts with at least one server, such as server 104, or other computing devices over network 102. Client devices 106-116 include a desktop computer 106, a mobile telephone or mobile device 108, such as a smartphone, a PDA 110, a laptop computer 112, a tablet computer 114, and a HMD 116. However, any other or additional client devices can be used in communication system 100. A smartphone represents a class of mobile devices 108 that are handheld devices with a mobile operating system and integrated mobile broadband cellular network connectivity for voice, short message service (SMS), and Internet data communications. A 3D display can display stereoscopic images including one or more stereoscopic rendered images. In certain embodiments, any client device 106-116 can include an encoder, a decoder, or both. For example, mobile device 108 can receive multiple frames from a linear camera array and then stereoscopic render the multiple frames to send to one of client devices 106-116.

[0026] In this example, some client devices 108-116 communicate indirectly with network 102. For example, electronic device 108 and PDA 110 communicate via one or more base stations 118, such as cellular base stations or eNodeBs (eNBs). Also, laptop computer 112, tablet computer 114, and HMD 116 communicate via one or more wireless access points 120, such as IEEE 802.11 wireless access points. It should be noted that these are for illustration only and each client device 106-116 can communicate directly with network 102 or indirectly with network 102 via any suitable intermediate device or network. In certain embodiments, server 104 or any client device 106-116 can be used to rectify and align multiple frames, transform each frame to an adjacent frame, insert patches from the transformed frame into the adjacent frame, and send a bitstream including the rectified multiple frames to another client device, such as any client device 106-116.

[0027] In certain embodiments, any client device 106-114 securely and efficiently transmits information to another device, such as server 104. Also, any client device 106-116 can trigger information transmission between itself and server 104. Any client device 106-114 can act as a VR display when attached to a headset via a stand and similarly function as HMD 116. For example, when mobile device 108 is attached to a stand system and worn over a user's eyes, its functionality can be similar to HMD 116. Mobile device 108 (or any other client device 106-116) can trigger information transmission between itself and server 104.

[0028] In certain embodiments, any of the client devices 106-116 or the server 104 can create stereoscopic frames, compress stereoscopic frames, transmit stereoscopic frames, receive stereoscopic frames, render stereoscopic frames, or a combination thereof. For example, the server 104 can then compress the stereoscopic frames to generate a bitstream, and then transmit the bitstream to one or more of the client devices 106-116. In another example, one of the client devices 106-116 can compress the stereoscopic frames to generate a bitstream, and then transmit the bitstream to another one of the client devices 106-116 or to the server 104.

[0029] Although Figure 1 one example of a communication system 100 is shown, various changes can be made to Figure 1 the system. For example, the communication system 100 could include any number of each component in any suitable arrangement. In general, computing and communication systems Figure 1 have a wide variety of configurations, and Figure 1 Although

[0030] Figure 2 and Figure 3 An example electronic device according to embodiments of the disclosure is shown. In particular, Figure 2 An example server 200 is shown, and the server 200 can be representative of the server 104 in Figure 1 The server 200 can be representative of one or more encoders, decoders, local servers, remote servers, cluster computers, and components that act as a single pool of seamless resources, cloud-based servers, and the like. The server 200 can be accessed by one or more of the client devices 106-116 of Figure 1 or another server.

[0031] As Figure 2 shown, the server 200 includes a bus system 205 that supports communication between at least one processing device (such as a processor 210), at least one storage device 215, at least one communication interface 220, and at least one input / output (I / O) unit 225. The server 200 can represent one or more local servers, one or more compression servers, or one or more encoding servers, such as an encoder. In certain embodiments, the encoder can perform decoding.

[0032] Processor 210 executes instructions that can be stored in memory 230. Processor 210 may include any suitable number and type of processors or other devices arranged in any suitable manner. Exemplary types of processor 210 include microprocessors, microcontrollers, digital signal processors, field-programmable gate arrays, application-specific integrated circuits, and discrete circuits. In some embodiments, processor 210 may encode stereo frames stored in storage device 215. In some embodiments, encoding and decoding of stereo frames are performed to ensure that when stereo frames are reconstructed, they match the stereo frames before encoding.

[0033] Memory 230 and permanent storage device 235 are examples of storage devices 215 representing any structure capable of storing and facilitating the retrieval of information (such as data, program code, or other suitable information on a temporary or permanent basis). Memory 230 may represent random access memory or any other suitable volatile or non-volatile storage device. For example, instructions stored in memory 230 may include instructions for receiving multiple frames of a scene captured from a linear camera array; instructions for selecting a first frame and a second frame from the multiple frames; instructions for correcting and aligning the first and second frames to a reference frame, wherein a blank area of ​​the second frame has a larger area than a blank area of ​​the first frame; transforming the first frame to have near-optimal overlap with the second frame; and inserting a patch from the transformed first frame into a blank area of ​​the second frame. Permanent storage device 235 may contain one or more components or devices supporting longer-term data storage, such as read-only memory, hard disk drive, flash memory, or optical disk.

[0034] Communication interface 220 supports communication with other systems or devices. For example, communication interface 220 may include features that facilitate communication with other systems or devices. Figure 1 The communication interface 220 is a network interface card or wireless transceiver for communication on network 102. The communication interface 220 can support communication over any suitable physical or wireless communication link. For example, the communication interface 220 can transmit a bit stream containing stereo frames to another device, such as one of client devices 106 to 116.

[0035] I / O unit 225 allows for data input and output. For example, I / O unit 225 can provide connectivity for user input via a keyboard, mouse, keypad, touchscreen, or other suitable input device. I / O unit 225 can also send output to a display, printer, or other suitable output device. However, it should be noted that I / O unit 225 can be omitted, such as when I / O interaction with server 200 occurs via a network connection.

[0036] It should be noted that, although Figure 2 Described as representing Figure 1The server 104, however, can have the same or similar architecture used in one or more of various client devices 106 to 116. For example, a desktop computer 106 or a laptop computer 112 may have the same architecture as the server 104. Figure 2 The structures shown are the same or similar.

[0037] Figure 3 An example electronic device 300 is shown, and electronic device 300 can represent Figure 1 One or more of client devices 106 to 116 are included. Electronic device 300 may be a mobile communication device, such as a mobile station, user station, wireless terminal, or desktop computer (similar to...). Figure 1 Desktop computers 106), portable electronic devices (similar to) Figure 1 Mobile devices 108, PDAs 110, laptops 112, tablets 114, or HMDs 116, etc. In some embodiments, Figure 1 One or more of the client devices 106 to 116 may contain the same or similar configuration as electronic device 300. In some embodiments, electronic device 300 is an encoder, decoder, or both. For example, electronic device 300 can be used for data transmission, image or video compression, image or video decompression, encoding, decoding, and media rendering applications.

[0038] like Figure 3 As shown, electronic device 300 includes an antenna 305, a radio frequency (RF) transceiver 310, a transmit (TX) processing circuitry 315, a microphone 320, and a receive (RX) processing circuitry 325. The RF transceiver 310 may include, for example, an RF transceiver, a Bluetooth transceiver, a Wi-Fi transceiver, a ZigBee transceiver, an infrared transceiver, and various other wireless communication signals. Electronic device 300 also includes a speaker 330, a processor 340, an input / output (I / O) interface (IF) 345, an input terminal 350, a display 355, a memory 360, and a sensor 365. The memory 360 includes an operating system (OS) 361 and one or more application programs 362.

[0039] The RF transceiver 310 receives, from the antennas 305, incoming RF signals transmitted by access points, such as base stations, WI FI routers or Bluetooth devices, or other devices of a network 102, such as a WI-FI, Bluetooth, cellular, 5G, LTE, LTE-A, Wi MAX, or any other type of wireless network. The RF transceiver 310 down-converts the incoming RF signals to generate intermediate frequency or baseband signals. The intermediate frequency or baseband signals are sent to the RX processing circuitry 325, which generates processed baseband signals by filtering, decoding, and / or digitizing the baseband or intermediate frequency signals. The RX processing circuitry 325 transmits the processed baseband signals to the speaker 330, such as for voice data, or to the processor 340 for further processing, such as for web browsing data.

[0040] The TX processing circuitry 315 receives analog or digital voice data from the microphone 320 or other outgoing baseband data from the processor 340. The outgoing baseband data can include web browsing data, e-mail, or interactive video game data. The TX processing circuitry 315 encodes, multiplexes, and / or digitizes the outgoing baseband data to generate processed baseband or intermediate frequency signals. The RF transceiver 310 receives the outgoing processed baseband or intermediate frequency signals from the TX processing circuitry 315 and up-converts the baseband or intermediate frequency signals to RF signals for transmission via the antennas 305.

[0041] The processor 340 can include one or more processors or other processing devices. The processor 340 can execute instructions stored in the memory 360, such as OS 361, to control the overall operation of the electronic device 300. For example, the processor 340 can control the reception of forward channel signals and the transmission of reverse channel signals by the RF transceiver 310, the RX processing circuitry 325, and the TX processing circuitry 315 in accordance with well-known principles. The processor 340 can include any suitable number and type of processors or other devices in any suitable arrangement. For example, in certain embodiments, the processor 340 includes at least one microprocessor or microcontroller. Exemplary types of processors 340 include microprocessors, microcontrollers, digital signal processors, field programmable gate arrays, application specific integrated circuits, and discrete circuits.

[0042] The processor 340 is also capable of executing other processes and programs resident in the memory 360, such as operations to receive and store data. The processor 340 can move data into or out of memory 360 as required by an executing process. In certain embodiments, the processor 340 is configured to execute the one or more applications 362 based on the OS 361 or in response to receiving signals from external sources or operators. For example, the applications 362 can include an encoder, a decoder, a VR or AR application, a camera application (for still images and video), a video telephony call application, an email client, a social media client, an SMS message client, and a virtual assistant, among others. In certain embodiments, the processor 340 is configured to receive and transmit media content.

[0043] The processor 340 is also coupled to the I / O interface 345, which provides the electronic device 300 with the ability to connect to other devices such as the client devices 106-114. The I / O interface 345 is the communication path between these accessories and the processor 340.

[0044] The processor 340 is also coupled to the input 350 and the display 355. The input 350 can be used by the operator of the electronic device 300 to enter data or inputs into the electronic device 300. The input 350 can be a keyboard, a touchscreen, a mouse, a trackball, a voice input, or other device capable of functioning as a user interface to allow a user to interact with the electronic device 300. For example, the input 350 can include voice recognition processing, thereby allowing a user to input voice commands. In another example, the input 350 can include a touch panel, a (digital) pen sensor, a key, or an ultrasonic input device. The touch panel can recognize, for example, a touch input such as at least one of a capacitive scheme, a pressure sensitive scheme, an infrared scheme, or an ultrasonic scheme. The input 350 can be associated with the sensor 365 and / or the camera by providing additional inputs to the processor 340. In certain embodiments, the sensor 365 includes one or more inertial measurement units (IMUs) such as accelerometers, gyroscopes, and magnetometers, motion sensors, optical sensors, cameras, pressure sensors, heart rate sensors, altimeters, etc. The input 350 can also contain control circuitry. In the capacitive scheme, the input 350 can recognize a touch or proximity.

[0045] The display 355 can be a liquid crystal display (LCD), a light-emitting diode (LED) display, an organic LED (OLED), an active matrix OLED (AMOLED), or other display capable of rendering text and / or graphics, such as from websites, videos, games, and images. The display 355 can be sized to fit within an HMD. The display 355 can be a single display screen or multiple display screens capable of forming a stereoscopic display. In certain embodiments, the display 355 is a heads-up display (HUD). The display 355 can display 3D objects, such as stereoscopic frames.

[0046] The memory 360 is coupled to the processor 340. Part of the memory 360 can include a RAM, and another part of the memory 360 can include a flash memory or other ROM. The memory 360 can include a persistent storage (not shown) that represents any structure capable of storing and facilitating the retrieval of information such as data, program code, and / or other suitable information on a more permanent basis. The memory 360 can contain one or more components or devices, such as a read only memory, a hard disk drive, a flash memory or other media, that support longer-term data storage. The memory 360 can also contain media content. The media content can include various types of media such as images, video, three-dimensional content, VR content, AR content, 3D point clouds, stereoscopic frames, etc.

[0047] The electronic device 300 also includes one or more sensors 365 that can measure physical quantities or detect activation states of the electronic device 300 and convert the measured or detected information into electrical signals. For example, the sensors 365 can include one or more buttons for touch input, a camera, a gesture sensor, an IMU sensor such as a gyroscope or gyro sensor and an accelerometer, an eye tracking sensor, a barometric sensor, a magnetic sensor or magnetometer, a grip sensor, a proximity sensor, a color sensor, a biophysical sensor, a temperature / humidity sensor, an illuminance sensor, an ultraviolet (UV) sensor, an electromyography (EMG) sensor, an electroencephalogram (EEG) sensor, an electrocardiogram (ECG) sensor, an IR sensor, an ultrasonic sensor, an iris sensor, a fingerprint sensor, a color sensor such as a red, green, blue (RGB) sensor, etc. The sensors 365 can also include a control circuit for controlling any of the sensors included therein.

[0048] Although Figure 2 and Figure 3 various changes can be made to Figure 2 and Figure 3 For example, various components in Figure 2 and Figure 3 may be combined, further subdivided, or omitted, and additional components can be added according to particular needs. As a particular example, the processor 340 can be divided into multiple processors, such as one or more central processing units (CPUs) and one or more graphics processing units (GPUs). Also, electronic devices and servers, when placed in computing and communication, can have a variety of configurations, and Figure 2 and Figure 3 The present disclosure is not limited to any particular electronic device or server.

[0049] Figure 4 An example end-to-end pipeline of a stereoscopic rendering system 400 using a camera array and a display according to the present disclosure is shown.Figure 4 The illustrated embodiment of the stereoscopic rendering system 400 is for illustration only. Figure 4 The scope of the disclosure is not limited to any particular implementation of the electronic device.

[0050] As Figure 4 The stereoscopic rendering of the image array 600 can be performed using the linear multi-camera array 402, the stereoscopic rendering processor 404, and the display 406, as illustrated. The camera array 402, the stereoscopic rendering processor 404, and the display 406 can correspond to the imaging sensor 365, the processor 340, and the display 355, respectively, of the electronic device 300, as illustrated. Figure 3 The imaging sensor 365, the processor 340, and the display 355 of the electronic device 300, as illustrated.

[0051] The camera array 402 can be composed of an array of imaging sensors. The orientation of the imaging sensors can be adjusted to align the FOV of each imaging sensor to a projection plane. In certain embodiments, the orientation of the imaging sensors can be fixed to a projection plane at a specified distance. The imaging sensors can be aligned linearly and spaced evenly apart. In certain embodiments, the imaging sensors can be spaced unevenly apart based on the distance from a center imaging sensor or the center of the camera array 402. The imaging sensors can capture image frames or video frames simultaneously.

[0052] The stereoscopic rendering processor 404 can process the captured image or video frames into stereoscopic frames for output on the display 406. The stereoscopic rendering processor 404 can include a rectification and alignment (RA) processor 408 and a conversion processor 410. The RA processor 408 can apply a geometric transformation to each frame from the camera array 402 to create a “convergence plane” or zero disparity plane at a particular depth in the captured scene. Pixels in the frame corresponding to the convergence plane have zero values, pixels in the frame in front of the convergence plane have negative values, and pixels in the frame behind the convergence plane have positive values. This type of rectification and alignment provides frames that can be displayed correctly on the display 406. Any loss of FOV is an undesirable side effect of this geometric transformation. The conversion processor 410 can convert the rectified and aligned frames into a suitable representation, such as a stitched video, for display.

[0053] The display 406 can display the processed frames from the stereoscopic rendering processor 404. The display 406 can be a 3D display, such as a loup display. The display 406 allows for viewing of different frames from each of the individual imaging sensors from an angle on the display 406.

[0054] Although Figure 4 The stereoscopic rendering system 400 is illustrated, but the system can be implemented on Figure 4Various modifications can be made. For example, the size, shape, and dimensions of the various components in the stereoscopic rendering system 400 can be varied as needed or desired. Furthermore, the number and placement of the various components described in the stereoscopic rendering system 400 can be varied as needed or desired. Moreover, although described as a series of steps, the individual steps in the stereoscopic rendering system 400 can overlap, occur in parallel, or occur any number of times. Furthermore, the stereoscopic rendering system 400 can be used for any other suitable imaging process and is not limited to the specific process described above.

[0055] Figure 5A and Figure 5B Example FOVs 500 and 502 from the first camera 504 and the second camera 506 according to this disclosure are shown. In particular, Figure 5A The available FOV for convergence is shown to be 500, and Figure 5B The available FOV 502 for non-convergent applications is shown. Figure 5A and Figure 5B The embodiments of available FOV 500, 502 shown are for illustrative purposes only. Figure 5A and Figure 5B This disclosure is not intended to limit the scope to any particular implementation of a camera array.

[0056] like Figure 5A and Figure 5B As shown, stereo and automated stereo rendering systems typically employ off-axis convergence mode, where stereo camera pairs, such as first camera 504 and second camera 506, or multiple cameras of the stereo rendering system 400, move inward so that the first FOV 508 corresponding to the first camera 504 and the second FOV 510 corresponding to the second camera 506 converge to a convergence usable FOV 500 at the projection plane 512, as... Figure 5A As shown. Off-axis convergence mode is also known as "parallel-axis asymmetric frustum perspective projection". Parallel-axis asymmetric frustum perspective projection mode is an ideal way to create stereo pairs because it is very close to how human vision works. Off-axis convergence mode produces a convergent usable FOV 500. The convergent usable FOV 500 represents the maximum FOV amount on the projection plane 512 from the first FOV 508 and the second FOV 510.

[0057] On the contrary, such as Figure 4 As shown in Figure B, several stereo and multi-camera capture devices are equipped with physical cameras whose sensors cannot be moved, such that the optical axes of the first camera 504 and the second camera 506 are parallel, producing a non-converging usable field of view 502 at the projection plane 512. These stereo image frames or video frames cannot be directly used for viewing through an automatic stereo rendering system, therefore translation and / or cropping transformations can be applied to the images. However, these techniques result in significant cropping of the first field of view 508 and / or the second field of view 510.

[0058] Although Figure 5A and Figure 5B The available FOVs 500, 502 are shown, but various changes can be made to Figure 5A and Figure 5B For example, Figure 5A and Figure 5B The size, shape, and dimensions of the available FOVs 500, 502 and corresponding components of

[0059] Figures 6A to 6F An example FOV restoration of an image array 600 according to the present disclosure is shown. Specifically, Figure 6A The image array 600 is shown, Figure 6B The transformed image array 602 is shown, Figure 6C The available FOV 604 of the transformed image array 602 is shown, Figure 6D The blank area 606 from the transformed image array 602 is shown, Figure 6E The hierarchical restoration 608 of the transformed image array 600 is shown, and Figure 6F The first transformed frame 00 is shown at a magnification. Figures 6A to 6F The embodiments shown are for illustration only. Figures 6A to 6F The scope of the present disclosure is not limited to any particular implementation of electronic devices.

[0060] As Figure 6AAs shown, a set of video frames is captured from a one-dimensional (ID) 25x1 linear camera array in which the same cameras are arranged along a horizontal axis. Frames 00-24 of image array 600 are captured by camera array 402 with parallel camera arrangement disposed in a linear rig. The raw FOV in each frame corresponding to the FOV of a camera or optical sensor can be approximately 66°. First frame 00 corresponds to an imaging sensor at a first end of linear camera array 402, and frame 24 corresponds to an imaging sensor at a second end of linear camera array 402 opposite the first end. The scene depicted in image array 600 is a table against a wall, the table is covered with stuffed toys, and there are potted plants on each side. First frame 00 captures the farthest point on the left side of the scene, and 25th frame 24 captures the farthest point on the right side of the scene. Reference frame 610 corresponds to the 13th frame 12 in the illustrated embodiment. In general, reference frame 610 for FOV restoration utilizes a frame captured from a camera towards the center of camera array 402. However, this is not limiting, and any frame in image array 600 can be selected as a default or reference frame. Reference frame 610 in the illustrated embodiment depicts the entire table and all stuffed toys and a portion of each potted plant on both sides of the table.

[0061] As Figure 6B shown, the translational and / or shear of the captured image frames 00-24 or video frames using the multi-camera arrangement with parallel optical axes wastes a large portion of the FOV. FOV restoration requires the corresponding frame after the translational transform. The cropping in each frame 00-24 along the horizontal direction is clearly visible except for reference frame 610 (frame 12). In addition, the amount of cropping of the image is proportional to the distance of the corresponding camera in camera array 402 from the center camera. This cropping results in a reduction of the actual or usable FOV of the rendered stereoscopic content. For the illustrative example given herein, the usable FOV is reduced from 66° to approximately 24° as Figure 6C shown. FOV restoration aims to restore the FOV lost due to the shift and / or shear transform applied to the image frames and / or video frames obtained from the physical parallel arrangement of cameras in multi-camera array 402 for rendering in an autostereoscopic display such as display 406.

[0062] As Figure 6CAs shown, FOV 612 is based on the region of the scene in reference frame 610 in camera array 402. From left FOV boundary 614 to right FOV boundary 616 of reference frame 610, FOV 612 includes a portion of a potted plant, a number of stuffed toys on a table, and a portion of another potted plant. Available FOV 604 is determined based on the region of the scene viewed from each camera. In other words, left available FOV boundary 618 of available FOV 604 is based on the imaging sensor of the camera located farthest to the right of camera array 402. Right available FOV boundary 620 of available FOV 604 is based on the imaging sensor on the camera located farthest to the left of camera array 402. As an illustrative example, Figure 6D A blank region 606 is shown that corresponds to a region not captured in the corresponding frame of reference frame 610. For example, frame 00, which corresponds to the leftmost camera in camera array 402, captures the scene to the left of reference frame 610 as shown. First blank region 606a corresponds to a region in reference frame 610 not captured in first frame 00. Similarly, second blank region 606b, third blank region 606c, fourth blank region 606d, and fifth blank region 606e correspond to regions in reference frame 610 not captured in second frame 01, third frame 02, fourth frame 03, and fifth frame 04, respectively. First blank region 606a determines right available FOV boundary 620, and twenty-fifth blank region determines left available FOV boundary 618. Figure 6A

[0063] As shown in Figure 6E and Figure 6F Hierarchical restoration 608 can be used to fill blank regions 606 in each frame of transformed image array 602. As an illustrative example, Figure 6E A process for filling first blank region 606a of first frame 00 is shown, and for ease of description, Figure 6F An enlarged blank region 606 is shown. Because each frame further from first frame 00 in the sequence has increasing differences in orientation and disparity, it is more apparent to fill blank region 606 with patches from frames captured by cameras further from the first camera. However, frames from cameras closer to the first camera also have blank regions 606. In certain embodiments, a difference between first blank region 606a in first frame 00 and second blank region 606b in second frame 01 can be determined as a blank portion 622 of blank region 606 of first frame 00 and patches in the region of frame 01.

[0064] ​In certain embodiments, the blank portions 622 and the patches 624 can be determined based on different criteria. For example, the blank portions 622 and the patches 624 can be determined based on an equalized area of the blank region 606 for each of the blank portions 622. In this case, the frame for the corresponding patch will be determined based on the difference between the blank region 606a of the first frame 00 and another frame that exceeds the equalized blank portion size. In another example, the size of the blank portions 622 can be determined based on skipping a certain number of frames for the patch 624. The number of frames can be the same, such as two frames, or different, such as Figure 6E The corresponding first blank patch 624a is determined after determining the first blank portion 622a. In the illustrative example, the first blank patch 624a is copied from the fifth frame 04 and inserted into the first blank portion 622a. The second blank patch 624b is copied from the eighth frame 07 and inserted into the second blank portion 622b. The third blank patch 624c is copied from the tenth frame 09 and inserted into the third blank portion 622c. The fourth blank patch 624d is copied from the twelfth frame 11 and inserted into the fourth blank portion 622d. The fifth blank patch 624e is copied from the fourteenth frame 13 and inserted into the fifth blank portion 622e. If the final patch exceeds the final blank portion, a partial blank patch can be used, or the remaining blank portion can be filled with the corresponding portion of the reference frame 610.

[0065] In certain embodiments, the frames can include feature points 624 on the convergence plane. The feature points 624 can be used to identify the translation between frames. The translation of the feature points 624 can be used to identify the size of the blank region 606. For example, the koala can be located on the convergence plane of the frame as shown in Figure 6A Accordingly, the lateral translation of the koala is shown as moving between frame 03 and frame 04. If frame 04 is used as the first frame or reference frame, the lateral translation of the koala can be used to determine the blank region of frame 03.

[0066] In certain embodiments, a feature point is identified in a first frame and a corresponding feature point is identified in a second frame. Then, a geometric transformation matrix, such as a homography matrix, is estimated between the frames using the feature points. The size of the blank region can be determined from the translation component of the geometric transformation matrix.

[0067] In certain embodiments, dense 3D scene geometry is first reconstructed using Structure from Motion (SfM) and Multi-View Stereo (MVS). If camera intrinsic and extrinsic parameters are known (e.g., via camera calibration), these parameters are used during 3D reconstruction. Otherwise, the positions and orientations of multiple cameras can be determined according to uncalibrated and / or unstructured 3D reconstruction techniques, such as the ones used in COLMAP. Depending on the best representation of local geometry in the scene, the reconstructed scene geometry can be represented using point clouds, meshes, or a hybrid of both point clouds and meshes. By projecting (re-imaging) the reconstructed 3D geometry from the same position and orientation as the real camera corresponding to each image in the sequence, virtual perspective cameras with appropriately wide FOVs or asymmetric frustums are used to generate patches corresponding to the blank regions 606 in each image. In addition, intra-image inpainting techniques can be used to fill any occlusion holes in the projected image patches. Finally, the image patches are enlarged into each of the panned and / or cropped images to fill the blank regions 606 of the FOV 612.

[0068] In certain embodiments, a complete set of images from different viewpoints can be generated by re-imaging the reconstructed 3D geometry with a virtual camera array, rather than just the blank regions 606. In certain embodiments, the virtual cameras have larger FOVs than the original physical cameras, but adopt the same type of parallel-axial configuration as the original physical cameras. The FOVs of the virtual cameras are determined based on the FOVs of the physical cameras, the FOV required for autostereoscopic display or the FOV, and the depth of the projection plane. The newly generated images (with larger FOVs) undergo the same type of panning and / or cropping transformations required for autostereoscopic viewing.

[0069] In certain embodiments, when generating a complete set of images, the virtual cameras can adopt virtual sensor shifts or asymmetric frustums. The amount of sensor shift or the degree of asymmetry of the frustum for a particular virtual camera in the virtual camera array is a function of the distance of the virtual camera from the center or reference camera (and increases as the distance increases). Due to the sensor shifts or asymmetric frustums, the generated images can be viewed directly through an autostereoscopic display.

[0070] In certain embodiments, the blank regions 606 in each of the panned and / or cropped images can be synthesized using deep learning based view synthesis techniques specifically for view extrapolation. However, rather than directly using a network pretrained for view extrapolation, a slightly modified architecture is used to enable the network to leverage scene information from other images in the sequence.

[0071] In another embodiment, incremental and iterative new view extrapolation can be used, where a new view is synthesized from a set of real images and previously synthesized views, thereby incrementally increasing the FOV overlap within that set of camera images. As with the methods discussed earlier, the view extrapolation algorithm does not need to be completely blind; instead, it can utilize information from other images in the set.

[0072] The patches generated in the above techniques may exhibit slightly different image features (such as color variations, size variations, etc.). Therefore, filtering can be applied at the boundaries to seamlessly combine the original image with the synthesized blank area 606.

[0073] In yet another embodiment, the original image (before translation and / or shearing transformations) is used to form a layered depth representation of the scene, such as a multi-plane image (MPI). The layered depth image can be obtained via a deep learning network, such as local light field fusion for new view synthesis. Based on the position and orientation of the virtual camera, a new view of the scene can be reconstructed from the MPI-like representation of the scene by combining the components of the layered depth. Therefore, MPI-based view synthesis techniques can be combined with the aforementioned geometric reprojection techniques to restore missing areas of the field of view (FOV) in each of the translation and / or shearing images.

[0074] In another embodiment, the rotation of several cameras in the physical camera array 402 can be intentionally perturbed to produce varying degrees of front-beam configuration, thereby increasing the FOV overlap between several subsets of cameras in the camera array 402. The angle of rotation can depend on the geometry of the scene and the distance of the scene from the camera array 402. Since cameras with converging optical axes will have significant overlap in the FOV, the loss of FOV can be minimized during translation and / or shearing transformations. In yet another embodiment, this camera front-beam technique can be combined with the foregoing methods discussed in this disclosure for FOV restoration.

[0075] although Figures 6A to 6F The image array 600's FOV restoration is shown, but it is possible to... Figures 6A to 6F Various modifications can be made. For example, the size, shape, and dimensions of the image array 600 and its individual components can be varied as needed or desired. Furthermore, the number and arrangement of various images in the image array 600 can be varied as needed or desired. Moreover, the image array 600 can be used for any other suitable imaging process, and is not limited to the specific process described above.

[0076] Figures 7A to 7D An example of graded FOV restoration according to this disclosure is shown. In particular, Figure 7A The graded restoration of 700 is shown; Figure 7B An exemplary correction and alignment frame 702 is shown prior to the graded restoration 700; Figure 7CAn example inpainting frame 714 is shown; and Figure 7D An example discontinuity 703 at the patch boundary is shown. Figures 7A to 7D The illustrated embodiment of hierarchical FOV inpainting 700 is for illustration only. Figures 7A to 7D The scope of the disclosure is not limited to any particular implementation of an electronic device.

[0077] As shown in FIG. 7, the hierarchical FOV inpainting 700 can be used to fill in the blank region 606 for each of frames 00-24. The hierarchical FOV inpainting 700 performs inpainting for each frame in order starting with the reference frame 610. However, the entire blank region 606 is patched from adjacent frames. However, because each frame is processed in order, the patch includes patches from each frame between the current frame and the reference frame, including the reference frame. The multiple cameras 704 in the camera array 402 capture the multiple frames.

[0078] In addition to the reference frame, each frame is rectified and aligned to the reference frame to generate a first frame 702a, a second frame 702b, a third frame 702c, etc. The image frames (or video frames) are rectified and aligned relative to the reference view 610. For example, if the image is captured using a ID linear rig, where calibration data is available, the image is first rectified using the intrinsic and extrinsic camera parameters. Then, the images can be aligned by identifying common features in each image that originate from a selected plane of convergence in the scene, and using the common features in image pairs to find a geometric transformation matrix, such as a homography matrix. The images are then aligned using the estimated transformation matrix to render them suitable for display via a 3D display. After the alignment process, these common feature points (image points) that lie on the plane of convergence have the same pixel coordinates in each image. This step of rectification and alignment also produces a “blank” region in the image, which results in a net loss of FOV in the light field. Feature points are identified at the plane of convergence in the first frame and the second frame. The identified feature points can be used to determine a transformation matrix between the first frame and the second frame. The size of the patch can be determined from the translation component of the transformation matrix.

[0079] After rectification and alignment, each frame includes a respective blank region 606. The first frame 702a has a first blank region 606a that is smaller than a second blank region 606b of the second frame 702b, which is smaller than a third blank region 606c of the third frame 702c.

[0080] A first transform 706a is performed on the reference frame 610 to generate a first transformed frame 708a having an orientation corresponding to the first frame 702a from the second camera 704b. In other words, the first transform adjusts the reference frame 610 to have approximately the best overlap with the first frame 702a. The first transform 706a adjusts the reference frame 610 between the parameters of the first camera 704a and the parameters of the second camera 704b. The first camera 704a is located at the center of the camera array 402 and the second camera 704b is located to one side of the first camera 704a. Whether the first camera 704a and the second camera 704b have parallel optical axes or asymmetric frustums, the reference frame 610 and the first frame 702a have slightly different perspectives. The first transform 706a is used to accommodate the difference between the perspective of the reference frame 610 and the perspective of the first frame 702a. A first patch region 710a is selected from the first transformed frame 708a corresponding to the first blank region 606a in the first frame 702a. A first inpainting function 712a is performed to insert the first patch region 710a from the first transformed frame 708a into the first blank region 606a of the first frame 702a to generate a first inpainted frame 714a.

[0081] A second transform 706b is performed on the first inpainted frame 714a to generate a second transformed frame 708b corresponding to the second frame 702b from the third camera 704c. The second transform 706b adjusts the first inpainted frame 714a between the parameters of the second camera 704b and the parameters of the third camera 704c. The third camera 704c is further from the center of the camera array 402 than the second camera 704b. Whether the second camera 704b and the third camera 704c have parallel optical axes or asymmetric frustums, the first frame 702a and the second frame 702b have slightly different perspectives. The second transform 706b is used to accommodate the difference between the perspective of the first frame 702a and the perspective of the second frame 702b. A second patch region 710b is selected from the second transformed frame 708b corresponding to the second blank region 606b in the second frame 702b. A second inpainting function 712b is performed to insert the second patch region 710b from the second transformed frame 708b into the second blank region 606b of the second frame 702b to generate a second inpainted frame 714b. Then, patches can be added to the currently selected image as shown in Table 1 below.

[0082] [Table 1]

[0083]

[0084]

[0085] A third transform 706c is performed on the second inpainted frame 714b to generate a third transformed frame 708c corresponding to the third frame 702c from the fourth camera 704d. The third transform 706c adjusts the second inpainted frame 714b between the parameters of the third camera 704c and the parameters of the fourth camera 704d. The fourth camera 704d is farther from the center of the camera array 402 than the third camera 704c. Whether the third camera 704c and the fourth camera 704d have parallel optical axes or asymmetric frustums, the second frame 702b and the third frame 702c have slightly different perspectives. The third transform 706c is used to accommodate the difference between the perspective of the second frame 702b and the perspective of the third frame 702c. A third patch region 710c is selected from the third transformed frame 708c corresponding to the third blank region 606c in the third frame 702c. A third inpainting function 712c is performed to insert the third patch region 710c from the third transformed frame 708c into the third blank region 606c of the third frame 702c to generate a third inpainted frame 714c. This process can be extended to any number of cameras in the camera array 402.

[0086] As Figure 7B and Figure 7C shown, the hierarchical inpainting 700 can be performed to inpaint the missing FOV of the blank region 606 in the frame 702. One or more patch regions 710 can be copied and inserted into the corresponding blank region 606.

[0087] In certain embodiments, the reference (or just inpainted FOV) images can be warped using depth-based image warping techniques, such as depth image based rendering (DIBR), to render the reference images from the viewpoint of the selected image whose FOV is to be inpainted. If a depth map is directly available from one of many depth sensing technologies, such as LiDAR, stereo cameras, structured light sensing, etc., the depth map can be used directly. Alternatively, a depth map for each view can be estimated using stereo-based depth estimation techniques.

[0088] Some advantages of hierarchical inpainting are that patches from the closest viewpoint have the least difference in perspective and occlusion relationships. For depth-based warping, the warped patches produce the least amount of de-occlusion holes. The luminance of these patches is also the closest. Thus, in the inpainted image 714, luminance discontinuities at the patch boundaries can be minimized.

[0089] As Figure 7DAs shown, although the hierarchical FOV restoration method has the above advantages, brightness and other discontinuities can still appear at the patch boundaries. In addition, if depth-based warping is not used to warp the reference (or the just-restored image), depth discontinuities can appear at the patch boundaries in areas far from the convergence plane. The convergence plane in the example image is set to be very close to the plane passing through the sitting person. The white vertical line on the right side of the image shows the patch boundaries. Insert A shows areas far from the convergence plane that exhibit depth discontinuities at the patch boundaries. Insert B and Insert C show areas very close to the convergence plane that do not exhibit depth discontinuities at the patch boundaries.

[0090] In certain embodiments, alpha blending can be used at the boundaries while adding patches to the image to restore the missing FOV, creating a smooth transition. Once the patch size is determined, a mask (MP) is generated that has a linear gradient (from 0 to 1) portion near the patch boundaries and then a constant value of 1 in the rest of the area. The mask size matches the patch size. The width of the gradient portion can be varied in proportion to the degree of depth discontinuity and the desired degree of smoothing. A complementary mask (MI) can also be generated by subtracting MP from 1. P Then, the patch can be added to the currently selected image as shown in Equation 1 and Equation 2.

[0091] Equation 1 patch = from_image[:, w - pw :, :]

[0092] Equation 2 to_image[:, w - pw :, :] = MI * to_image[:, w - pw :, :] + M P *patch

[0093] In certain embodiments, if the corresponding depth map is available, a variable amount of blur can be applied to the patch based on the depth to reduce the impact of depth discontinuity. All embodiments of the reconstruction of the missing FOV discussed in this disclosure employ automatic detection of the FOV missing portion (i.e., FOV loss in each view during rectification and alignment).

[0094] In some embodiments, depending on a trade-off between speed and complexity, the size and location of missing portions of the field of view (FOV) in the corrected and aligned view can be determined using one of two methods. The first method determines the size of the missing area of ​​the FOV based on the translation components of a geometric transformation matrix estimated using feature points derived from a convergence plane in a reference view and corresponding feature points in the target view, where the FOV of the corresponding feature points will be restored before alignment. The second method determines the size of the missing area of ​​the FOV by comparing the target view with an adjacent previously restored view (or reference view) and finding non-overlapping areas in the target view. While the first method is simple and fast, it is less accurate than the more complex second method.

[0095] although Figures 7A to 7D The graded FOV recovery of 700 is shown, but it is possible to... Figures 7A to 7D Various modifications can be made. For example, the size, shape, and dimensions of the various components in the graded FOV restoration 700 can be varied as needed or desired. Furthermore, the number and placement of the various components in the graded FOV restoration 700 can be varied as needed or desired. Moreover, the graded FOV restoration 700 can be used for any other suitable imaging process, and is not limited to the specific processes described above.

[0096] Figure 8A and Figure 8B An example method for graded FOV restoration according to this disclosure is shown. Specifically, Figure 8A An example method 800 for graded FOV restoration is shown; and Figure 8B An example method 801 for graded FOV restoration is shown. For ease of explanation, Figure 8A and Figure 8B Methods 800 and 801 are described as using Figure 4 The stereo rendering processor 404 is used to execute this. However, methods 800 and 801 can be used with any other suitable system and any other suitable processor. Methods 800 and 801 describe the acquisition of an image array from the camera array 402, such as the FOV restoration of image array 600. Reference frame 610 is a frame corresponding to camera 704a at the center of camera array 402. Processor 404 processes the remaining frames of camera array 402 sequentially.

[0097] like Figure 8A As shown, in step 802, processor 404 receives a frame and determines whether the frame is adjacent to a reference frame. "Receive" can mean wirelessly receiving from a remote electronic device, receiving from an external electronic device via a wired connection, or loading from a storage device in the memory of an electronic device. Frames are received sequentially starting from the reference frame. When frames are captured from multiple directions of the reference frame, processor 404 can determine the first direction from which to process the frame.

[0098] At step 804, when the received frame is determined to be adjacent to the reference frame in step 802, the processor 404 selects the reference frame as the first frame. Since the frames are processed sequentially starting from the reference frame, the reference frame is the initial first frame selected when the FOV is restored directly adjacent frames. The reference frame is selected as the first frame for steps 810-816.

[0099] At step 806, when the received frame is determined not to be adjacent to the reference frame in step 802, the processor 404 selects the previous stereoscopic rendered frame as the first frame. The received frame not being directly adjacent to the reference frame means that at least one frame has been processed previously. The processor 404 determines the most recently processed frame, which is selected as the first frame for steps 810-816.

[0100] At step 808, the processor 404 selects the received frame as the second frame. For steps 810-816, the un-restored view immediately adjacent to the first image is designated as the second image. The received frame includes the blank region 606. The size of the blank region 606 increases in each frame sequentially removed from the reference frame 610. The blank region 606 is based on the difference between the region of the scene captured in the received frame or second frame and the region of the scene captured in the reference frame.

[0101] At step 810, the processor 404 estimates a 2D-2D transformation matrix that relates the first image to the second image. Due to the slightly different orientation of the first image and the second image, a 2D-2D transformation matrix is generated to transform the first frame to appear in a similar orientation as the second frame. In other words, the first frame is warped to have near-optimal overlap with the second frame. Examples of geometric transformations can include a homography matrix or an affine transformation matrix that warps the reference image such that there is near-optimal overlap (as measured by the Procrustes overlap) between the warped reference image and the selected image in the overlap region. It should be noted that due to the substantial difference in perspective between the two images, the overlap will not be exact everywhere except for points near or at the plane of convergence.

[0102] At step 812, the processor 404 warps the first image using the estimated transformation matrix. The 2D-2D transformation matrix is applied to the first frame to generate a warped version of the first frame. The first frame in the series of frames is not affected. That is, the warped version of the frame is temporarily stored. In certain embodiments, the 2D-2D transformation matrix can be applied to a region of the first frame that corresponds to the region of the blank region 606 in the second frame. In certain embodiments, the 2D-2D transformation matrix can be applied to a region of the first frame in a manner that produces the blank region 606 of the second frame. In other words, to properly warp the first frame, the warp can require regions outside of the region of the blank region 606.

[0103] In step 814, processor 404 determines the size of the patch to be copied from the distorted first image. The size of the blank area 606 in the second frame can be determined as the size of the patch from the first frame. Processor 404 can determine the area of ​​the blank area 606 as the difference between the area of ​​the scene captured in the second frame and the area of ​​the scene captured in the reference frame. The size of the patch can be derived from the translation component of the geometric transformation matrix used to align the images during correction and alignment.

[0104] In step 816, processor 404 copies the patch from the distorted first image and inserts or adds the patch to the second image. Processor 404 selects the patch based on the patch size determined in step 814. Processor 404 inserts the patch from the first frame into the blank area of ​​the second frame.

[0105] In step 818, processor 404 determines whether the field of view (FOV) has been restored for all frames on the current side of the reference frame. Since more than one frame has been captured on each side of the reference frame, processor 404 processes each frame sequentially in steps 810 to 816. When it is determined that a frame exists after the second frame, method 800 returns to step 806. When the second frame is the last frame in a series of frames on one side of the reference frame, method 800 proceeds to step 820. The pseudocode for this step is provided in Table 1 above, as well as Equations 1 and 2.

[0106] In step 820, processor 404 determines whether the field of view (FOV) has been restored on all sides of the reference frame. Processor 404 may determine whether blank areas have not yet been processed on frames directly adjacent to the reference frame. If blank areas exist in frames directly adjacent to the reference frame, method 800 returns to step 802. If blank areas no longer exist in frames directly adjacent to the reference frame, FOV restoration is complete. For 1D camera array equipment, FOV can be restored on both sides, and for 2D camera array equipment, FOV can be restored on all four sides.

[0107] like Figure 8B As shown, in step 822, the processor 404 may select a view as a reference view. The reference view may be a frame corresponding to a camera at the center of the camera array, a frame corresponding to a camera at the end of the camera array, or any other selection of a reference frame.

[0108] In step 824, processor 404 corrects and aligns each frame from the cameras in the camera array with a selected reference frame. The FOV of each frame is determined based on the number of scenes captured by the corresponding frame compared to the reference frame. This creates blank areas in each frame other than the reference frame. Since the reference frame is compared to itself, the entire scene captured by the reference frame is the FOV.

[0109] In step 826, the processor 404 determines whether the number and density of real cameras are sufficient to produce suitable geometric results. When the processor 404 determines that the number and density of real cameras are not sufficient to produce suitable geometric results, the method 801 continues the operations of the method 800. When the processor 404 determines that the number and density of real cameras are sufficient to produce suitable geometric results, the method 801 proceeds to step 828.

[0110] In step 828, the processor 404 estimates the internal and external camera parameters. Examples of parameters can include height, orientation, etc., which can be different for each camera in the camera array.

[0111] In step 830, the processor 404 reconstructs a dense geometric representation of the scene. The processor 404 can use the current frame to construct the dense geometric representation. Frames captured from the camera array can be used to construct a point cloud or other 3D model.

[0112] In step 832, the processor 404 can place virtual cameras around the reconstructed scene at respective locations and respective orientations of real cameras in the scene. In step 834, the processor 404 can move the image planes of the virtual cameras laterally to re-image some portions of the scene that were lost during rectification and alignment in each virtual camera. In step 836, the processor 404 can use 2D image inpainting techniques to fill in holes in the re-imaged patches. In step 840, the processor 404 can add the patches to the corresponding rectified and aligned views to restore the FOV.

[0113] Although Figure 8A and Figure 8B example methods 800, 801 for restoring FOV are shown, various changes can be made Figure 8A and Figure 8B For example, although shown as a series of steps, various steps in Figure 8A and Figure 8B may overlap, occur in parallel, or occur any number of times.

[0114] Figure 9 An example method 900 for restoring FOV of images captured via a parallel camera setup for stereoscopic rendering according to the present disclosure is shown. For ease of explanation, Figure 9 the method 900 is described as being performed using the processor 404 of Figure 9 However, the method 900 can be used with any other suitable system and any other suitable processor.

[0115] As Figure 9As shown, at step 902, the processor 404 receives a plurality of frames captured of a scene from a linear camera array. "Receive" can refer to receiving wirelessly from a remote electronic device, receiving from an external electronic device through a wired connection, or loading from storage from a memory of the electronic device. The frames are received in order starting from a reference frame. When the frames are captured from multiple directions from the reference frame, the processor 404 can determine a first direction to process the frames starting from the reference frame.

[0116] At step 904, the processor 404 selects a first frame and a second frame from the plurality of frames. The first frame can be the reference frame. The second frame can be directly adjacent to the first frame with respect to the position of the cameras in the linear camera array 402 that captured each of the first frame and the second frame.

[0117] At step 906, the processor 404 rectifies and aligns the first frame and the second frame to the reference frame, where the blank region of the second frame has a greater area than the blank region of the first frame. The rectification and alignment of each of the first frame and the second frame results in a frame having original information of the first frame and the second frame that is within the FOV of the reference frame, respectively. The resulting first frame and second frame can each have a blank region, each blank region having a different size. When the first frame is the reference frame, the first frame can not have a blank region.

[0118] At step 908, the processor 404 transforms the first frame to match the orientation of the second frame. The processor 404 can estimate a transformation matrix between the first frame and the second frame. The processor 404 can warp the first frame to the orientation of the second frame using the transformation matrix. The warping modifies the first frame to have near-optimal overlap with the second frame. Applying the transformation matrix creates a convergence plane of zero disparity planes at a particular depth in the scene.

[0119] At step 910, the processor 404 inserts a patch from the transformed first frame into the blank region of the second frame. The size of the patch is determined based on the blank region of the second frame. Once the patch is inserted into the blank region of the second frame, the second frame is immediately restored to the FOV of the reference frame.

[0120] Steps 904-910 can be repeated for each frame in order starting from the reference frame including a third frame. Steps 904-910 can also be repeated for each frame in order starting from the reference frame in a second direction from the linear camera array, including a fourth frame.

[0121] Although Figure 9 One example of a method 900 for restoring the FOV of images captured via a parallel camera setup for stereoscopic rendering is shown, various changes can be made Figure 9 For example, although shown as a series of steps, various steps in Figure 9 may overlap, occur in parallel, or occur any number of times.

[0122] While the present disclosure has been described with an example embodiment, various changes and modifications can be suggested to one skilled in the art. It is intended that the present disclosure encompass such changes and modifications as fall within the scope of the appended claims. None of the description in this application should be read in the limitations of any particular embodiments described. The scope of the patent should be limited only by the claims and their equivalents.

Claims

1. An apparatus comprising: at least one memory configured to receive a plurality of frames of a scene captured from a camera array; and at least one processor operably coupled to the at least one memory, the processor configured to: select a first frame and a second frame from the plurality of frames; rectify and align the first frame and the second frame to a reference frame, wherein a blank region of the second frame has a greater area than a blank region of the first frame, and wherein the blank region of the first frame corresponds to a region in the reference frame not captured by the first frame and the blank region of the second frame corresponds to a region in the reference frame not captured by the second frame; identify feature points at a convergence plane in the first frame and the second frame; determine a transformation matrix between the first frame and the second frame using the identified feature points; transform the first frame to have near-optimal overlap with the second frame using the transformation matrix; and insert a patch in a first patch region of the transformed first frame into the blank region of the second frame, wherein the first patch region of the transformed first frame corresponds to the blank region of the second frame, and wherein a size of the patch is determined according to a translation component of the transformation matrix.

2. The apparatus of claim 1, wherein the processor is further configured to: select a third frame from the plurality of frames, rectify and align the third frame to a reference frame, wherein a blank region of the third frame has a greater area than a blank region of the second frame, and the blank region of the third frame corresponds to a region in the reference frame not captured by the third frame, transform the second frame including the patch to have near-optimal overlap with the third frame, and inserting a second patch from a patch region of the transformed second frame and the transformed patch into a blank region of the third frame, wherein, a patch region of the transformed second frame and the transformed patch correspond to the blank region of the third frame.

3. The apparatus of claim 1, wherein: the first frame is a reference frame located at a center of the plurality of frames, and the processor is further configured to: select a fourth frame, the fourth frame being on an opposite side of the first frame from the second frame, rectify and align the fourth frame to a reference frame, wherein a blank region of the fourth frame has a greater area than a blank region of the first frame, and the blank region of the fourth frame corresponds to a region in the reference frame not captured by the fourth frame, transform the first frame to have near-optimal overlap with the fourth frame, and insert a third patch in a second patch region of the transformed first frame into the blank region of the fourth frame, wherein the second patch region of the transformed first frame corresponds to the blank region of the fourth frame.

4. The apparatus of claim 1, wherein, To transform the first frame to have near-optimal overlap with the second frame, the processor is further configured to: estimate a transformation matrix between the first frame and the second frame, and warp the first frame using the transformation matrix.

5. The apparatus of claim 1, wherein the first frame is a reference frame.

6. The apparatus of claim 1, wherein the processor is further configured to: A size of the patch is determined from the transformed first frame by finding a non-overlapping region between the first frame and the second frame.

7. A method comprising: receiving a plurality of frames of a scene captured from an array of cameras; selecting a first frame and a second frame from the plurality of frames; rectifying and aligning the first frame and the second frame to a reference frame, wherein a blank region of the second frame has a larger area than a blank region of the first frame, and wherein the blank region of the first frame corresponds to a region in the reference frame not captured by the first frame and the blank region of the second frame corresponds to a region in the reference frame not captured by the second frame; identifying feature points at a convergence plane in the first frame and the second frame; determining a transformation matrix between the first frame and the second frame using the identified feature points; transforming the first frame to have near-optimal overlap with the second frame using the transformation matrix; and inserting a patch from a first patch region of the transformed first frame into the blank region of the second frame, wherein the first patch region of the transformed first frame corresponds to the blank region of the second frame, and wherein a size of the patch is determined from a translation component of the transformation matrix.

8. The method of claim 7, further comprising: selecting a third frame from the plurality of frames; rectifying and aligning the third frame to a reference frame, wherein a blank region of the third frame has a larger area than the blank region of the second frame, and the blank region of the third frame corresponds to a region in the reference frame not captured by the third frame; transforming the second frame including the patch to have near-optimal overlap with the third frame; and inserting a second patch from a patch region of the transformed second frame and the transformed patch into the blank region of the third frame, wherein the patch region of the transformed second frame and the transformed patch correspond to the blank region of the third frame.

9. The method of claim 7, wherein: the first frame is a reference frame located at a center of the plurality of frames, and the method further comprises: selecting a fourth frame, the fourth frame being on an opposite side of the first frame from the second frame, rectifying and aligning the fourth frame to a reference frame, wherein a blank region of the fourth frame has a larger area than the blank region of the first frame, and the blank region of the fourth frame corresponds to a region in the reference frame not captured by the fourth frame; transforming the first frame to have near-optimal overlap with the fourth frame; and inserting a third patch from a second patch region of the transformed first frame into the blank region of the fourth frame, wherein the second patch region of the transformed first frame corresponds to the blank region of the fourth frame. To transform the first frame to have near-optimal overlap with the second frame, the method comprises:

10. The method of claim 7, wherein, estimating a transformation matrix between the first frame and the second frame, and warping the first frame using the transformation matrix.

11. The method of claim 7, wherein the first frame is a reference frame. ​ 12. The method of claim 7, further comprising: determining a size of the patch from the transformed first frame by finding a non-overlapping region between the first frame and the second frame.

13. A non-transitory computer-readable medium comprising instructions that, when executed, cause a processor to: receive a plurality of frames of a scene captured from an array of cameras; select a first frame and a second frame from the plurality of frames; rectify and align the first frame and the second frame to a reference frame, wherein a blank region of the second frame has a greater area than a blank region of the first frame, and wherein the blank region of the first frame corresponds to a region in the reference frame that is not captured by the first frame and the blank region of the second frame corresponds to a region in the reference frame that is not captured by the second frame; identify feature points at a convergence plane in the first frame and the second frame; determine a transformation matrix between the first frame and the second frame using the identified feature points; transform the first frame to have near-optimal overlap with the second frame using the transformation matrix; and insert a patch from a first patch region of the transformed first frame into the blank region of the second frame, wherein the first patch region of the transformed first frame corresponds to the blank region of the second frame, and wherein a size of the patch is determined from a translation component of the transformation matrix.