Method and apparatus for optimal overlap rate estimation for three-dimensional (3D) reconstruction
By dynamically adjusting the overlap rate, the problem of insufficient or excessive keyframes in three-dimensional reconstruction is solved, and the balance between reconstruction quality and computing resources is optimized, which is suitable for scanning objects of different sizes.
Patent Information
- Application Number
- CN202380087838.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-02-02
- Filing Date
- 2023-12-13
- Publication Date
- 2025-07-18
AI Technical Summary
In the three-dimensional reconstruction process, the number of keyframes selected for reconstruction is insufficient or too many, resulting in the balance of reconstruction quality and computing resources, especially when scanning objects of different sizes, it is difficult to optimize.
By determining the pose information of the image sensor, fitting the curve characteristics of the viewpoint set, dynamically adjusting the overlap rate to select the best keyframe, and optimizing the reconstruction process.
It realizes the effective balance of reconstruction quality and the use of computing resources when scanning objects in different sizes, and improves the efficiency and accuracy of three-dimensional reconstruction.
Smart Images

Figure CN120344997A_ABST
Abstract
Description
Technical Field
[0001] This disclosure generally relates to image processing. For example, aspects of this disclosure relate to systems and techniques for performing optimal overlap rate estimation for three-dimensional (3D) reconstruction. Background Art
[0002] Many devices and systems allow a scene to be captured by generating frames (also referred to as images) and / or video data (including multiple frames) of the scene. For example, a camera or a computing device including a camera (e.g., a mobile device including one or more cameras such as a mobile phone or a smartphone) can capture a sequence of frames of a scene. Image and / or video data can be captured and processed by such devices and systems (e.g., mobile devices, IP cameras, etc.) and can be output for consumption (e.g., displayed on the device and / or other devices). In some cases, image and / or video data can be captured by such devices and systems and output for processing and / or consumption by other devices.
[0003] Frames or images can be processed (e.g., using object detection, recognition, segmentation, etc.) to determine any objects present in the frame, which is useful for many applications. For example, a model for representing an object in a frame can be determined and used to facilitate the efficient operation of various systems. Examples of such applications and systems include extended reality (XR), robotics, automotive and aviation, three-dimensional scene understanding, object grasping, object tracking, and many other applications and systems. The term "XR" can encompass virtual reality (VR), augmented reality (AR), mixed reality (MR), etc. Summary of the Invention
[0004] Systems and techniques for performing optimal overlap rate estimation for three-dimensional (3D) reconstruction are described herein. The following presents a simplified summary of the invention related to one or more aspects disclosed herein. Accordingly, the following summary is neither to be considered an exhaustive overview of all contemplated aspects nor to identify critical or decisive elements of all contemplated aspects or to delineate the scope associated with any particular aspect. Accordingly, the following summary presents in simplified form certain concepts related to one or more aspects of the mechanisms disclosed herein prior to the detailed description presented below.
[0005] In an illustrative example, a method for image processing is provided. The method includes: obtaining pose information of an image sensor, the pose information indicating a set of viewpoints of an environment, wherein the image sensor is configured to capture frames for reconstructing an object; fitting a curve to the set of viewpoints; determining one or more features of the curve; determining the orientation of the image sensor based on the one or more features of the curve and the set of viewpoints; determining an overlap rate based on the one or more features of the curve and the determined orientation of the image sensor; and outputting the determined overlap rate to select frames to be used for reconstructing the object.
[0006] As another example, an apparatus for processing sensor data is provided. The apparatus includes at least one memory and at least one processor coupled to the at least one memory. The at least one processor is configured to: obtain pose information of an image sensor, the pose information indicating a set of viewpoints of an environment, wherein the image sensor is configured to capture frames for reconstructing an object; fit a curve to the set of viewpoints; determine one or more features of the curve; determine the orientation of the image sensor based on the one or more features of the curve and the set of viewpoints; determine an overlap rate based on the one or more features of the curve and the determined orientation of the image sensor; and output the determined overlap rate to select frames to be used for reconstructing the object.
[0007] In another example, a non-transitory computer-readable medium storing instructions is provided. The instructions, when executed by at least one processor, cause the at least one processor to: obtain pose information of an image sensor, the pose information indicating a set of viewpoints of an environment, wherein the image sensor is configured to capture frames for reconstructing an object; fit a curve to the set of viewpoints; determine one or more features of the curve; determine the orientation of the image sensor based on the one or more features of the curve and the set of viewpoints; determine an overlap rate based on the one or more features of the curve and the determined orientation of the image sensor; and output the determined overlap rate to select frames to be used for reconstructing the object.
[0008] As another example, an apparatus for image processing is provided. The apparatus includes: means for obtaining pose information of an image sensor, the pose information indicating a set of viewpoints of an environment, wherein the image sensor is configured to capture frames for reconstructing an object; means for fitting a curve to the set of viewpoints; determining one or more features of the curve; means for determining the orientation of the image sensor based on the one or more features of the curve and the set of viewpoints; means for determining an overlap rate based on the one or more features of the curve and the determined orientation of the image sensor; and means for outputting the determined overlap rate to select frames to be used for reconstructing the object.
[0009] In some aspects, one or more of the devices described herein may include or be part of the following: an extended reality device (e.g., a virtual reality (VR) device, an augmented reality (AR) device, or a mixed reality (MR) device, such as a head-mounted display (HMD) or glasses), a mobile device (e.g., a mobile phone or other mobile device), a vehicle or a computing system or component of a vehicle, a wearable device (e.g., a network-connected watch or other wearable device), a personal computer, a laptop computer, a server computer, a television, a video game console, or other device. In some aspects, the device further includes at least one camera for capturing one or more images or video frames. For example, the device may include one camera (e.g., an RGB camera) or multiple cameras for capturing one or more images and / or one or more videos including video frames. In some aspects, the device includes a display for displaying one or more images, videos, notifications, or other displayable data. In some aspects, the device includes a transmitter configured to send data or information to at least one device via a transmission medium. In some aspects, the processor includes a central processing unit (CPU), a graphics processing unit (GPU), a neural processing unit (NPU), or other processing device or component.
[0010] This summary is not intended to identify key or essential features of the claimed subject matter, nor is it intended to be used in isolation to determine the scope of the claimed subject matter. The subject matter should be understood by reference to the appropriate portions of the entire specification of this patent, any or all of the drawings, and each claim.
[0011] The foregoing and other features and examples will become more apparent after reference to the following specification, claims, and drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] Exemplary examples of the present application are described in detail below with reference to the following drawings:
[0013] Figure 1 is a block diagram illustrating the architecture of an image capture and processing system according to some examples;
[0014] Figure 2 is a diagram illustrating the architecture of an exemplary extended reality (XR) system according to some examples;
[0015] Figure 3 is a block diagram illustrating an example of a model generation system according to some examples;
[0016] Figure 4 is a diagram illustrating an example operation of an image capture device capturing an input frame according to some examples;
[0017] Figure 5A andFigure 5B , Figure 5C and Figure 5D illustrate examples of overlap rates in accordance with aspects of the present disclosure;
[0018] Figure 6 is a top view of a room in accordance with aspects of the present disclosure, which illustrates a scanned object for reconstruction;
[0019] Figure 7 is a block diagram illustrating an architecture of an example overlap rate adjustment system for optimizing an overlap rate in accordance with aspects of the present disclosure;
[0020] Figure 8A and Figure 8B illustrate a set of viewpoints on a fitted circle in accordance with aspects of the present disclosure;
[0021] Figure 9 is a flowchart illustrating a process for image processing in accordance with aspects of the present disclosure;
[0022] Figure 10A is a perspective view of a head-mounted display (HMD) implementing feature tracking and / or visual simultaneous localization and mapping (VSLAM) in accordance with some examples.
[0023] Figure 10B is an illustration of a head-mounted display (HMD) being worn by a user in accordance with some examples Figure 10A of.
[0024] Figure 11A is a perspective view of a front surface of a mobile device implementing feature tracking and / or visual simultaneous localization and mapping (VSLAM) using one or more front cameras in accordance with some examples.
[0025] Figure 11B is a perspective view of a back surface of a mobile device in accordance with aspects of the present disclosure.
[0026] Figure 12 is a diagram illustrating an example of a system for implementing certain aspects of the present technology. DETAILED DESCRIPTION
[0027] Certain aspects and examples of the present disclosure are provided below. As will be apparent to those skilled in the art, some of these aspects and examples may be applied independently, and some of them may be applied in combination. In the following description, for purposes of explanation, specific details are set forth in order to provide a thorough understanding of the subject matter of the present application. However, it will be apparent that the various examples may be practiced without these specific details. The accompanying drawings and description are not intended to be restrictive.
[0028] The following description provides only illustrative examples and is not intended to limit the scope, applicability, or configuration of the present disclosure. Rather, the following description will provide a viable description for those skilled in the art to implement the illustrative examples. It should be understood that various changes can be made to the functions and arrangements of the elements without departing from the essence and scope of the present application as set forth in the appended claims.
[0029] The generation of three-dimensional (3D) models of physical objects can be used in many systems and applications, such as extended reality (XR) (e.g., including augmented reality (AR), virtual reality (VR), mixed reality (MR), etc.), robotics, automotive, aviation, 3D scene understanding, object grasping, object tracking, and many other systems and applications. In an AR environment, for example, a user can view a frame or image that includes the integration of artificial or virtual graphics with the user's natural environment. As used herein, the terms "frame" and "image" can be used interchangeably. For example, a frame or image can be captured by a device's camera and can include pixel data that defines objects, backgrounds, and / or other information in the scene captured by the image. AR applications allow processing of the frames to add virtual objects to these frames and align or register the virtual objects with these frames in multiple dimensions. For example, real-world objects that exist in reality can be represented using models that are similar to or an exact match for the real-world objects. In one example, a model of a virtual airplane that represents a real airplane located on or moving on a runway can be presented in the view of an AR device (e.g., a mobile device, AR glasses, an AR head-mounted display (HMD), or other device), while the user continues to view their natural environment in the AR environment. A viewer may be able to manipulate the model while viewing the real-world scene. In another example, models with different colors or different physical properties in the AR environment can be used to identify and render actual objects located on or moving on a table. In some cases, artificial virtual objects that do not exist in reality or computer-generated copies of actual objects or structures in the user's natural environment can also be added to the AR environment.
[0030] A 3D object scanning application is provided to allow users to build high-quality 3D models with short processing times. The 3D models can include 3D meshes with points of different depths. Various devices are capable of performing 3D object scanning functions. By combining novel sensors with cutting-edge tracking algorithms, device manufacturers (e.g., original equipment manufacturers or OEMs) can provide 3D object scanning capabilities for consumer-grade devices (e.g., mobile phones such as smartphones, XR devices such as AR glasses and VR HMDs, and other devices). By providing 3D object scanning capabilities to consumer-grade devices, more users with different skills can generate novel content for the virtual world.
[0031] To perform 3D object scanning on a target object, a device may capture a sequence of frames (e.g., a series of frames or a video) of the target object from different views (e.g., from different positions and angles). The sequence of frames may then be used to generate a 3D model (also referred to as 3D reconstruction) of the target object. In some cases, not every frame may be usable for reconstruction because the difference between one frame and the next may not be sufficient to justify the processing cost of using that frame. However, reducing the number of key frames used for reconstruction may differently affect the quality of reconstruction depending on the size of the object being reconstructed. Therefore, using a single overlap rate to select the frames to be used as key frames (e.g., frames for the reconstruction of an object) may not be an optimal practice.
[0032] This document describes systems, apparatuses, electronic devices, methods (also referred to as processes), and computer-readable media for performing optimal overlap rate estimation (collectively referred to herein as "systems and techniques"). For example, since the use cases for reconstruction may vary based on the size of the object being reconstructed, the overlap rate may be selected based on the size of the object being reconstructed. To help account for different sizes of objects, the sizes may be divided into size categories, and overlap rates may be predefined for the size categories. In some cases, to scan an object, multiple images of the object may be captured at different angles relative to the object, and when scanning the object, the camera may move around (e.g., around the object or around the user of the camera).
[0033] In some cases, how the camera moves around when scanning an object may be used to determine what size category the object being scanned should be in. Generally, when scanning a small object, the camera may move in a smaller arc (e.g., with a smaller radius) than when scanning a larger object. Additionally, when scanning a large object, the camera may point to a position away from the center of the arc, such as when scanning a room. Based on the size of the arc and the position the camera points to, an estimate of the size category of the object and the corresponding overlap rate may be determined.
[0034] In some cases, adjusting the overlap rate based on the size category of the object being scanned may allow multiple key frames to be adapted to the size of the object being reconstructed. Since the optimal number of key frames for reconstructing an object varies based on the size of the object, adjusting the number of key frames based on size helps balance the computational resources used for reconstruction with the quality of the reconstruction.
[0035] Various aspects of the present application will be described with reference to the drawings. Figure 1is a block diagram illustrating the architecture of an image capture and processing system 100. The image capture and processing system 100 includes various components for capturing and processing images of a scene (e.g., an image of scene 110). The image capture and processing system 100 can capture individual images (or photos), and / or can capture video including multiple images (or video frames) in a particular sequence. In some cases, the lens 115 and the image sensor 130 can be associated with an optical axis. In one illustrative example, both the photosensitive area of the image sensor 130 (e.g., a photodiode) and the lens 115 can be centered on the optical axis. The lens 115 of the image capture and processing system 100 faces the scene 110 and receives light from the scene 110. The lens 115 bends the incoming light from the scene towards the image sensor 130. The light received by the lens 115 passes through an aperture. In some cases, the aperture (e.g., aperture size) is controlled by one or more control mechanisms 120 and is received by the image sensor 130. In some cases, the aperture can have a fixed size.
[0036] One or more control mechanisms 120 can control exposure, focus, and / or zoom based on information from the image sensor 130 and / or based on information from the image processor 150. One or more control mechanisms 120 can include multiple mechanisms and components; for example, the control mechanism 120 can include one or more exposure control mechanisms 125A, one or more focus control mechanisms 125B, and / or one or more zoom control mechanisms 125C. One or more control mechanisms 120 can also include additional control mechanisms other than those illustrated, such as control mechanisms for controlling analog gain, flash, HDR, depth of field, and / or other image capture attributes.
[0037] The focusing control mechanism 125B of the control mechanism 120 can obtain a focusing setting. In some examples, the focusing control mechanism 125B stores the focusing setting in a memory register. Based on the focusing setting, the focusing control mechanism 125B can adjust the positioning of the lens 115 relative to the positioning of the image sensor 130. For example, based on the focusing setting, the focusing control mechanism 125B can move the lens 115 closer to or farther away from the image sensor 130 by actuating a motor or a servo system (or other lens mechanism) to adjust the focus. In some cases, additional lenses may be included in the image capture and processing system 100, such as one or more microlenses located above each photodiode of the image sensor 130, each of the one or more microlenses bending the light received from the lens 115 towards the corresponding photodiode before the light reaches the photodiode. The focusing setting can be determined via contrast detection autofocus (CDAF), phase detection autofocus (PDAF), hybrid autofocus (HAF), or some combination thereof. The control mechanism 120, the image sensor 130, and / or the image processor 150 can be used to determine the focusing setting. The focusing setting can be referred to as an image capture setting and / or an image processing setting. In some cases, the lens 115 can be fixed relative to the image sensor, and the focusing control mechanism 125B can be omitted without departing from the scope of the present disclosure.
[0038] The exposure control mechanism 125A of the control mechanism 120 can obtain an exposure setting. In some cases, the exposure control mechanism 125A stores the exposure setting in a memory register. Based on the exposure setting, the exposure control mechanism 125A can control the size of the aperture (e.g., aperture size or aperture value), the duration for which the aperture is open (e.g., exposure time or shutter speed), the duration for which the sensor collects light (e.g., exposure time or electronic shutter speed), the sensitivity of the image sensor 130 (e.g., ISO speed or film speed), the analog gain applied by the image sensor 130, or any combination thereof. The exposure setting can be referred to as an image capture setting and / or an image processing setting.
[0039] The zoom control mechanism 125C of the control mechanism 120 can obtain a zoom setting. In some examples, the zoom control mechanism 125C stores the zoom setting in a memory register. Based on the zoom setting, the zoom control mechanism 125C can control the focal length of an assembly of lens elements (lens assembly) including the lens 115 and one or more additional lenses. For example, the zoom control mechanism 125C can control the focal length of the lens assembly by actuating one or more motors or servo systems (or other lens mechanisms) to move one or more lenses relative to each other. The zoom setting can be referred to as an image capture setting and / or an image processing setting. In some examples, the lens assembly can include a parfocal zoom lens or a variable focal length zoom lens. In some examples, the lens assembly can include a focusing lens (in some cases, the focusing lens can be the lens 115), which first receives light from the scene 110, and the light then passes through an afocal zoom system between the focusing lens (e.g., the lens 115) and the image sensor 130 before reaching the image sensor 130. In some cases, the afocal zoom system can include two positive (e.g., converging, convex) lenses with equal or similar focal lengths (e.g., within a threshold difference from each other), with a negative (e.g., diverging, concave) lens between them. In some cases, the zoom control mechanism 125C moves one or more of the lenses in the afocal zoom system, such as the negative lens and one or both of the positive lenses. In some cases, the zoom control mechanism 125C can control the zoom by capturing an image from an image sensor (e.g., including the image sensor 130) among a plurality of image sensors with a zoom corresponding to the zoom setting. For example, the image capture and processing system 100 can include a wide-angle image sensor with a relatively low zoom and a telephoto image sensor with a larger zoom. In some cases, based on the selected zoom setting, the zoom control mechanism 125C can capture an image from the corresponding sensor.
[0040] The image sensor 130 includes one or more arrays of photodiodes or other photosensitive elements. Each photodiode measures the amount of light that ultimately corresponds to a specific pixel in the image produced by the image sensor 130. In some cases, different photodiodes may be covered by different filters. In some cases, different photodiodes may be covered in different color filters and may thus measure light that matches the color of the filter covering the photodiode. Various color filter arrays can be used, including a Bayer color filter array, a four-color filter array (also known as a four-color Bayer color filter array or QCFA), and / or any other color filter array. For example, a Bayer color filter includes a red color filter, a blue color filter, and a green color filter, where each pixel of the image is generated based on red light data from at least one photodiode covered in the red color filter, blue light data from at least one photodiode covered in the blue color filter, and green light data from at least one photodiode covered in the green color filter.
[0041] Back to Figure 1 , other types of color filters can use yellow, magenta, and / or cyan (also known as "emerald") color filters as an alternative or supplement to the red, blue, and / or green color filters. In some cases, some photodiodes may be configured to measure infrared (IR) light. In some embodiments, the photodiodes that measure IR light may not be covered by any filter, thus allowing the IR photodiodes to measure both visible light (e.g., color) and IR light. In some examples, the IR photodiodes may be covered by an IR filter, thus allowing IR light to pass through and blocking light from other parts of the spectrum (e.g., visible light, color). Some image sensors (e.g., image sensor 130) may lack filters altogether (e.g., color, IR, or any other part of the spectrum) and may instead use different photodiodes (vertically stacked in some cases) throughout the pixel array. The different photodiodes throughout the pixel array may have different spectral sensitivity curves, thereby responding to light of different wavelengths. Monochrome image sensors may also lack filters and thus lack color depth.
[0042] In some cases, image sensor 130 may alternatively or additionally include opaque and / or reflective masks that block light from reaching certain photodiodes or portions of certain photodiodes at certain times and / or from certain angles. In some cases, the opaque and / or reflective masks may be used for phase detection autofocus (PDAF). In some cases, the opaque and / or reflective masks may be used to block portions of the electromagnetic spectrum from reaching the photodiodes of the image sensor (e.g., IR cut-off filter, UV cut-off filter, band-pass filter, low-pass filter, high-pass filter, etc.). Image sensor 130 may also include an analog gain amplifier for amplifying the analog signals output by the photodiodes and / or an analog-to-digital converter (ADC) for converting the analog signals output by the photodiodes (and / or the analog signals amplified by the analog gain amplifier) into digital signals. In some cases, certain components or functions discussed with respect to one or more of the control mechanisms 120 may alternatively or additionally be included in image sensor 130. Image sensor 130 may be a charge-coupled device (CCD) sensor, an electron-multiplying CCD (EMCCD) sensor, an active pixel sensor (APS), complementary metal-oxide semiconductor (CMOS), N-type metal-oxide semiconductor (NMOS), a hybrid CCD / CMOS sensor (e.g., sCMOS), or some other combination thereof.
[0043] Image processor 150 may include one or more processors, such as one or more image signal processors (ISPs) (including ISP 154), one or more host processors (including host processor 152), and / or with respect to Figure 12One or more processors of any other type of processor 1210 discussed in the computing system 1200. The host processor 152 can be a digital signal processor (DSP) and / or other types of processors. In some specific implementations, the image processor 150 is a single integrated circuit or chip (e.g., referred to as a system-on-chip or SoC) that includes the host processor 152 and the ISP 154. In some cases, the chip may also include one or more input / output ports (e.g., input / output (I / O) port 156), a central processing unit (CPU), a graphics processing unit (GPU), a broadband modem (e.g., 3G, 4G or LTE, 5G, etc.), a memory, connection components (e.g., Bluetooth™, Global Positioning System (GPS), etc.), any combination thereof and / or other components. The I / O port 156 can include any suitable input / output port or interface according to one or more protocols or specifications, such as an Inter-Integrated Circuit 2 (I2C) interface, an Inter-Integrated Circuit 3 (I3C) interface, a Serial Peripheral Interface (SPI) interface, a Serial General Purpose Input / Output (GPIO) interface, a Mobile Industry Processor Interface (MIPI) (such as a MIPI CSI-2 Physical (PHY) layer port or interface, an Advanced High-Performance Bus (AHB) bus, any combination thereof and / or other input / output ports. In an illustrative example, the host processor 152 can communicate with the image sensor 130 using an I2C port, and the ISP 154 can communicate with the image sensor 130 using a MIPI port.
[0044] The image processor 150 can perform multiple tasks, such as demosaicking, color space conversion, image frame downsampling, pixel interpolation, automatic exposure (AE) control, automatic gain control (AGC), CDAF, PDAF, automatic white balance, merging image frames to form an HDR image, image recognition, object recognition, feature recognition, receiving inputs, managing outputs, managing memory, or some combination thereof. The image processor 150 can store the image frames and / or the processed images in a random access memory (RAM) 140 / 1025, a read-only memory (ROM) 145 / 1020, a cache, a memory cell, another storage device, or some combination thereof.
[0045] A variety of input / output (I / O) devices 160 may be connected to the image processor 150. The I / O devices 160 may include a display screen, a keyboard, a keypad, a touch screen, a touchpad, a touch-sensitive surface, a printer, any other output device, any other input device, or some combination thereof. In some cases, captions may be input into the image processing device 105B through the physical keyboard or keypad of the I / O device 160, or through the virtual keyboard or keypad of the touch screen of the I / O device 160. The I / O 160 may include one or more ports, jacks, or other connectors that implement a wired connection between the image capture and processing system 100 and one or more peripheral devices, through which the image capture and processing system 100 may receive data from and / or send data to one or more peripheral devices. The I / O 160 may include one or more wireless transceivers that implement a wireless connection between the image capture and processing system 100 and one or more peripheral devices, through which the image capture and processing system 100 may receive data from and / or send data to one or more peripheral devices. The peripheral devices may include any type of I / O device 160 discussed previously, and once they are coupled to a port, jack, wireless transceiver, or other wired and / or wireless connector, they themselves may be considered I / O devices 160.
[0046] In some cases, the image capture and processing system 100 may be a single device. In some cases, the image capture and processing system 100 may be two or more separate devices, including an image capture device 105A (e.g., a camera) and an image processing device 105B (e.g., a computing device coupled to the camera). In some specific embodiments, the image capture device 105A and the image processing device 105B may be coupled together, for example, via one or more wires, cables, or other electrical connectors, and / or wirelessly via one or more wireless transceivers. In some specific embodiments, the image capture device 105A and the image processing device 105B may be disconnected from each other.
[0047] As Figure 1 shown, the vertical dashed line divides Figure 1 the image capture and processing system 100 into two parts, representing the image capture device 105A and the image processing device 105B, respectively. The image capture device 105A includes a lens 115, a control mechanism 120, and an image sensor 130. The image processing device 105B includes an image processor 150 (including an ISP 154 and a host processor 152), a RAM 140, a ROM 145, and an I / O device 160. In some cases, certain components illustrated in the image capture device 105A (such as the ISP 154 and / or the host processor 152) may be included in the image capture device 105A.
[0048] The image capture and processing system 100 may include an electronic device, such as a mobile or fixed telephone handset (e.g., a smart phone, a cellular phone, etc.), a desktop computer, a laptop or notebook computer, a tablet computer, a set-top box, a television, a camera, a display device, a digital media player, a video game console, a video streaming device, an Internet Protocol (IP) camera, or any other suitable electronic device. In some examples, the image capture and processing system 100 may include one or more wireless transceivers for wireless communication (such as cellular network communication, 802.10 Wi-Fi communication, wireless local area network (WLAN) communication, or some combination thereof). In some embodiments, the image capture device 105A and the image processing device 105B may be different devices. For example, the image capture device 105A may include a camera device, and the image processing device 105B may include a computing device, such as a mobile phone, a desktop computer, or other computing device.
[0049] Although the image capture and processing system 100 is shown as including certain components, those of ordinary skill in the art should understand that the image capture and processing system 100 may include more components than Figure 1 those shown therein. The components of the image capture and processing system 100 may include software, hardware, or one or more combinations of software and hardware. For example, in some embodiments, the components of the image capture and processing system 100 may include electronic circuits or other electronic hardware, and / or may be implemented using electronic circuits or other electronic hardware, which may include one or more programmable electronic circuits (e.g., a microprocessor, a GPU, a DSP, a CPU, and / or other suitable electronic circuits), and / or may include computer software, firmware, or any combination thereof, and / or may be implemented using computer software, firmware, or any combination thereof, to perform the various operations described herein. The software and / or firmware may include one or more instructions stored on a computer-readable storage medium and executable by one or more processors of the electronic device implementing the image capture and processing system 100.
[0050] Figure 2 is a schematic diagram illustrating the architecture of an example extended reality (XR) system 200 in accordance with some aspects of the present disclosure. In some examples, Figure 2 the extended reality (XR) system 200 may include the image capture and processing system 100, the image capture device 105A, the image processing device 105B, or a combination thereof. In some examples, Figure 3 the model generation system 300 may include the image capture and processing system 100, the image capture device 105A, the image processing device 105B, or a combination thereof.
[0051] The XR system 200 can run (or execute) XR applications and implement XR operations. In some examples, as part of the XR experience, the XR system 200 can perform tracking and positioning, map building of the environment in the physical world (e.g., a scene), and / or positioning and rendering of virtual content on a display 209 (e.g., a screen, a visible plane / area, and / or other displays). For example, the XR system 200 can generate a map of the environment in the physical world (e.g., a three-dimensional (3D) map), track the pose (e.g., position and orientation) of the XR system 200 relative to the environment (e.g., relative to the 3D map of the environment), position and / or anchor virtual content at a specific location on the map of the environment, and render the virtual content on the display 209 such that the virtual content appears to be at a location in the environment corresponding to the specific location on the map of the scene, where the virtual content is positioned and / or anchored at the map of the scene. The display 209 can include glass, a screen, lenses, projectors, and / or other display mechanisms that allow a user to see the real-world environment and also allow XR content to be overlaid, overlapped, blended, or otherwise displayed thereon.
[0052] In this illustrative example, the XR system 200 includes one or more image sensors 202, an accelerometer 204, a gyroscope 206, a storage device 207, a computing component 210, an XR engine 220, an image processing engine 224, a rendering engine 226, a communication engine 228, and a model generation engine 230. It should be noted that Figure 2 the components 202 to 230 shown are non-limiting examples provided for illustrative and explanatory purposes, and other examples may include more, fewer, or different components compared to Figure 2 the components shown. For example, in some cases, the XR system 200 can include one or more other sensors (e.g., one or more inertial measurement units (IMUs), radar, light detection and ranging (LIDAR) sensors, radio detection and ranging (RADAR) sensors, sound detection and ranging (SODAR) sensors, sound navigation and ranging (SONAR) sensors, audio sensors, etc.), one or more display devices, one or more other processing engines, one or more other hardware components, and / or Figure 2 one or more other software and / or hardware components not shown. Although the various components of the XR system 200 (such as the image sensor 202) may be referred to in the singular form herein, it should be understood that the XR system 200 can include multiple components of any of the components discussed herein (e.g., multiple image sensors 202).
[0053] The XR system 200 includes an input device 208 or communicates (wired or wirelessly) with the input device. The input device 208 can include any suitable input device, such as a touch screen, a pen or other pointer device, a keyboard, a mouse, buttons or keys, a microphone for receiving voice commands, a gesture input device for receiving gesture commands, a video game controller, a steering wheel, a joystick, a set of buttons, a trackball, a remote control, any other input device 1045 discussed herein, or any combination thereof. In some cases, the image sensor 202 can capture images that can be processed for gesture command interpretation.
[0054] The XR system 200 can also communicate (wired or wirelessly) with one or more other electronic devices. For example, the communication engine 228 can be configured to manage connections and communicate with one or more electronic devices. In some cases, the communication engine 228 can correspond to Figure 12 the communication interface 1240.
[0055] In some cases, the XR system 200 can generate a 3D reconstruction of an object using, for example, a sequence of frames of the target object. For example, the model generation engine 230 can be configured to obtain images and generate a 3D reconstruction of the object based on the obtained images.
[0056] In some specific implementations, one or more of the image sensor 202, accelerometer 204, gyroscope 206, storage device 207, computing component 210, XR engine 220, image processing engine 224, rendering engine 226, communication engine 228, and model generation engine 230 can be part of the same computing device. For example, in some cases, one or more of the image sensor 202, accelerometer 204, gyroscope 206, storage device 207, computing component 210, XR engine 220, image processing engine 224, and rendering engine 226 can be integrated into an HMD, extended reality glasses, a smart phone, a laptop computer, a tablet computer, a gaming system, and / or any other computing device. However, in some specific implementations, one or more of the image sensor 202, accelerometer 204, gyroscope 206, storage device 207, computing component 210, XR engine 220, image processing engine 224, rendering engine 226, communication engine 228, and model generation engine 230 can be part of two or more separate computing devices. For example, in some cases, some of the components 202 - 230 can be part of or implemented by one computing device, and the remaining components can be part of or implemented by one or more other computing devices.
[0057] The storage device 207 can be any storage device for storing data. Additionally, the storage device 207 can store data from any of the components in the XR system 200. For example, the storage device 207 can store data from the image sensor 202 (e.g., image or video data), data from the accelerometer 204 (e.g., measurements), data from the gyroscope 206 (e.g., measurements), data from the computing component 210 (e.g., processing parameters, preferences, virtual content, rendered content, scene maps, tracking and positioning data, object detection data, privacy data, XR application data, face recognition data, occlusion data, etc.), data from the XR engine 220, data from the image processing engine 224, data from the rendering engine 226 (e.g., output frames), data from the communication engine, and / or data from the model generation engine 230 (e.g., 3D reconstruction). In some examples, the storage device 207 can include a buffer for storing frames to be processed by the computing component 210.
[0058] One or more computing components 210 can include a central processing unit (CPU) 212, a graphics processing unit (GPU) 214, a digital signal processor (DSP) 216, an image signal processor (ISP) 218, and / or other processors (e.g., a neural processing unit (NPU) implementing one or more trained neural networks). The computing component 210 can perform various operations such as image enhancement, computer vision, graphics rendering, extended reality operations (e.g., tracking, positioning, pose estimation, map building, content anchoring, content rendering, etc.), image and / or video processing, sensor processing, recognition (e.g., text recognition, face recognition, object recognition, feature recognition, tracking or pattern recognition, scene recognition, occlusion detection, etc.), trained machine learning operations, filtering, and / or any of the various operations described herein. In some examples, the computing component 210 can implement (e.g., control, operate, etc.) the XR engine 220, the image processing engine 224, the rendering engine 226, and the model generation engine 230. In other examples, the computing component 210 can also implement one or more other processing engines.
[0059] The image sensor 202 may include any image and / or video sensor or capture device. In some examples, the image sensor 202 may be part of a multi-camera assembly (such as a dual-camera assembly). The image sensor 202 may capture image and / or video content (e.g., raw image and / or video data), which may then be processed by the computing component 210, the XR engine 220, the image processing engine 224, the rendering engine 226, and / or the model generation engine 230, as described herein. In some examples, the image sensor 202 may include the image capture and processing system 100, the image capture device 105A, the image processing device 105B, or a combination thereof.
[0060] In some examples, the image sensor 202 may capture image data and may generate an image (also referred to as a frame) based on the image data and / or may provide the image data or frame to the XR engine 220, the image processing engine 224, and / or the rendering engine 226 for processing. The image or frame may include a video frame or a static image in a video sequence. The image or frame may include an array of pixels representing a scene. For example, the image may be a red, green, and blue (RGB) image having red, green, and blue color components per pixel; a luminance, chroma red, chroma blue (YCbCr) image having a luminance component and two chroma (color) components (chroma red and chroma blue) per pixel; or any other suitable type of color or monochrome image.
[0061] In some cases, the image sensor 202 (and / or other cameras of the XR system 200) may also be configured to capture depth information. For example, in some implementations, the image sensor 202 (and / or other cameras) may include an RGB-depth (RGB-D) camera. In some cases, the XR system 200 may include one or more depth sensors (not shown) that are separate from the image sensor 202 (and / or other cameras) and may capture depth information. For example, such depth sensors may obtain depth information independently of the image sensor 202. In some examples, the depth sensor may be physically mounted in the same general location as the image sensor 202 but may operate at a different frequency or frame rate than the image sensor 202. In some examples, the depth sensor may take the form of a light source that projects a structured or textured light pattern (which may include one or more narrowband lights) onto one or more objects in the scene. Depth information may then be obtained by leveraging the geometric deformation of the projected pattern caused by the surface shape of the object. In one example, depth information may be obtained from a stereo sensor, such as a combination of an infrared structured light projector and an infrared camera registered to a camera (e.g., an RGB camera).
[0062] The XR system 200 may also include other sensors in one or more of its sensors. The one or more sensors may include one or more accelerometers (e.g., accelerometer 204), one or more gyroscopes (e.g., gyroscope 206), and / or other sensors. The one or more sensors may provide velocity, orientation, and / or other position-related information to the computing component 210. For example, the accelerometer 204 may detect the acceleration of the XR system 200 and may generate an acceleration measurement based on the detected acceleration. In some cases, the accelerometer 204 may provide one or more translation vectors (e.g., up / down, left / right, forward / backward), which may be used to determine the position or pose of the XR system 200. The gyroscope 206 may detect and measure the orientation and angular velocity of the XR system 200. For example, the gyroscope 206 may be used to measure the pitch, roll, and yaw of the XR system 200. In some cases, the gyroscope 206 may provide one or more rotation vectors (e.g., pitch, yaw, roll). In some examples, the image sensor 202 and / or the XR engine 220 may use the measurements obtained by the accelerometer 204 (e.g., one or more translation vectors) and / or the measurements obtained by the gyroscope 206 (e.g., one or more rotation vectors) to calculate the pose of the XR system 200. As previously mentioned, in other examples, the XR system 200 may also include other sensors, such as an inertial measurement unit (IMU), a magnetometer, a gaze and / or eye tracking sensor, a machine vision sensor, an intelligent scene sensor, a speech recognition sensor, a shock sensor, a vibration sensor, a positioning sensor, a tilt sensor, etc.
[0063] As described above, in some cases, the one or more sensors may include at least one IMU. An IMU is an electronic device that uses a combination of one or more accelerometers, one or more gyroscopes, and / or one or more magnetometers to measure the specific force, angular velocity, and / or orientation of the XR system 200. In some examples, the one or more sensors may output the measured information associated with the capture of an image captured by the image sensor 202 (and / or other cameras of the XR system 200) and / or depth information obtained using one or more depth sensors of the XR system 200.
[0064] The XR engine 220 may use the outputs of one or more sensors (e.g., accelerometer 204, gyroscope 206, one or more IMUs, and / or other sensors) to determine the pose (also referred to as the head pose) of the XR system 200 and / or the pose of the image sensor 202 (or other cameras of the XR system 200). In some cases, the pose of the XR system 200 and the pose of the image sensor 202 (or other cameras) may be the same. The pose of the image sensor 202 refers to the positioning and orientation of the image sensor 202 relative to a reference frame (e.g., with respect to the scene 110). In some embodiments, the camera pose may be determined for six degrees of freedom (6DoF), which refers to three translational components (e.g., which may be given by the X (horizontal), Y (vertical), and Z (depth) coordinates relative to a reference frame such as the image plane) and three angular components (e.g., roll, pitch, and yaw relative to the same reference frame). In some embodiments, the camera pose may be determined for three degrees of freedom (3DoF), which refers to three angular components (e.g., roll, pitch, and yaw).
[0065] In some cases, a device tracker (not shown) may use measurements from one or more sensors and image data from the image sensor 202 to track the pose of the XR system 200 (e.g., 6DoF pose). For example, the device tracker may fuse visual data from the image data (e.g., using a visual tracking solution) with inertial data from the measurements to determine the positioning and movement of the XR system 200 relative to the physical world (e.g., the scene) and a map of the physical world. As described below, in some examples, when tracking the pose of the XR system 200, the device tracker may generate a three-dimensional (3D) map of the scene (e.g., the real world) and / or generate an update to the 3D map of the scene. The 3D map update may include, for example but not limited to, new or updated features and / or features or fiducial points associated with the scene and / or the 3D map of the scene, a positioning update that identifies or updates the positioning of the XR system 200 within the scene and the 3D map of the scene, etc. The 3D map may provide a digital representation of the scene in the real / physical world. In some examples, the 3D map may anchor location-based objects and / or content to real-world coordinates and / or objects. The XR system 200 may use the map-constructed scene (e.g., the scene in the physical world represented by and / or associated with the 3D map) to merge the physical and virtual worlds and / or to merge virtual content or objects with the physical environment.
[0066] In some aspects, computing component 210 may use a vision tracking solution to determine and / or track the pose of image sensor 202 and / or XR system 200 as a whole based on images captured by image sensor 202 (and / or other cameras of XR system 200). For example, in some examples, computing component 210 may use computer vision-based tracking, model-based tracking, and / or simultaneous localization and mapping (SLAM) techniques to perform tracking. For example, computing component 210 may perform SLAM or may communicate (wired or wirelessly) with a SLAM engine (not shown). SLAM refers to a class of techniques in which a camera (e.g., image sensor 202) and / or the pose of XR system 200 relative to a map of an environment (e.g., a map of the environment modeled by XR system 200) are tracked simultaneously while creating a map of the environment. This map may be referred to as a SLAM map and may be three-dimensional (3D). SLAM techniques may use color or grayscale image data captured by image sensor 202 (and / or other cameras of XR system 200) and may be used to generate an estimate of the 6DoF pose measurement of image sensor 202 and / or XR system 200. Such SLAM techniques configured to perform 6DoF tracking may be referred to as 6DoF SLAM. In some cases, the output of one or more sensors (e.g., accelerometer 204, gyroscope 206, one or more IMUs, and / or other sensors) may be used to estimate, correct, and / or otherwise adjust the estimated pose.
[0067] In some cases, 6DoF SLAM (e.g., 6DoF tracking) may associate features observed from certain input images from image sensor 202 (and / or other cameras) to the SLAM map. For example, 6DoF SLAM may use feature point associations from the input images to determine the pose (position and orientation) of image sensor 202 and / or XR system 200 for that input image. 6DoF mapping may also be performed to update the SLAM map. In some cases, the SLAM map maintained using 6DoF SLAM may contain 3D feature points triangulated from two or more images. For example, key frames may be selected from the input images or video stream to represent the observed scene. For each key frame, the corresponding 6DoF camera pose associated with the image may be determined. The pose of image sensor 202 and / or XR system 200 may be determined by projecting features from the 3D SLAM map into the image or video frame and updating the camera pose based on the verified 2D-3D correspondences.
[0068] In an illustrative example, computing component 210 may extract feature points from certain input images (e.g., each input image, a subset of input images, etc.) or from each key frame. Feature points (also referred to as registration points) as used herein are unique or identifiable portions of an image, such as a part of a hand, an edge of a table, and other examples. Features extracted from the captured images may represent different feature points along a three-dimensional space (e.g., coordinates on the X, Y, and Z axes), and each feature point may have an associated feature location. Feature points in a key frame match (are the same as or correspond to) or fail to match feature points of a previously captured input image or key frame. Feature detection may be used to detect feature points. Feature detection may include image processing operations for examining one or more pixels of an image to determine whether a feature exists at a particular pixel. Feature detection may be used to process the entire captured image or certain portions of the image. For each image or key frame, once features have been detected, local image patches around the feature may be extracted. Any suitable technique may be used to extract features, such as Scale-Invariant Feature Transform (SIFT) (which localizes features and generates their descriptions), Learned Invariant Feature Transform (LIFT), Speeded-Up Robust Features (SURF), Gradient Location-Orientation Histogram (GLOH), Oriented FAST and Rotated BRIEF (ORB), Binary Robust Invariant Scalable Keypoints (BRISK), Fast Retina Keypoint (FREAK), KAZE, Accelerated-KAZE (AKAZE), Normalized Cross-Correlation (NCC), descriptor matching, another suitable technique, or a combination thereof.
[0069] As an illustrative example, computing component 210 may extract feature points corresponding to a mobile device (e.g., Figure 2 an XR system of Figure 10A the HMD 1010 of Figure 11A the mobile device 1150 of), etc. In some cases, feature points corresponding to the mobile device may be tracked to determine the pose of the mobile device. As described in more detail below, the pose of the mobile device may be used to determine the position of a projection of AR media content that may enhance media content displayed on a display of the mobile device.
[0070] In some cases, XR system 200 may also track a user's hand and / or fingers to allow the user to interact with and / or control virtual content in a virtual environment. For example, XR system 200 may track the pose and / or movement of a user's hand and / or fingertips to identify or translate the user's interaction with the virtual environment. User interaction may include, for example but not limited to, moving a virtual content item, resizing a virtual content item, selecting an input interface element in a virtual user interface (e.g., a virtual representation of a mobile phone, a virtual keyboard, and / or other virtual interfaces), providing input through the virtual user interface, etc.
[0071] Figure 3 is a block diagram illustrating an example model generation system 300 in accordance with aspects of the present disclosure. The model generation system 300 provides a pipeline for closed object scanning. The model generation system 300 can be used as a stand-alone solution or integrated into an existing 3D scanning solution. As Figure 3 shown, the model generation system 300 includes an object tracking engine 306, a segmentation engine 308, and a model generation engine 310. As described in more detail below, the various components of the model generation system 300 can be used to perform object scanning by processing frames of an object (e.g., input frame 302) and generating one or more 3D models of the object.
[0072] For example, the object tracking engine 306 and the segmentation engine 308 can perform a tracking-based object segmentation process. The segmentation engine 308 segments the object from other objects (such as planes), thereby allowing the model generation engine 310 to generate a 3D model of the object (of the one or more output 3D models 314) without planes associated with the planar surface. Using the techniques described below, the model generation system 300 can detect irregular segmentation results, be robust to drift that may occur during tracking, and can recover from segmentation failures.
[0073] The model generation system 300 can be a single computing device or part of or implemented by multiple computing devices. In some examples, the model generation system 300 can include or be part of the following devices: a single electronic device, such as a mobile or cellular phone (e.g., a smart phone, a cellular phone, etc.); an XR device, such as an HMD or AR glasses; a camera system (e.g., a digital camera, an IP camera, a video camera, a security camera, etc.); a desktop computer; a laptop or notebook computer; a tablet computer; an Internet of Things (IoT) device; a set-top box; a television (e.g., a network or Internet-connected television) or other display device; a digital media player; a game console; a video streaming device; a drone or unmanned aerial vehicle; or any other suitable electronic device. In some examples, the model generation system 300 can include one or more wireless transceivers for wireless communication, such as cellular network communication, 802.11 wi-fi communication, wireless local area network (WLAN) communication, or some combination thereof. In some embodiments, the model generation system 300 can be implemented as Figure 12 part of the computing system 1200 as shown.
[0074] Although the model generation system 300 is shown as including certain components, those of ordinary skill in the art will understand that the model generation system 300 can include more than Figure 3The components shown and more components. The components of the model generation system 300 can include software, hardware, or one or more combinations of software and hardware. For example, in some specific implementations, the components of the model generation system 300 can include electronic circuits or other electronic hardware and / or can be implemented using electronic circuits or other electronic hardware, which can include one or more programmable electronic circuits (e.g., microprocessors, GPUs, DSPs, CPUs, and / or other suitable electronic circuits), and / or can include computer software, firmware, or any combination thereof and / or can be implemented using computer software, firmware, or any combination thereof to perform the various operations described herein. The software and / or firmware can include one or more instructions stored on a computer-readable storage medium and executable by one or more processors of the electronic device implementing the model generation system 300.
[0075] Although not shown in Figure 3 the model generation system 300 can include various computing components. The computing components can include, for example but not limited to, a central processing unit (CPU), a graphics processing unit (GPU), a digital signal processor (DSP) (such as a host processor or an application processor), and / or an image signal processor (ISP). In some cases, one or more computing components can include other electronic circuits or hardware, computer software, firmware, or any combination thereof to perform any of the various operations described herein. The computing components can also include computing device memory, such as read-only memory (ROM), random access memory (RAM), dynamic random access memory (DRAM), one or more cache memory devices (e.g., CPU cache or other cache components), and other memory components.
[0076] The model generation system 300 may also include one or more input / output (I / O) devices. The I / O devices may include a display screen, a keyboard, a keypad, a touch screen, a touchpad, a touch-sensitive surface, a printer, any other output device, any other input device, or any combination thereof. In some examples, the I / O devices may include one or more ports, jacks, or other connectors that implement a wired connection between the model generation system 300 and one or more peripheral devices, through which the system 300 may receive data from and / or send data to one or more peripheral devices. In some examples, the I / O devices may include one or more wireless transceivers that implement a wireless connection between the model generation system 300 and one or more peripheral devices, through which the system 300 may receive data from and / or send data to one or more peripheral devices. The peripheral devices may include any type of I / O device discussed previously, and once they are coupled to a port, jack, wireless transceiver, or other wired and / or wireless connector, they themselves may be considered I / O devices.
[0077] As Figure 3 shown, the input frame 302 is input into the model generation system 300. Each frame in the input frame 302 captures an object positioned on a surface in a scene. When the image capture device moves around the object, the image capture device may capture the input frame 302 from different angles during the image capture process. For example, when the input frame 302 is captured, the user may move the image capture device around the object.
[0078] Each frame includes a plurality of pixels, and each pixel corresponds to a set of pixel values, such as depth values, photometric values (e.g., red-green-blue (RGB) values, intensity values, chromaticity values, saturation values, etc.), or a combination thereof. In some examples, in addition to or instead of photometric values (e.g., RGB values), the input frame 302 may include depth information. For example, the input frame 302 may include a depth map (e.g., captured by a 3D sensor such as a depth sensor or a camera), a red-green-blue depth (RGB-D) frame or image, and other types of frames that include depth information. In addition to color and / or luminance information, RGB-D frames are also used to record depth information. In an illustrative example, a depth sensor may be used to capture multiple depth maps of an object from different angles. A depth map refers to an image or an image channel (e.g., the depth channel in an RGB-D frame) that contains information indicating the distance of the surface of an object in a scene from a viewpoint such as a camera.
[0079] Figure 4FIG. 0 is a diagram illustrating an example operation of an exemplary image capture device 402 capturing an input frame (e.g., input frame 302). In some cases, the image capture device can be part of system 300 or include part of the computing device of system 300. In some cases, the image capture device can be part of a separate computing device different from the computing device including system 300. As shown, the image capture device 402 moves around or about an object 410 (a cup) on a planar surface 411 along a path 404 (e.g., shown as an arc). During the movement of the image capture device 402 along the path 404, the image capture device 402 is located at various positions, which are illustrated as camera poses 406A, 406B, 406C, 406D, 406E, and 406F in Figure 4 As should be noted, Figure 4 the number, spacing, and orientation of the camera poses 406A through 406F shown in
[0080] are shown for illustrative purposes only and should not be considered restrictive. For example, more camera poses or fewer camera poses can be used.
[0081] Based on how the image capture device 402 moves around the object 410, the path 404 can have any configuration. In some examples, as the image capture device 402 moves along the path 404 from the position associated with camera pose 406A to the position associated with camera pose 406F, various frames of the object can be captured and used as the input frame 302. For example, at camera pose 406A (which represents the initial camera pose of the image capture device 402 at the first position along the path 404), the image capture device 402 can capture a first frame. As the image capture device 402 continues to move along the path 404, additional frames can be captured. In some examples, frames can be captured continuously at the frame rate of the image capture device. For example, if the frame rate of the image capture device is 30 frames per second (fps), the image capture device can capture 30 frames every 1 second of time. Then the input frame 302 can be provided to the model generation system 300.
[0082] In some examples, all input frames 302 are used to perform object tracking, while the model generation engine 310 uses only certain frames (referred to herein as key frames) to generate a 3D model of the objects in the input frames. For example, when capturing frames of an object, an image capture device that captures frames at a frame rate of 30 fps or any other frame rate can be used to scan the object.
[0083] One or more of the captured frames can be designated as key frames. The frames other than the key frames in the plurality of frames can be non-key frames. In some examples, non-key frames can be used for segmentation based on tracking and / or object detection, but not for 3D model generation. In some cases, not every frame is available for reconstruction because there are differences between one frame and the next, and at the full frame rate of the camera (e.g., 30 fps to 60 fps), the features of the environment typically do not change enough to justify a full reconstruction. The number of key frames in the plurality of captured frames can affect the reconstruction quality (e.g., how detailed the reconstructed virtual object is, how accurately the reconstructed virtual object matches the real object, etc.) at the cost of computing resources (e.g., load, time, power, etc.). For example, using a relatively large number of key frames can result in an accurate and detailed reconstruction while using a relatively large amount of computing resources. Similarly, using a relatively small number of key frames for reconstruction can minimize the amount of computing resources used, but can result in a lower quality reconstruction that may not be usable. For example, there may not be enough object angles and / or perspectives to obtain an acceptable amount of detail, and features of the object may be missed. Therefore, it may be useful to optimize the number of key frames to obtain a high-quality reconstruction while balancing the amount of computing resources used.
[0084] In some cases, the optimal number of key frames can vary based on the scale (e.g., size) of the object being reconstructed. For example, the level of detail used to reconstruct a smaller object may be relatively high because interactions with smaller objects can be performed at a closer range compared to larger objects. In some cases, the number of key frames selected for reconstructing an object can be described based on an overlap rate.
[0085] Figure 5A , Figure 5B , Figure 5C and Figure 5D illustrate examples of overlap rates in accordance with aspects of the present disclosure. In Figures 5A to 5D , each triangle represents the frustum of a camera (not shown) used to capture the scene. The frustum of the image represents the portion of the environment (e.g., scene) that is visible in the image. Figure 5AIncludes two frustums from two different positions: 502 and 504. As shown, a portion of frustum 502 and frustum 504 overlap (e.g., the same portion of the environment appears in the images captured with frustum 502 and frustum 504). The overlap rate refers to the amount of the overlapping portion 506 of the frustums.
[0086] As Figures 5B to 5D shown, in some cases, the overlap rate may refer to variable X, which represents the amount of overlap 520 between the frustums in the case where all frames of the video are used for reconstruction. Thus, Figure 5B the overlap rate shown is X. Figure 5C Illustrates that half of the frames of the video (e.g., every other frame) are used for reconstruction and thus have an overlap rate of X / 2 530. Similarly, Figure 5D illustrates that one out of four frames is used for reconstruction and thus has an overlap rate of X / 4 540. Thus, the lower the overlap rate, the fewer key frames are used. Thus, optimizing the overlap rate can be used to adjust the amount of key frames used for reconstruction. Dynamically determining the overlap rate during image capture allows for the real-time selection and adjustment of the number of selected key frames used to reconstruct an object when capturing an image of the object being reconstructed. It is worth noting that Figure 5C and Figure 5D only show the frames used for reconstruction (e.g., the frustums of the key frames), and omit the frames not used for reconstruction. For Figure 5C and Figure 5D , frames (both frames used for reconstruction and frames not used for reconstruction) are captured at the rate shown in Figure 5B (e.g., based on the frame rate of the camera).
[0087] In some cases, the overlap rate can be directly specified. For example, if the overlap rate is specified as 50%, then once the camera has moved sufficiently such that 50% of the threshold of the frustum captured in the current image was not captured in the previous image, a key frame can be captured. In some cases, when the overlap rate is directly specified, the overlap rate can be independent of the frame rate. That is, the overlap rate is determined by the change in the viewing angle rather than the change in the frame rate. Thus, as long as the camera (relative to the environment) does not move, no additional key frames will be generated, regardless of how many frames are captured while the camera is stationary. If the camera then moves by a sufficient angle such that half of the view is new, additional key frames can be generated. In some cases, it can be determined based on the pose of the camera (e.g., 6DoF (or 3DoF) data) whether the camera has moved sufficiently such that additional key frames should be generated.
[0088] As indicated above, the optimal number of key frames may vary based on the scale of the object being reconstructed, since a constant overlap rate may produce a good quality reconstruction for an object of one size while providing a lower quality reconstruction for another object of a different size. For example, the level of detail suitable for an object sized to be held in the hand may be higher than the level of detail suitable for another object approximately the size of a table. Additionally, reducing the overlap rate (e.g., from 50% to 25%) has a greater impact on the reconstruction of smaller objects than on the reconstruction of larger objects.
[0089] To help balance computational resources and reconstruction quality, the overlap rate (and thus the number of key frames) used to reconstruct an object may vary based on the size of the object. To help account for different object sizes, the sizes may be divided into size categories such as small-sized objects, medium-sized objects, and large-sized objects. In some cases, a small-sized object may be an object having a size that can be approximated as being held in the hand (e.g., approximately 5 cm to 30 cm), a medium-sized object may be an object larger than a small-sized object up to approximately the size of a table (e.g., approximately 2 m), and a large-sized object may be a larger object such as a room, a car, a house, etc. It should be understood that the categories of small, medium, and large objects are merely examples, and fewer categories or additional categories may be defined. In some cases, the number of categories and the sizes of the objects within the categories may be determined experimentally based on, for example, how different overlap rates affect the reconstruction quality of objects of a particular size.
[0090] Figure 6 is a top view of a room 600 according to aspects of the present disclosure, which illustrates scanning an object for reconstruction. In some cases, how an object is scanned for reconstruction may indicate the size of the object being scanned. Generally, to scan an object, multiple images of the object may be captured at different angles relative to the object. In some cases, this may be performed by moving a scanning device (such as a camera) around the object being scanned. For example, to scan a medium-sized object 602 such as a laptop computer, the camera may move around the medium-sized object 602 along a first arc 604. The first arc may be associated with a first curvature.
[0091] Then, the camera can move around the small-sized object 606 along the second arc 608. In some cases, when attempting to scan a smaller object for reconstruction, the camera can move closer compared to a medium-sized object or a large-sized object to better capture details of the smaller object. In some cases, hints or instructions can also be used to encourage the user of the camera to position the camera such that the object being scanned fills a particular portion of the image (e.g., as displayed and captured on the screen). Thus, in some cases, the second arc 608 around the small-sized object 606 can have a greater curvature (e.g., a smaller radius) than the first curvature of the first arc 604. In some cases, the scanning distance around the medium-sized object 602 is approximately constant, and the camera points towards the interior of the arc. Similarly, the scanning distance around the small-sized object 606 is approximately constant, and the camera points towards the interior of the arc. In some cases, the camera can prompt the user of the camera to maintain the scanning distance during scanning, for example, via a user interface. In some cases, when attempting to scan a large-sized object (such as the room 600) for reconstruction, the camera can move around the user of the camera such that the camera moves along the third arc 610, where the camera points towards the exterior of the arc. Based on the observations regarding how scanning can be performed as discussed with respect to Figure 6 the observations of how scanning can be performed, techniques for optimizing the overlap rate can be designed based on the scanning use case.
[0092] Figure 7 is a block diagram illustrating the architecture of an example overlap rate adjustment system 700 for optimizing the overlap rate according to aspects of the present disclosure. In some cases, the overlap rate adjustment system 700 can be part of or a component of a model generation system (such as the model generation system 300). In the overlap rate adjustment system 700, the pose collector 702 can collect pose data 704 regarding the pose of the camera. The pose data 704 can indicate the viewpoint of the camera (e.g., direction / angle / orientation), which indicates where the camera is facing to capture an image. In some cases, as discussed above with respect to Figure 2 the pose data 704 can be determined based on the images captured by the camera and / or the outputs of one or more sensors.
[0093] The pose collector 702 can collect a set of N viewpoints and pass the set of the N viewpoints to the arc fitting engine 706. In some cases, N can be determined in advance. In some cases, N can be determined in advance based on the number of points used by the arc fitting engine 706 to fit an arc. The arc fitting engine 706 attempts to fit an arc to the viewpoints. For example, the viewpoints may include relative position information indicating how the position of a viewpoint changes relative to other viewpoints, and a path between these positions can be drawn. In some cases, the arc fitting engine 706 attempts to find a circle that is most similar to the path between the N viewpoints. For example, the arc fitting engine 706 can apply the Taubin curve fitting algorithm to fit a circle to the viewpoints. Other arc or curve fitting techniques can also be used. If the arc fitting engine 706 finds a fitted circle, the center and radius of the fitted circle can be found. This information about the fitted circle, together with the set of N viewpoints, can be passed to the view direction fitting engine 708. If the arc fitting engine 706 cannot fit a circle to the path between the N viewpoints (for example, if the camera moves in a straight line), the set of the N viewpoints can be discarded, and the overlap rate can be set to the default overlap rate or the large-size object overlap rate.
[0094] The view direction fitting engine 708 can attempt to determine an indication of the camera direction. For example, the view direction fitting engine 708 can determine whether the camera is pointing inward towards the center of the fitted circle or outward away from the fitted circle. The view direction fitting engine 708 can pass the indication of the camera direction, together with the information about the fitted circle, to the use case prediction engine 710.
[0095] In some cases, the use case prediction engine 710 may predict the overlap rate of the output 712. In some cases, the use case prediction engine 710 may predict the overlap rate based on, for example, an indication of the camera orientation and the characteristics of the fitted circle. In some cases, the characteristics of the fitted circle may include the radius of the fitted circle, the diameter of the fitted circle, the center of the fitted circle, etc. For example, if the radius of the fitted circle is within a first length range (e.g., radius), such as from 0 meters to N meters, and the camera is pointing inward (e.g., pointing in the expected direction, which is inward here), then the use case prediction engine 710 may predict that the camera is attempting to reconstruct a small-sized object. The use case prediction engine 710 may output 712 an overlap rate R suitable for such small-sized objects. Similarly, if the radius of the fitted circle is within a second length range, such as from N meters to M meters, and the camera is pointing inward, then the use case prediction engine 710 may predict that the camera is attempting to reconstruct a medium-sized object. The use case prediction engine 710 may output 712 an overlap rate S suitable for such medium-sized objects (e.g., an overlap rate where S < R). As another example, if the radius of the fitted circle is within a third length range (or any length), such as from 0 meters to M meters, and the camera is pointing outward, then the use case prediction engine 710 may predict that the camera is attempting to reconstruct a large-sized object. In some cases, when the camera is pointing outward from the fitted circle, the use case prediction engine 710 may predict that the camera is attempting to reconstruct a large-sized object. The use case prediction engine 710 may output 712 an overlap rate T suitable for large-sized objects (e.g., an overlap rate where T < S). In some cases, the length ranges, the expected directions of the camera pointing, and the corresponding overlap rates are predetermined and stored in, for example, a table.
[0096] In some cases, after the overlap rate is output 712, the distance measurement engine 714 may estimate whether the distance to the scanned object deviates to determine whether the scanning of the object has stopped or whether a new arc may need to be fitted. For example, the distance measurement engine 714 may receive the pose data 704. The distance measurement engine 714 may also receive information about the fitted circle. In some cases, the pose data 704 may be separate from the set of N viewpoints collected from the pose collector 702. Based on the pose data 704, the distance measurement engine 714 may determine whether the distance between the camera and the center of the fitted circle deviates (e.g., increases or decreases). If the distance deviation exceeds a threshold amount, then the distance measurement engine 714 may cause the arc fitting engine 706 to attempt to refit the arc to redetermine the overlap rate.
[0097] Figure 8A and Figure 8B Illustrates a set of viewpoints on a fitted circle according to aspects of the present disclosure. As indicated above, the view direction fitting engine 708 may determine whether the camera is pointing inward towards the center of the fitted circle. As an example of this determination, the view direction fitting engine 708 may receive information about the fitted circle, including the center 802 of the fitted circle and the set of N viewpoints. Figure 8AShows a set of N viewpoints drawn along a fitted circle 804. A view direction fitting engine 708 can determine a center vector 808 from a viewpoint 806A to the center 802 of the fitted circle 804. Then, the direction of the determined center vector 808 can be compared with the direction 810 corresponding to the pose of the viewpoint 806A to determine the angle Ɵ between the direction of the determined center vector 808 and the direction 810 of the pose. If Ɵ < 90°, it can be determined that the camera is pointing inward.
[0098] Similarly, Figure 8B Shows a set of N viewpoints drawn along a fitted circle 854. In the fitted circle 854, a center vector 858 from a viewpoint 856A to the center 852 of the fitted circle 854. Then, the direction of the determined center vector 858 can be compared with the direction 860 corresponding to the pose of the viewpoint 856A to determine the angle . In Figure 8B , because > 90° and < 180, it can be determined that the camera is pointing outward.
[0099] Figure 9 Is a flowchart illustrating a process 900 for image processing according to aspects of the present disclosure. The process 900 can be executed by a computing device (or apparatus) or a component of a computing device (e.g., a chipset, codec, etc.) (e.g., Figure 1 's image processor 150, Figure 2 's computing component 210, Figure 12 's processor 1210, or other computing device). The computing device can be a mobile device (e.g., a mobile phone), a network-connected wearable device such as a watch, an extended reality (XR) device such as a virtual reality (VR) device or an augmented reality (AR) device (e.g., Figure 10A , Figure 10B 's HMD 1010, Figure 11A , Figure 11B 's mobile phone 1150), a vehicle or a component or system of a vehicle, or other types of computing devices. The operations of the process 900 can be implemented as software components executed and run on one or more processors (e.g., Figure 1 's image processor 150, Figure 2 's CPU 212, GPU 214, DSP 216, and / or ISP 218, Figure 12 's processor 1210 or other processors). In some cases, the operations of the process 900 can be implemented by Figure 12 's computing system 1200.
[0100] At block 902, a computing device (or its components) may obtain pose information of an image sensor (e.g., Figure 1 image capture device 105A of Figure 2 image sensor 202 of Figure 10A , Figure 10B cameras 1030A, 1030B of Figure 11A , Figure 11B cameras 1130A, 1130B of ). In some cases, the pose information indicates a set of viewpoints of the environment (e.g., Figure 8A , Figure 8B viewpoint 806 of ). In some cases, the image sensor is configured to capture frames for reconstructing an object. In some cases, the viewpoints in the set of viewpoints include position information and orientation information of the image sensor. In some cases, the pose information includes six degrees of freedom (6DOF) information.
[0101] At block 904, the computing device (or its components) may fit a curve to the set of viewpoints. In some cases, to fit the curve, the computing device (or its components) may determine a path for the set of viewpoints based on the position information. In some cases, to fit the curve, the computing device (or its components) may apply a curve fitting algorithm to the path.
[0102] At block 906, the computing device (or its components) may determine one or more characteristics of the curve. In some cases, to determine one or more characteristics of the curve, the computing device (or its components) may determine the center of the curve and the radius of the curve. In some cases, the orientation of the image sensor is determined based on the determined center of the curve. In some cases, the overlap rate is determined based on the radius of the curve.
[0103] At block 908, the computing device (or its components) may determine the orientation of the image sensor based on the one or more characteristics of the curve and the set of viewpoints. In some cases, to determine the orientation of the image sensor, the computing device (or its components) may determine a vector between the position of the viewpoints around the curve and the center of the curve. In some cases, to determine the orientation of the image sensor, the computing device (or its components) may compare the vector with the orientation associated with the viewpoints.
[0104] At block 910, the computing device (or its component) may determine an overlap rate based on the one or more features of the curve and the determined orientation of the image sensor. In some cases, to determine the overlap rate, the computing device (or its component) may determine that the radius of the curve is within a radius range. In some cases, to determine the overlap rate, the computing device (or its component) may determine that the orientation of the image sensor is pointed in an expected direction. In some cases, the overlap rate is determined based on a set of size categories. In some cases, the overlap rate is pre-determined for the size categories in the set of size categories. In some cases, the number of viewpoints in the set of viewpoints is pre-determined.
[0105] At block 912, the computing device (or its component) may output the determined overlap rate to select a frame to be used for reconstructing an object. The computing device (or its component) may obtain additional pose information of the image sensor. The computing device (or its component) may determine that the distance from the center of the curve to the viewpoint indicated by the additional pose information has deviated from the radius of the curve by more than a threshold amount. The computing device (or its component) may re-design another curve based on the viewpoint. Figure 10A FIG. 1000 is a perspective view illustrating a head-mounted display (HMD) 1010 that performs feature tracking and / or visual simultaneous localization and mapping (VSLAM) according to some examples. The HMD 1010 may be, for example, an augmented reality (AR) head-mounted headset, a virtual reality (VR) head-mounted headset, a mixed reality (MR) head-mounted headset, an extended reality (XR) head-mounted headset, or some combination thereof. The HMD 1010 may be an example of an XR system 200, a model generation system 300, or a combination thereof. The HMD 1010 includes a first camera 1030A and a second camera 1030B along the front of the HMD 1010. The first camera 1030A and the second camera 1030B may be two cameras among one or more cameras. In some examples, the HMD 1010 may have only a single camera. In some examples, in addition to the first camera 1030A and the second camera 1030B, the HMD 1010 may further include one or more additional cameras. In some examples, in addition to the first camera 1030A and the second camera 1030B, the HMD 1010 may further include one or more additional sensors.
[0106] Figure 10B illustrates according to some examples Figure 10APerspective view 1030 of a head-mounted display (HMD) 1010 being worn by a user 1020. The user 1020 wears the HMD 1010 on the user 1020's head, above the user 1020's eyes. The HMD 1010 can capture images using a first camera 1030A and a second camera 1030B. In some examples, the HMD 1010 displays one or more display images based on the images captured by the first camera 1030A and the second camera 1030B towards the user 1020's eyes. The display images can provide a stereoscopic view of the environment, in some cases with overlaid information and / or with other modifications. For example, the HMD 1010 can display a first display image to the user 1020's right eye, the first display image being based on the image captured by the first camera 1030A. The HMD 1010 can display a second display image to the user 1020's left eye, the second display image being based on the image captured by the second camera 1030B. For example, the HMD 1010 can provide overlay information in the display image that is overlaid on the images captured by the first camera 1030A and the second camera 1030B.
[0107] The HMD 1010 may not include wheels, propellers, or other conveyance means of its own. Instead, the HMD 1010 relies on the movement of the user 1020 to move the HMD 1010 back and forth in the environment. In some cases, such as when the HMD 1010 is a VR headset, the environment can be fully or partially virtual. If the environment is at least partially virtual, the movement through the virtual environment can also be virtual. For example, the movement through the virtual environment can be controlled by an input device 208. The movement actuator can include any such input device 208. The movement through the virtual environment may not require wheels, propellers, legs, or any other form of conveyance means. Even if the environment is virtual, SLAM technology can still be valuable because the virtual environment can be deconstructed and / or generated by a device other than the HMD 1010, such as a remote server or console associated with a video game or video game platform. In some cases, feature tracking and / or SLAM can even be performed by a vehicle or other device in the virtual environment, the vehicle or other device having its own physical conveyance system that allows it to move physically back and forth in the physical environment. For example, SLAM can be performed in the virtual environment to test whether the model generation system 300 is working properly without wasting time or energy during movement and without wearing out the physical conveyance system.
[0108] Figure 11AFIG. 1100 is a perspective view of a front surface 1155 of a mobile device 1150 that illustrates using one or more front cameras 1130A-B to perform features described herein, including, for example, feature tracking and / or visual simultaneous localization and mapping (VSLAM). The mobile device 1150 can be, for example, a cellular phone, a satellite phone, a portable game console, a music player, a health tracking device, a wearable device, a wireless communication device, a laptop computer, a mobile device, any other type of computing device or computing system 1300 discussed herein, or a combination thereof. The front surface 1155 of the mobile device 1150 includes a display screen 1145. The front surface 1155 of the mobile device 1150 includes a first camera 1130A and a second camera 1130B. The first camera 1130A and the second camera 1130B are illustrated in a bezel around the display screen 1145 on the front surface 1155 of the mobile device 1150. In some examples, the first camera 1130A and the second camera 1130B can be positioned in a notch or cutout cut from the display screen 1145 on the front surface 1155 of the mobile device 1150. In some examples, the first camera 1130A and the second camera 1130B can be under-display cameras located between the display screen 1145 and the remainder of the mobile device 1150 such that light passes through a portion of the display screen 1145 before reaching the first camera 1130A and the second camera 1130B. The first camera 1130A and the second camera 1130B of the perspective view 1100 are front cameras. The first camera 1130A and the second camera 1130B face in a direction perpendicular to a planar surface of the front surface 1155 of the mobile device 1150. The first camera 1130A and the second camera 1130B can be two of the one or more cameras. In some examples, the front surface 1155 of the mobile device 1150 can have only a single camera. In some examples, in addition to the first camera 1130A and the second camera 1130B, the mobile device 1150 can include one or more additional cameras. In some examples, in addition to the first camera 1130A and the second camera 1130B, the mobile device 1150 can include one or more additional sensors.
[0109] Figure 11BFIG. 1130 is a perspective view of the rear surface 1165 of the exemplary mobile device 1150. The mobile device 1150 includes a third camera 1130C and a fourth camera 1130D on the rear surface 1165 of the mobile device 1150. The third camera 1130C and the fourth camera 1130D of the perspective view 1190 are rear-facing. The third camera 1130C and the fourth camera 1130D are oriented in a direction perpendicular to the planar surface of the rear surface 1165 of the mobile device 1150. Although the rear surface 1165 of the mobile device 1150 does not have a display screen 1145 as illustrated in the perspective view 1190, in some examples, the rear surface 1165 of the mobile device 1150 may have a second display screen. If the rear surface 1165 of the mobile device 1150 has a display screen 1145, any positioning of the third camera 1130C and the fourth camera 1130D with respect to the display screen 1145 may be used, as discussed with respect to the first camera 1130A and the second camera 1130B at the front surface 1155 of the mobile device 1150. The third camera 1130C and the fourth camera 1130D may be two of the one or more cameras. In some examples, the rear surface 1165 of the mobile device 1150 may have only a single camera. In some examples, in addition to the first camera 1130A, the second camera 1130B, the third camera 1130C, and the fourth camera 1130D, the mobile device 1150 may further include one or more additional cameras. In some examples, in addition to the first camera 1130A, the second camera 1130B, the third camera 1130C, and the fourth camera 1130D, the mobile device 1150 may further include one or more additional sensors.
[0110] Like the HMD 1010, the mobile device 1150 does not include wheels, propellers, or other conveyance means of its own. Instead, the mobile device 1150 relies on the movement of the user who holds or wears the mobile device 1150 to move the mobile device 1150 back and forth in the environment. In some cases, such as when the mobile device 1150 is used for AR, VR, MR, or XR, the environment may be fully or partially virtual. In some cases, the mobile device 1150 may be inserted into a head-mounted device (HMD) (e.g., inserted into a cradle of the HMD) such that the mobile device 1150 serves as a display of the HMD, where the display screen 1145 of the mobile device 1150 serves as the display of the HMD. If the environment is at least partially virtual, the movement through the virtual environment may also be virtual. For example, the movement through the virtual environment may be controlled by one or more joysticks, buttons, video game controllers, mice, keyboards, touchpads, and / or other input devices coupled to the mobile device 1150 in a wired or wireless manner.
[0111] Figure 12FIG. is an example diagram illustrating a system for implementing some aspects of the present technology. Specifically, Figure 12 An example of a computing system 1200 is illustrated, which can be any computing device that constitutes an internal computing system, a remote computing system, a camera, or any component thereof. Components of the system communicate with each other using connection 1205. Connection 1205 can be a physical connection using a bus or a direct connection to the processor 1210, such as in a chipset architecture. Connection 1205 can also be a virtual connection, a networking connection, or a logical connection.
[0112] In some examples, the computing system 1200 is a distributed system, where the functions described in the present disclosure can be distributed within a data center, multiple data centers, a peer-to-peer network, etc. In some examples, one or more of the described system components represent many such components, each component performing some or all of the functions that the component is described for. In some cases, these components can be physical or virtual devices.
[0113] The example computing system 1200 includes at least one processing unit (CPU or processor) 1210 and a connection 1205 that couples various system components including a system memory 1215 such as a read-only memory (ROM) 1220 and a random access memory (RAM) 1225 to the processor 1210. The computing system 1200 can include a cache 1212 that is directly connected to, in close proximity to, or integrated as part of the processor 1210 and is a high-speed memory.
[0114] The processor 1210 can include any general-purpose processor and hardware services or software services such as services 1232, 1234, and 1236 stored in a storage device 1230, which are configured to control the processor 1210 and a dedicated processor in which software instructions are incorporated into the actual processor design. The processor 1210 can be a fully independent computing system that includes multiple cores or processors, buses, memory controllers, caches, etc. The multi-core processor can be symmetric or asymmetric.
[0115] To enable user interaction, computing system 1200 includes an input device 1245 that can represent any number of input mechanisms, such as a microphone for voice, a touch-sensitive screen for gesture or graphical input, a keyboard, a mouse, motion input, speech, a camera, an accelerometer, a gyroscope, and the like. Computing system 1200 may also include an output device 1235 that can be one or more of a plurality of output mechanisms. In some instances, a multimodal system may enable a user to provide multiple types of input / output to communicate with computing system 1200. Computing system 1200 may include a communication interface 1240 that generally may govern and manage user input and system output. The communication interface may perform or facilitate receiving and / or sending wired or wireless communications using wired and / or wireless transceivers, including using audio jack / plug, microphone jack / plug, universal serial bus (USB) port / plug, Apple ® Lightning ® port / plug, Ethernet port / plug, fiber optic port / plug, proprietary wired port / plug, Bluetooth ® wireless signaling, Bluetooth ® low energy (BLE) wireless signaling, iBeacon ® wireless signaling, radio frequency identification (RFID) wireless signaling, near field communication (NFC) wireless signaling, dedicated short range communication (DSRC) wireless signaling, 802.10 Wi-Fi wireless signaling, wireless local area network (WLAN) signaling, visible light communication (VLC), worldwide interoperability for microwave access (WiMAX), infrared (IR) communication wireless signaling, public switched telephone network (PSTN) signaling, integrated services digital network (ISDN) signaling, 3G / 4G / 5G / LTE cellular data network wireless signaling, ad hoc network signaling, radio wave signaling, microwave signaling, infrared signaling, visible light signaling, ultraviolet light signaling, wireless signaling along the electromagnetic spectrum, or some combination thereof. The communication interface 1240 may also include one or more global navigation satellite system (GNSS) receivers or transceivers for determining the location of computing system 1200 based on one or more signals received from one or more satellites associated with one or more GNSS systems. GNSS systems include, but are not limited to, the United States' Global Positioning System (GPS), Russia's Global Navigation Satellite System (GLONASS), China's BeiDou Navigation Satellite System (BDS), and Europe's Galileo GNSS. There are no restrictions on operating on any particular hardware arrangement, and thus the underlying features here can be easily replaced to obtain improved hardware or firmware arrangements as they are developed.
[0116] The storage device 1230 can be a non-volatile and / or non-transitory and / or computer-readable memory device, and can be a hard disk or other types of computer-readable media that can store data accessible by a computer, such as a tape cassette, a flash memory card, a solid-state memory device, a digital versatile disc, a cassette tape, a floppy disk, a flexible disk, a hard disk, a magnetic tape, a magnetic stripe / strip, any other magnetic storage media, flash memory, memristor memory, any other solid-state memory, a compact disc read-only memory (CD-ROM) optical disc, a rewritable compact disc (CD) optical disc, a digital video disc (DVD) optical disc, a Blu-ray disc (BDD) optical disc, a holographic optical disc, another optical media, a Secure Digital (SD) card, a micro Secure Digital (microSD) card, a Memory Stick ® card, a smart card chip, an EMV chip, a subscriber identity module (SIM) card, a mini / micro / nano / pico SIM card, another integrated circuit (IC) chip / card, a random access memory (RAM), a static RAM (SRAM), a dynamic RAM (DRAM), a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a flash EPROM (FLASHEPROM), a cache memory (L1 / L2 / L3 / L4 / L5 / L#), a resistive random access memory (RRAM / ReRAM), a phase change memory (PCM), a spin transfer torque RAM (STT-RAM), another memory chip or cartridge and / or a combination thereof.
[0117] The storage device 1230 can include software services, servers, services, etc., and when the code defining such software is executed by the processor 1210, the code causes the system to perform functions. In some examples, the hardware services that perform specific functions can include software components stored in a computer-readable medium connected to the necessary hardware components (such as the processor 1210, the connection 1205, the output device 1235, etc.) to perform functions.
[0118] As used herein, the term "computer-readable medium" includes, but is not limited to, portable or non-portable storage devices, optical storage devices, and various other media capable of storing, containing, or carrying instructions and / or data. The computer-readable medium may include non-transitory media in which data can be stored and which do not include carrier waves and / or transient electronic signals propagated wirelessly or over a wired connection. Examples of non-transitory media may include, but are not limited to, magnetic disks or tapes, optical storage media such as compact discs (CDs) or digital versatile discs (DVDs), flash memory, memory, or memory devices. The computer-readable medium may have code and / or machine-executable instructions stored thereon, which may represent procedures, functions, subroutines, programs, routines, subroutines, modules, software packages, classes, or any combination of instructions, data structures, or program statements. A code segment may be coupled to another code segment or a hardware circuit by passing and / or receiving information, data, arguments, parameters, or memory contents. The information, arguments, parameters, data, etc. may be transferred, forwarded, or transmitted using any suitable means, including memory sharing, message passing, token passing, network transmission, etc.
[0119] In some examples, computer-readable storage devices, media, and memories may include cables or wireless signals that contain bitstreams, etc. However, when mentioned, non-transitory computer-readable storage media specifically exclude media such as power consumption, carrier signals, electromagnetic waves, and signals themselves.
[0120] Specific details are provided in the above description to provide a thorough understanding of the examples provided herein. However, those of ordinary skill in the art will understand that the examples may be implemented without these specific details. For clarity, in some cases, the present technology may be presented as including separate functional blocks, including functional blocks that include devices, device components, steps or routines in a method embodied in software or a combination of hardware and software. Additional components other than those shown in the figures and / or described herein may be used. For example, circuits, systems, networks, processes, and other components may be shown as components in block diagram form so as not to obscure the examples with unnecessary details. In other cases, well-known circuits, processes, algorithms, structures, and techniques may be shown without unnecessary details so as to avoid obscuring the examples.
[0121] The above may describe various examples as processes or methods, which are depicted as flowcharts, process flow diagrams, data flow diagrams, structure diagrams, or block diagrams. Although a flowchart may describe operations as a sequential process, many of the operations in the operations can be performed in parallel or concurrently. In addition, the order of the operations can be rearranged. When the operations of a process are completed, the process is terminated, but the process may have additional steps not included in the figures. A process may correspond to a method, function, procedure, subroutine, subprogram, etc. When a process corresponds to a function, the termination of the process may correspond to the function returning to the calling function or the main function.
[0122] The processes and methods according to the above examples can be implemented using computer-executable instructions stored or otherwise obtained from a computer-readable medium. Such instructions can include, for example, instructions and data that cause or otherwise configure a general-purpose computer, a special-purpose computer, or a processing device to perform a certain function or group of functions. Part of the computer resources used can be accessed through a network. The computer-executable instructions can be, for example, binary, intermediate format instructions, such as assembly language, firmware, source code, etc. Examples of computer-readable media that can be used to store instructions, the information used, and / or the information created during the methods according to the described examples include magnetic disks or optical disks, flash memory, USB devices with non-volatile memory, networked storage devices, etc.
[0123] Devices implementing the processes and methods according to these disclosures can include hardware, software, firmware, middleware, microcode, hardware description language, or any combination thereof, and can take any of a variety of form factors. When implemented in software, firmware, middleware, or microcode, the program code or code segments (e.g., a computer program product) for performing the necessary tasks can be stored in a computer-readable or machine-readable medium. The processor can execute the necessary tasks. Typical examples of form factors include laptop computers, smartphones, mobile phones, tablet devices, or other small form factor personal computers, personal digital assistants, rack-mounted devices, stand-alone devices, etc. The functionality described herein can also be embodied in peripheral devices or plug-in cards. By additional example, such functionality can also be implemented on a circuit board among different chips or different processes executed on a single device.
[0124] Instructions, the media for conveying such instructions, the computing resources for executing them, and other structures for supporting such computing resources are example components for providing the functions described in this disclosure.
[0125] In the foregoing description, aspects of the present application have been described with reference to specific examples of the present application, but those skilled in the art will recognize that the present application is not limited thereto. Thus, although illustrative examples of the present application have been described in detail herein, it should be understood that the inventive concept may be implemented and adopted in various other ways, and the appended claims are intended to be construed to include such variations, unless limited by the prior art. The various features and aspects of the foregoing application may be used singly or in combination. Further, the examples may be utilized in any number of environments and applications other than those described herein, without departing from the broader spirit and scope of the specification. Accordingly, the specification and drawings are to be regarded as illustrative rather than restrictive. For purposes of illustration, the methods are described in a particular order. It should be understood that in alternative examples, the methods may be performed in an order different from that described.
[0126] One of ordinary skill in the art should understand that, without departing from the scope of this specification, the less than (“<”) and greater than (“>”) symbols or terms used herein may be replaced, respectively, with the less than or equal to (“ ”) and greater than or equal to (“ ”) symbols.
[0127] In cases where a component is described as “configured to” perform certain operations, such a configuration may be implemented, for example, by designing electronic circuitry or other hardware to perform the operations, by programming a programmable electronic circuit (e.g., a microprocessor or other suitable electronic circuit) to perform the operations, or any combination thereof.
[0128] The phrase “coupled to” means that any component is directly or indirectly physically connected to another component, and / or any component directly or indirectly communicates with another component (e.g., is connected to another component via a wired or wireless connection and / or other suitable communication interface).
[0129] Claim language or other language reciting “at least one of” a set and / or “one or more” of a set indicates that one member of the set or multiple members of the set (in any combination) satisfy the claim. For example, claim language reciting “at least one of A and B” means A, B, or A and B. In another example, claim language reciting “at least one of A, B, and C” means A, B, C, or A and B, or A and C, or B and C, or A and B and C. The language “at least one of” a set and / or “one or more” of a set does not limit the set to the items listed in the set. For example, claim language reciting “at least one of A and B” may mean A, B, or A and B, and may additionally include items not listed in the set of A and B.
[0130] The various illustrative logical blocks, modules, circuits, and algorithmic steps described in connection with the examples disclosed herein may be implemented in electronic hardware, computer software, firmware, or any combination thereof. To clearly illustrate this interchangeability of hardware and software, various illustrative components, blocks, modules, circuits, and steps have been described above generally in terms of their functionality. Whether such functionality is implemented as hardware or software depends upon the particular application and design constraints imposed on the overall system. Skilled artisans may implement the described functionality in varying ways for each particular application, but such implementation decisions should not be interpreted as causing a departure from the scope of the present application.
[0131] The techniques described herein may also be implemented in electronic hardware, computer software, firmware, or any combination thereof. Such techniques may be implemented in any of a variety of devices, such as a general-purpose computer, a cellular telephone as a wireless communication device, or an integrated circuit device having multiple uses, including applications in cellular telephones and other devices. Any features described as modules or components may be implemented together in an integrated logic device or separately as discrete but interoperable logic devices. If implemented in software, the techniques may be at least partially implemented by a computer-readable data storage medium comprising program code, including instructions that, when executed, perform one or more of the methods described above. The computer-readable data storage medium may form part of a computer program product, which may include packaging materials. The computer-readable medium may include memory or data storage media, such as random access memory (RAM) (such as synchronous dynamic random access memory (SDRAM)), read-only memory (ROM), non-volatile random access memory (NVRAM), electrically erasable programmable read-only memory (EEPROM), flash memory, magnetic or optical data storage media, and the like. Additionally or alternatively, the techniques may be at least partially implemented by a computer-readable communication medium that carries or conveys program code in the form of instructions or data structures that can be accessed, read, and / or executed by a computer, such as a propagated signal or wave.
[0132] The program code can be executed by a processor, which can include one or more processors, such as one or more digital signal processors (DSPs), general-purpose microprocessors, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other equivalent integrated or discrete logic circuits. Such a processor can be configured to perform any of the techniques described in this disclosure. A general-purpose processor can be a microprocessor; however, in an alternative, the processor can be any conventional processor, controller, microcontroller, or state machine. The processor can also be implemented as a combination of computing devices, such as a combination of a DSP and a microprocessor, multiple microprocessors, one or more microprocessors in conjunction with a DSP core, or any other such configuration. Thus, as used herein, the term "processor" can refer to any of the foregoing structures, any combination of the foregoing structures, or any other structure or device suitable for implementing the techniques described herein. Additionally, in some aspects, the functionality described herein can be provided within a dedicated software module or hardware module configured for encoding and decoding, or incorporated into a combined video encoder-decoder (CODEC).
[0133] Exemplary aspects of the present disclosure include:
[0134] Aspect 1. A method for image processing, the method comprising: obtaining pose information of an image sensor, the pose information indicating a set of viewpoints of an environment, wherein the image sensor is configured to capture frames for reconstructing an object; fitting a curve to the set of viewpoints; determining one or more features of the curve; determining a direction of the image sensor based on the one or more features of the curve and the set of viewpoints; determining an overlap rate based on the one or more features of the curve and the determined direction of the image sensor; and outputting the determined overlap rate to select frames to be used for reconstructing the object.
[0135] Aspect 2. The method according to aspect 1, wherein determining one or more features of the curve includes determining a center of the curve and a radius of the curve, wherein the direction of the image sensor is determined based on the determined center of the curve, and wherein the overlap rate is determined based on the radius of the curve.
[0136] Aspect 3. The method according to aspect 2, further comprising: obtaining additional pose information of the image sensor; determining that a distance from the center of the curve to a viewpoint indicated by the additional pose information has deviated from the radius of the curve by more than a threshold amount; and refitting another curve based on the viewpoint.
[0137] Aspect 4. The method according to any one of aspects 2 or 3, wherein the viewpoints in the set of viewpoints include position information and direction information of the image sensor.
[0138] Aspect 5. The method according to aspect 4, wherein fitting the curve comprises: determining a path for the set of viewpoints based on the position information; and applying a curve fitting algorithm to the path.
[0139] Aspect 6. The method according to any one of aspects 4 or 5, wherein determining the orientation of the image sensor comprises: determining a vector between the position of the viewpoints around the curve and the center of the curve; and comparing the vector with the orientation associated with the viewpoints.
[0140] Aspect 7. The method according to any one of aspects 2 to 6, wherein determining the overlap rate comprises: determining that the radius of the curve is within a radius range; and determining that the orientation of the image sensor points in an expected direction.
[0141] Aspect 8. The method according to any one of aspects 1 to 7, wherein the pose information comprises six degrees of freedom information.
[0142] Aspect 9. The method according to any one of aspects 1 to 8, wherein the overlap rate is determined based on a set of size categories.
[0143] Aspect 10. The method according to aspect 9, wherein the overlap rate is predetermined for the size categories in the set of size categories.
[0144] Aspect 11. The method according to any one of aspects 1 to 10, wherein the number of viewpoints in the set of viewpoints is predetermined.
[0145] Aspect 12. An apparatus for processing sensor data, the apparatus comprising: at least one memory; and at least one processor, the at least one processor coupled to the at least one memory and configured to: obtain pose information of an image sensor, the pose information indicating a set of viewpoints of an environment, wherein the image sensor is configured to capture frames for reconstructing an object; fit a curve to the set of viewpoints; determine one or more features of the curve; determine the orientation of the image sensor based on the one or more features of the curve and the set of viewpoints; determine an overlap rate based on the one or more features of the curve and the determined orientation of the image sensor; and output the determined overlap rate to select frames to be used for reconstructing the object.
[0146] Aspect 13. The apparatus according to aspect 12, wherein, in order to determine the one or more features of the curve, the at least one processor is configured to determine the center of the curve and the radius of the curve, wherein the orientation of the image sensor is determined based on the determined center of the curve, and wherein the overlap rate is determined based on the radius of the curve.
[0147] Aspect 14. The apparatus according to aspect 13, wherein the at least one processor is further configured to: obtain additional pose information of the image sensor; determine that a distance from the center of the curve to a viewing point indicated by the additional pose information has deviated from the radius of the curve by more than a threshold amount; and refit another curve based on the viewing point.
[0148] Aspect 15. The apparatus according to any one of aspects 13 or 14, wherein the viewing points in the set of viewing points include position information and orientation information of the image sensor.
[0149] Aspect 16. The apparatus according to aspect 15, wherein to fit the curve, the at least one processor is configured to: determine a path for the set of viewing points based on the position information; and apply a curve fitting algorithm to the path.
[0150] Aspect 17. The apparatus according to any one of aspects 15 or 16, wherein to determine the orientation of the image sensor, the at least one processor is configured to: determine a vector between the position of the viewing point around the curve and the center of the curve; and compare the vector with the orientation associated with the viewing point.
[0151] Aspect 18. The apparatus according to any one of aspects 13 to 17, wherein to determine the overlap rate, the at least one processor is configured to: determine that the radius of the curve is within a radius range; and determine that the orientation of the image sensor points in an expected direction.
[0152] Aspect 19. The apparatus according to any one of aspects 12 to 18, wherein the pose information includes six degrees of freedom information.
[0153] Aspect 20. The apparatus according to any one of aspects 12 to 19, wherein the overlap rate is determined based on a set of size categories.
[0154] Aspect 21. The apparatus according to aspect 20, wherein the overlap rate is predetermined for the size categories in the set of size categories.
[0155] Aspect 22. The apparatus according to any one of aspects 12 to 21, wherein the number of viewing points in the set of viewing points is predetermined.
[0156] Aspect 23. A non-transitory computer-readable medium storing instructions that, when executed by at least one processor, cause the at least one processor to: obtain pose information of an image sensor, the pose information indicating a set of viewpoints of an environment, wherein the image sensor is configured to capture frames for reconstructing an object; fit a curve to the set of viewpoints; determine one or more features of the curve; determine a direction of the image sensor based on the one or more features of the curve and the set of viewpoints; determine an overlap rate based on the one or more features of the curve and the determined direction of the image sensor; and output the determined overlap rate to select frames to be used for reconstructing the object.
[0157] Aspect 24. The non-transitory computer-readable medium according to aspect 23, wherein, in order to determine the one or more features of the curve, the instructions cause the at least one processor to determine a center of the curve and a radius of the curve, wherein the direction of the image sensor is determined based on the determined center of the curve, and wherein the overlap rate is determined based on the radius of the curve.
[0158] Aspect 25. The non-transitory computer-readable medium according to aspect 24, wherein the at least one processor is further configured to: obtain additional pose information of the image sensor; determine that a distance from the center of the curve to a viewpoint indicated by the additional pose information has deviated from the radius of the curve by more than a threshold amount; and refit another curve based on the viewpoint.
[0159] Aspect 26. The non-transitory computer-readable medium according to any one of aspects 24 or 25, wherein the viewpoints in the set of viewpoints include position information and direction information of the image sensor.
[0160] Aspect 27. The non-transitory computer-readable medium according to aspect 26, wherein, in order to fit the curve, the instructions cause the at least one processor to: determine a path for the set of viewpoints based on the position information; and apply a curve fitting algorithm to the path.
[0161] Aspect 28. The non-transitory computer-readable medium according to any one of aspects 26 or 27, wherein, in order to determine the direction of the image sensor, the instructions cause the at least one processor to: determine a vector between a position of a viewpoint around the curve and the center of the curve; and compare the vector with the direction associated with the viewpoint.
[0162] Aspect 29. The non-transitory computer-readable medium according to any one of aspects 24 to 28, wherein, in order to determine the overlap rate, the instructions cause the at least one processor to: determine that the radius of the curve is within a radius range; and determine that the orientation of the image sensor points to an expected direction.
[0163] Aspect 30. The non-transitory computer-readable medium according to any one of aspects 23 to 29, wherein the pose information includes six degrees of freedom information.
[0164] Aspect 31. The non-transitory computer-readable medium according to any one of aspects 23 to 30, wherein the overlap rate is determined based on a set of size categories.
[0165] Aspect 32. The non-transitory computer-readable medium according to aspect 31, wherein the overlap rate is predetermined for the size categories in the set of size categories.
[0166] Aspect 33. The non-transitory computer-readable medium according to any one of aspects 23 to 32, wherein the number of viewpoints in the set of viewpoints is predetermined.
[0167] Aspect 34: An apparatus for image generation, the apparatus including components for performing one or more of the operations according to any one of aspects 1 to 11.
Claims
1. A method for image processing, the method comprising: Obtaining pose information of an image sensor, the pose information indicating a set of viewpoints of an environment, wherein the image sensor is configured to capture frames for reconstructing an object; Fitting a curve to the set of viewpoints; Determining one or more features of the curve; Determining the direction of the image sensor based on the one or more features of the curve and the set of viewpoints; Determining an overlap rate based on the one or more features of the curve and the determined direction of the image sensor; And Outputting the determined overlap rate to select frames to be used for reconstructing the object.
2. The method according to claim 1, wherein determining one or more features of the curve comprises determining the center of the curve and the radius of the curve, wherein the direction of the image sensor is determined based on the determined center of the curve, and wherein the overlap rate is determined based on the radius of the curve.
3. The method according to claim 2, the method further comprising: Obtaining additional pose information of the image sensor; Determining that the distance from the center of the curve to a viewpoint indicated by the additional pose information has deviated from the radius of the curve by more than a threshold amount; and Refitting another curve based on the viewpoint.
4. The method according to claim 2, wherein the viewpoints in the set of viewpoints include position information and direction information of the image sensor.
5. The method according to claim 4, wherein fitting the curve comprises: Determining a path for the set of viewpoints based on the position information; And Applying a curve fitting algorithm to the path.
6. The method according to claim 4, wherein determining the direction of the image sensor comprises: Determining a vector between the position of a viewpoint around the curve and the center of the curve; And Comparing the vector with the direction associated with the viewpoint.
7. The method according to claim 2, wherein determining the overlap rate comprises: Determining that the radius of the curve is within a radius range; And Determining that the direction of the image sensor points to an expected direction.
8. The method according to claim 1, wherein the pose information includes 6 - degree - of - freedom information.
9. The method according to claim 1, wherein the overlap rate is determined based on a set of size categories.
10. The method according to claim 9, wherein the overlap rate is pre - determined for size categories in the set of size categories.
11. The method according to claim 1, wherein the number of viewpoints in the set of viewpoints is pre - determined.
12. An apparatus for processing sensor data, the apparatus comprising: At least one memory; And At least one processor, the at least one processor being coupled to the at least one memory and configured to: Obtain pose information of an image sensor, the pose information indicating a set of viewpoints of an environment, wherein the image sensor is configured to capture frames for reconstructing an object; Fit a curve to the set of viewpoints; Determine one or more features of the curve; Determine the orientation of the image sensor based on the one or more features of the curve and the set of viewpoints; Determine the overlap rate based on the one or more features of the curve and the determined orientation of the image sensor; And Output the determined overlap rate to select the frames to be used for reconstructing the object.
13. The apparatus according to claim 12, wherein in order to determine the one or more features of the curve, the at least one processor is configured to determine the center of the curve and the radius of the curve, wherein the orientation of the image sensor is determined based on the determined center of the curve, and wherein the overlap rate is determined based on the radius of the curve.
14. The apparatus according to claim 13, wherein the at least one processor is further configured to: Obtain additional pose information of the image sensor; Determine that the distance from the center of the curve to the viewpoint indicated by the additional pose information has deviated from the radius of the curve by more than a threshold amount; and Re-fit another curve based on the viewpoint.
15. The apparatus according to claim 13, wherein the viewpoints in the set of viewpoints include position information and orientation information of the image sensor.
16. The apparatus according to claim 15, wherein in order to fit the curve, the at least one processor is configured to: Determine a path for the set of viewpoints based on the position information; and Apply a curve fitting algorithm to the path.
17. The apparatus according to claim 15, wherein in order to determine the orientation of the image sensor, the at least one processor is configured to: Determine a vector between the positions of the viewpoints around the curve and the center of the curve; and Compare the vector with the orientation associated with the viewpoint.
18. The apparatus according to claim 13, wherein in order to determine the overlap rate, the at least one processor is configured to: Determine that the radius of the curve is within a radius range; and Determine that the orientation of the image sensor points to an expected direction.
19. The apparatus according to claim 12, wherein the pose information includes six degrees of freedom information.
20. The apparatus according to claim 12, wherein the overlap rate is determined based on a set of size categories.
21. The apparatus according to claim 20, wherein the overlap rate is pre-determined for the size categories in the set of size categories.
22. The apparatus according to claim 12, wherein the number of viewpoints in the set of viewpoints is pre-determined.
23. A non-transitory computer-readable medium having instructions stored thereon, which when executed by at least one processor, cause the at least one processor to: Obtain pose information of an image sensor, the pose information indicating a set of viewpoints of an environment, wherein the image sensor is configured to capture frames for reconstructing an object; Fit a curve to the set of viewpoints; Determine one or more features of the curve; Determine the orientation of the image sensor based on the one or more features of the curve and the set of viewpoints; Determine an overlap rate based on the one or more features of the curve and the determined orientation of the image sensor; and Output the determined overlap rate to select frames to be used for reconstructing the object.
24. The non-transitory computer-readable medium according to claim 23, wherein, in order to determine the one or more features of the curve, the instructions cause the at least one processor to determine the center of the curve and the radius of the curve, wherein the orientation of the image sensor is determined based on the determined center of the curve, and wherein the overlap rate is determined based on the radius of the curve.
25. The non-transitory computer-readable medium according to claim 24, wherein the at least one processor is further configured to: Obtain additional pose information of the image sensor; Determine that the distance from the center of the curve to the viewpoint indicated by the additional pose information has deviated from the radius of the curve by more than a threshold amount; and Re-fit another curve based on the viewpoint.
26. The non-transitory computer-readable medium according to claim 24, wherein the viewpoints in the set of viewpoints include position information and orientation information of the image sensor.
27. The non-transitory computer-readable medium according to claim 26, wherein, in order to fit the curve, the instructions cause the at least one processor to: Determine a path for the set of viewpoints based on the position information; and Apply a curve fitting algorithm to the path.
28. The non-transitory computer-readable medium according to claim 26, wherein, in order to determine the orientation of the image sensor, the instructions cause the at least one processor to: Determine a vector between the position of the viewpoints around the curve and the center of the curve; and Compare the vector with the orientation associated with the viewpoint.
29. The non-transitory computer-readable medium according to claim 24, wherein, in order to determine the overlap rate, the instructions cause the at least one processor to: Determine that the radius of the curve is within a radius range; and Determine that the orientation of the image sensor points to an expected direction.
30. The non-transitory computer-readable medium according to claim 23, wherein the pose information includes six degrees of freedom information.