3D surface estimation and prediction for increasing the fidelity of real-time LiDAR model generation

A gimbal-controlled LiDAR sensor with an elevator algorithm and neural networks enhances 3D modeling fidelity by adjusting orientation for comprehensive space capture.

JP7796535B2Active Publication Date: 2026-01-09INTERNATIONAL BUSINESS MACHINE CORPORATION
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2021540810
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2019-02-06
Filing Date
2020-02-04
Publication Date
2026-01-09
Estimated Expiration
2040-02-04

AI Technical Summary

Technical Problem

Existing LiDAR sensors struggle to capture a space in all dimensions with high fidelity, requiring multiple sensors or mechanical rotation to achieve comprehensive 3D modeling.

Method used

A gimbal-mounted image sensor controlled by an elevator algorithm adjusts its orientation along multiple degrees of freedom to capture images, combined with SLAM and deep neural networks to generate high-fidelity 3D models.

Benefits of technology

Enables high-fidelity 3D capture of spaces using a single sensor, improving efficiency and reducing the need for multiple sensors.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007796535000001
    Figure 0007796535000001
  • Figure 0007796535000002
    Figure 0007796535000002
  • Figure 0007796535000003
    Figure 0007796535000003
Patent Text Reader

Abstract

A system and method for capturing images of a target area includes controlling an image sensor mounted on a gimbal to capture images of the target area over time (S202), and controlling the gimbal to adjust the orientation of the image sensor using an elevator algorithm (S204).
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates generally to computers and computer applications, and more particularly to a computer-implemented method for capturing various images of a target. [Background technology]

[0002] Time-of-flight (ToF) cameras and sensors, such as LiDAR (Light Detection and Ranging, also known as 3D laser scanning), are now becoming economical and effective for real-time 3D capture. LiDAR can be used indoors, for example, to track people in a room, and outdoors, for tracking various moving or fixed elements.

[0003] By using LiDAR in addition to SLAM (Simultaneous Localization and Mapping), it is also possible to build a detailed 3D model across time and space into a single frame of reference.

[0004] To date, a single LiDAR sensor cannot capture a space in all dimensions with high fidelity. Typically, a rotating laser or array of laser sensors that tracks a 360-degree circle or a portion of a 360-degree circle along an axis of rotation is used to capture a space in all dimensions. Multi-sensor LiDARs can also capture several such circles in a row.

[0005] What is desired is a sensor or camera coupled to a mechanical gimbal for automatically moving the sensor or camera along additional degrees of freedom (DoF), thereby enabling a high fidelity 3D capture of a space to be obtained. Summary of the Invention

[0006] In one embodiment, a computer-implemented method for capturing images of a target area can be provided, including controlling an image sensor mounted on a gimbal to capture images of the target area over time, and controlling the gimbal to adjust the orientation of the image sensor using an elevator algorithm.

[0007] A system including one or more processors operable to perform one or more of the methods described herein may also be provided.

[0008] A computer-readable storage medium storing a program of machine-executable instructions for performing one or more of the methods described herein may also be provided.

[0009] Further features, as well as the structure and operation of various embodiments, are described in detail below with reference to the accompanying drawings, where like reference numbers indicate identical or functionally similar elements. [Brief explanation of the drawings]

[0010] Embodiments of the present invention will now be described, by way of example only, with reference to the accompanying drawings, in which:

[0011] [Figure 1] 1 is a gimbal according to an embodiment of the present disclosure. [Figure 2] An example graphic diagram of the elevator algorithm. [Figure 3] FIG. 1 is a graphical representation of the gimbal and target area. [Figure 4] 1 is a flowchart including some steps of the disclosed method. [Figure 5] 1 illustrates a cloud computing environment according to an embodiment of the present invention. [Figure 6] 1 illustrates abstraction model layers according to an embodiment of the present invention. [Figure 7]1 illustrates a schematic of an exemplary computer or processing system according to an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION

[0012] The present disclosure is directed to computer systems and computer-implemented methods for capturing images of a target area. One embodiment includes a computer-implemented method for image capture of a target area, including controlling a gimbal-mounted image sensor to capture images of the target area over time, and further controlling a gimbal to adjust the orientation of the gimbal-mounted image sensor using an elevator algorithm.

[0013] FIG. 1 is a diagram of one embodiment of a gimbal assembly 110, which can be attached to any suitable structure via a connection at a base 136.

[0014] Gimbal assembly 110 includes a gimbal 120 on which an image sensor 122 is mounted. As used herein, the term gimbal refers to a number of annular rings and bearings that are actuatable and assembled to provide three-axis rotation and / or stabilization of the image sensor.

[0015] In the present disclosure, the image sensor 122 may be a LiDAR device or any other suitable imaging device capable of capturing an image of a target area.

[0016] The LiDAR device can include a LiDAR light beam source, at least one transceiver for directing the LiDAR beam from the source to a target area and collecting a beam reflected from the target area, a LiDAR detector for time-resolving the reflected beam, and one or more optical devices for routing the reflected beam. The LiDAR device is configured to generate multiple "shots" of laser light (e.g., distance measurements from the LiDAR device to the target area).

[0017] In combination with a LiDAR device, a controller 130 operating the LiDAR device can use simultaneous localization and mapping (SLAM) to combine two-dimensional range and position information collected from the LiDAR device to generate a detailed three-dimensional model spanning time and space into a single frame of reference. SLAM allows the controller to generate a map of the unknown target area (and the position of the image sensor 122 relative to the target area) using a variety of suitable algorithms, such as particle filters, extended Kalman filters, and GraphSLAM techniques.

[0018] The image sensor may be any suitable imaging device capable of capturing electro-optical (EO) radiation reflected and / or emitted from a target area. The image sensor may be an optical camera capable of detecting various forms of active and / or passive electro-optical radiation, including visible wavelength radiation, infrared radiation, broadband radiation, x-ray radiation, or any other type of radiation. The imaging device may be implemented as a video camera, such as a high-resolution video camera, or other video camera type. In other embodiments, the imaging device may include a still camera, a high-speed digital camera, a computer vision camera, etc.

[0019] The gimbal 120 is connected via an interface 124 to a controller 130, which is also part of the gimbal assembly 110. Although the controller 130 is shown as being disposed between the gimbal 120 and the base 136, in other embodiments, the controller 130 may be located anywhere else capable of sending signals to the gimbal 120, the image sensor 122, or both. The interface 124 includes mechanical and electrical interfaces for connecting the gimbal 120 and the image sensor 122 to the controller 130 (although the controller 130 may be located in other suitable locations outside the gimbal assembly 110).

[0020] Controller 130 can control gimbal 120 or image sensor 122, or both. By controlling gimbal 120, image sensor 122 mounted on the gimbal is indirectly controlled by controller 130. However, instead or additionally, image sensor 122 can also be directly controlled by controller 130, for example, by directing it to turn switches on and off, capture various images, and generate and receive laser light.

[0021] To control the gimbal 120, the controller 130 can use an elevator algorithm to adjust the orientation of the image sensor 122 (by actuating the gimbal 120) along one or more degrees of freedom so that the image sensor 122 can capture an image of the target area. These one or more degrees of freedom are the roll, pitch, and yaw of the image sensor 122. Using the elevator algorithm to adjust the orientation of the image sensor 122 allows a 3D model to be generated with a single image sensor 122, which differs from typical systems that require multiple sensors to generate a 3D model.

[0022] Actuation of the gimbal 120 along a single axis can be analogous to the movement of the read arm of a rotating hard disk drive (HDD). The SCAN algorithm (also called the elevator algorithm) can schedule a buffer of pending read or write requests to the disk and optimize the movement of the arm between the inner and outer edges of the disk as it moves in one dimension.

[0023] As an example, the elevator algorithm will be described with reference to Figure 2 in the context of a disk scheduling algorithm for moving the read / write head of a HDD. Given the following queues - 95, 180, 34, 119, 11, 123, 62, 64 - with the read / write head initially on track 50 and the tail track on 199, the elevator algorithm approach will function in the manner that the elevator does, as shown in Figure 2.

[0024] As seen in Figure 2, the elevator algorithm causes the read / write head to scan down (from a starting point of 50) towards the closest end (where zero is closer than 199), and when the read / write head hits the bottom end of the scale, it reverses and begins scanning up towards the top end of the scale to service any requests that have not been serviced before.

[0025] Similarly, points of interest with the highest priority for the disclosed gimbal 120 can be buffered as read / write requests to memory in communication with the controller 130. In this application, these buffered points of interest are along three dimensions of movement. The 3D angular position is related to the one-dimensional Substitute Value Then, instead of ordering a buffer of prioritized points of interest, we can use the elevator algorithm to Substitute Value can be applied to.

[0026] 3D angular position to 1D Substitute value forSeveral methods can be used to reduce it to: For example, one method is to use an appropriate Z-order curve, which maps multidimensional data into one dimension while preserving the locality of the data points. The Z value of a point in the multidimensional space is calculated by interleaving the binary representations of its coordinate values.

[0027] An example of a suitable Z-order curve is the application of the geohash algorithm to map multidimensional data into one dimension. The geohash algorithm can encode geographic locations into relatively short strings of numbers (one-dimensional values). The geohash algorithm is a hierarchical spatial data structure that can be used to convert values ​​onto a Z-order curve that subdivides space into grid-shaped buckets.

[0028] For purposes of describing the geohash algorithm as relevant to this disclosure, 3D points captured by image sensor 122 are represented as angles in degrees of pan and tilt, which can be converted by controller 130 to latitude / longitude positions (as if the attached sensor were the center of a virtual "Earth"). This method treats any pan reference point as zero longitude and any tilt reference point as zero latitude.

[0029] The one-dimensional decimal output of the geohash for a latitude / longitude pair (derived from pan / tilt) is sufficient as input to the 3D angle elevator algorithm, and based on this input, controller 130 can control the actuation of gimbal 120 to adjust the orientation of the image sensor along three degrees of freedom and build a 3D image model of the target area over time.

[0030] A method for capturing an image of a target area is further described with reference to Figure 3, which shows the gimbal 120 as well as the target area 138. In this embodiment, the target area 138 is 3D, i.e., substantially spherical, although in other embodiments, the target area 138 can be any 3D object, including, but not limited to, a landscape and / or waterscape (including features above and below the waterline), a building, a structure, an interior space, an exterior space, etc., and portions thereof.

[0031] As can be seen in Figure 3, image sensor 122 transmits captured images 139 along the lines shown, thus obtaining an image of target area 138. For each captured image 139, controller 130 controls image sensor 122 to collect the captured image 139. Although Figure 3 shows three captured images 139 collected over time, in other embodiments, two, four, or more captured images 139 can be collected over time by image sensor 122.

[0032] As more capture images 139 are collected by the image sensor 122, the controller 130 controls the gimbal 120 to adjust the orientation of the image sensor 122 using the elevator algorithm described above, thereby adjusting the orientation of the captured images 139.

[0033] In one embodiment, controller 130 uses an elevator algorithm to orient gimbal 120 to collect captured images 139 over the entire target area 138. In another embodiment, controller 130 uses an elevator algorithm to orient gimbal 120 to collect captured images 141 of a sub-area 140 of target area 138.

[0034] This sub-region 140 can be an area of ​​the target region 138 that is stationary (such as a house in a landscape), or the sub-region 140 can be an area of ​​the object or target region 138 that is moving (such as an animal in a landscape). The controller 130 can control the gimbal 120 to continue to capture images of the object or sub-region 140 as the object or sub-region 140 moves. In one embodiment, the controller 130 can control the gimbal 120 to continue to capture images of the moving object or moving sub-region 140 even if the object or sub-region 140 stops for a period of time.

[0035] In some embodiments, the controller 130 may utilize a deep neural network (DNN) to control the image sensor 122.

[0036] Additionally, the disclosed method can train and test a deep neural network using captured images of the target region 138. In one embodiment, the DNN can be tested using 10-fold cross validation. In one embodiment, the output of the network represents the number of combined captured images of the target region 138.

[0037] Specifically, in this disclosure, DNNs can be used to increase the fidelity of a 3D image model of target region 138. To increase the fidelity of the 3D image model of target region 138 (or any subregion 140 thereof), controller 130 can use DNNs to cause image sensor 122 to capture more images of target region 138. When using DNNs, controller 130 can increase the fidelity of the 3D image model by continuing to add and / or combine the captured images with additional images captured by image sensor 122. Controller 130 can then output a 3D image model with DNN increased fidelity.

[0038] In conjunction with the use of DNNs, controller 130 can determine a confidence interval for the DNN-based increased-fidelity 3D image model and control the operation of gimbal 120 based on the determined confidence interval. Controller 130 can determine regions of the DNN-based increased-fidelity 3D image model where the confidence interval is lower (e.g., where fewer images are captured by image sensor 122 and therefore fewer synthesized pixels in the region have a lower confidence interval), as well as regions of the DNN-based increased-fidelity 3D image model where the confidence interval is higher (e.g., where several images are captured by image sensor 122 and therefore more synthesized pixels in the region have a higher confidence interval). This confidence determination is described further below.

[0039] To use the DNN, the controller 130 pre-trains it on an input dataset of many 3D models before capturing images of the target region 138 (or any subregion 140 thereof), so that the DNN is trained to predict a dense 3D point cloud or mesh from a sparse sampling. This prediction can be constructed by intentionally downsampling a high-resolution point cloud or mesh.

[0040] Once trained, the DNN is used by controller 130 to predict a dense point cloud or mesh from images captured by image sensor 122. Segments of the captured image for which the DNN predictions return high confidence, as determined by controller 130, are then considered a low priority for rescanning. Segments for which the DNN predictions return low confidence, as determined by controller 130, are then considered a high priority for rescanning. Once the low-confidence and high-confidence regions have been determined by controller 130, controller 130 can actuate gimbal 120 to adjust the orientation of image sensor 122 to capture one or more additional images of the low-confidence regions.

[0041] To capture images of the target area 138, in step S202, the controller 130 controls the image sensor 122 mounted on the gimbal 120 to capture images of the target area 138 over time, as shown in Figure 4. In step S204, the controller 130 controls the gimbal 120 to adjust the orientation of the image sensor 122 using an elevator algorithm.

[0042] Next, in step S206, the controller 130 controls the gimbal 120 by actuating it to adjust the orientation of the image sensor 122 along three degrees of freedom using an elevator algorithm, and building a 3D image model of the target area 138 from the 3D angular position of the image sensor 122 over time.

[0043] Next, in step S208, the controller 130 controls the gimbal 120 by actuating the gimbal 120 to acquire an image of a sub-region 140 within the target region 138.

[0044] Alternatively, or in addition to step S208, after step S206, in step S210, the controller 130 controls the gimbal 120 by operating the gimbal 120 to acquire images of an object moving within the target area 138, thereby capturing images of the object over time as the object moves within the target area 138.

[0045] Alternatively, or in addition to steps S208 and / or S210, after step S206, in step S212, the controller 130 uses an elevator algorithm to convert the 3D angular position of the image sensor 122 into a one-dimensional coordinate system using a Z-order curve. Substitute Value Reduce to.

[0046] Alternatively, or in addition to steps S208 and / or S210 and / or S212, after step S206, in step S214, the controller 130 uses a deep neural network (DNN) to increase the fidelity of the 3D image model of the target region 138.

[0047] After S214, in step S216, the controller 130 controls the operation of the gimbal 120 based on the confidence of the DNN-based increased-fidelity 3D image model, such that image capture of areas with lower confidence is prioritized over areas with higher confidence.

[0048] Although this disclosure includes detailed descriptions of cloud computing, it should be understood that implementation of the teachings described herein is not limited to cloud computing environments. Rather, embodiments of the present invention can be implemented with any other type of computing environment now known or later developed.

[0049] Cloud computing is a service delivery model that enables convenient, on-demand network access to a shared pool of configurable computing resources (e.g., networks, network bandwidth, servers, processing, memory, storage, applications, virtual machines, and services) that can be rapidly provisioned and released with minimal administrative effort or interaction with a service provider. This cloud model can include at least five characteristics, at least three service models, and at least four deployment models.

[0050] The features are as follows: On-Demand Self-Service: Cloud consumers can automatically and unilaterally provision computing capacity, such as server time and network storage, as needed, without the need for human interaction with the service provider. Broad network access: Functionality is available over the network and accessed through standard mechanisms that facilitate use by heterogeneous thin or thick client platforms (e.g., cell phones, laptops, and PDAs). Resource Pooling: A provider's computing resources are pooled to serve multiple consumers using a multi-tenant model, with different physical and virtual resources dynamically allocated and reallocated on demand. Consumers are location-independent in that they generally have no control or knowledge of the exact location of the resources provided, although they may be able to identify a location at a higher level of abstraction (e.g., country, state, or data center). Rapid Elasticity: Capabilities can be rapidly and elastically provisioned and quickly scaled out, and rapidly released and quickly scaled in, sometimes automatically. To the consumer, the capacity available for provisioning often appears unlimited, and can be purchased in any quantity at any time. Service Metering: Cloud systems automatically control and optimize resource usage by using metering capabilities at some level of abstraction appropriate to the type of service (e.g., storage, processing, bandwidth, and active user accounts). Resource usage can be monitored, controlled, and reported, providing transparency to both providers and consumers of the services used.

[0051] The service model is as follows: Software as a Service (SaaS): The functionality offered to the consumer is the use of the provider's applications running on a cloud infrastructure. These applications are accessible from a variety of client devices through thin-client interfaces such as web browsers (e.g., web-based email). The consumer does not manage or control the underlying cloud infrastructure, including the network, servers, operating systems, storage, or individual application capabilities, with the expected exception of limited user-specific application configuration settings. Platform as a Service (PaaS): The functionality offered to consumers is the deployment onto a cloud infrastructure of applications they create or acquire, written using programming languages ​​and tools supported by the provider. The consumer does not manage or control the underlying cloud infrastructure, including the network, servers, operating systems, or storage, but does control the deployed applications and, in some cases, the configuration of the environment that hosts the applications. Infrastructure as a Service (IaaS): The capability offered to consumers is to provision processing, storage, network, and other basic computing resources on which the consumer can deploy and run any software, which may include operating systems and applications. The consumer does not manage or control the underlying cloud infrastructure, but does have control over the operating systems, storage, deployed applications, and in some cases limited control over the selection of network components (e.g., host firewalls).

[0052] The deployment model is as follows: Private Cloud: Cloud infrastructure is operated exclusively for an organization. It can be managed by the organization or a third party and can reside on-premise or off-premise. Community Cloud: Cloud infrastructure is shared by several organizations to support a specific community with common concerns (e.g., mission, security requirements, policies, and compliance considerations). It can be managed by those organizations or a third party and can reside on-premises or off-premises. Public Cloud: Cloud infrastructure is available to the general public or large industry groups and is owned by organizations that sell cloud services. Hybrid Cloud: A cloud infrastructure is a blend of two or more clouds (private, community, or public) that remain unique entities but are tied together by standardized or proprietary technologies (e.g., cloud bursting for load balancing between clouds) that enable data and application portability.

[0053] Cloud computing environments are service-oriented and focus on statelessness, low coupling, modularity, and semantic interoperability. At the heart of cloud computing is an infrastructure that includes a network of interconnected nodes.

[0054] Referring now to FIG. 5, an exemplary cloud computing environment 50 is shown. As shown, the cloud computing environment 50 includes one or more cloud computing nodes 10 with which local computing devices used by cloud consumers can communicate, such as, for example, a personal digital assistant (PDA) or cellular phone 54A, a desktop computer 54B, a laptop computer 54C, or an automobile computer system 54N, or a combination thereof. The nodes 10 can communicate with each other. The nodes 10 can be physically or virtually grouped in one or more networks (not shown), such as the private cloud, community cloud, public cloud, or hybrid cloud described above, or a combination thereof. This enables the cloud computing environment 50 to provide infrastructure as a service, platform as a service, or software as a service, or a combination thereof, without requiring the cloud consumer to maintain resources on a local computing device. It will be understood that the types of computing devices 54A-54N shown in FIG. 5 are intended to be exemplary only, and that computing node 10 and cloud computing environment 50 are capable of communicating with any type of computerized device (e.g., using a web browser) over any type of network or network-addressable connection or both.

[0055] Referring now to Figure 6, an exemplary set of functional abstraction layers provided by cloud computing environment 50 (Figure 5) is shown. It should be understood in advance that the components, layers, and functions shown in Figure 6 are intended to be merely exemplary, and embodiments of the present invention are not limited thereto. As shown, the following layers and corresponding functions are provided:

[0056] Hardware and software layer 60 includes hardware and software components. Examples of hardware components include mainframe 61, RISC (Reduced Instruction Set Computer) architecture-based servers 62, servers 63, blade servers 64, storage devices 65, and networks and networking components 66. In some embodiments, software components include network application server software 67 and database software 68.

[0057] The virtualization layer 70 provides an abstraction layer through which the following examples of virtual entities can be provided: virtual servers 71, virtual storage 72, virtual networks including virtual private networks 73, virtual applications and operating systems 74, and virtual clients 75.

[0058] In one example, the management layer 80 can provide the following functions: Resource provisioning 81 provides dynamic procurement of computing and other resources utilized to execute tasks within the cloud computing environment. Metering and pricing 82 provides cost tracking as resources are utilized within the cloud computing environment and billing or invoicing for the consumption of these resources. In one example, these resources may include application software licenses. Security provides identity verification for cloud consumers and tasks, as well as protection for data and other resources. User portal 83 provides access to the cloud computing environment for consumers and system administrators. Service level management 84 provides allocation and management of cloud computing resources to ensure required service levels are met. Service level agreement (SLA) planning and fulfillment 85 provides pre-allocation and procurement of cloud computing resources for anticipated future needs according to SLAs.

[0059] The workload tier 90 provides examples of functionality that can utilize a cloud computing environment. Examples of workloads and functionality that can be provided from this tier include mapping and navigation 91, software development and lifecycle management 92, virtual classroom instruction delivery 93, data analytics processing 94, transaction processing 95, and personalized recipe generation 96.

[0060] FIG. 7 illustrates a schematic diagram of an exemplary computer or processing system capable of implementing a method for generating personalized recipes in accordance with one embodiment of the present disclosure. The computer system is merely one example of a suitable processing system and is not intended to suggest any limitation as to the scope of use or functionality of the method embodiments described herein. The illustrated processing system may be operational with numerous other general-purpose or special-purpose computing system environments or configurations. Examples of well-known computing systems, environments, or configurations, or combinations thereof, that may be suitable for use with the processing system illustrated in FIG. 7 include, but are not limited to, personal computer systems, server computer systems, thin clients, thick clients, handheld or laptop devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics devices, network PCs, minicomputer systems, mainframe computer systems, and distributed cloud computing environments that include any of the above systems or devices.

[0061] A computer system may be described in the general context of computer system-executable instructions, such as program modules, being executed by a computer system. Generally, program modules may include routines, programs, objects, components, logic, data structures, etc. that perform particular tasks or implement particular abstract data types. The computer system may also be practiced in distributed cloud computing environments where tasks are performed by remote processing devices that are linked through a communications network. In a distributed cloud computing environment, program modules may be located in both local and remote computer system storage media, including memory storage devices.

[0062] Components of a computer system may include, but are not limited to, one or more processors or processing units 12, a system memory 16, and a bus 14 coupling various system components, including the system memory 16, to the processor 12. The processor 12 may include software modules 11 that perform the methods described herein. The modules 11 may be programmed into integrated circuits in the processor 12, or loaded from the memory 16, a storage device 18, or a network 24, or a combination thereof.

[0063] Bus 14 may represent any one or more of several types of bus structures, including a memory bus or memory controller, a peripheral bus, an accelerated graphics port, and a processor or local bus using any of a variety of bus architectures, including, by way of example and not limitation, an Industry Standard Architecture (ISA) bus, a Micro Channel Architecture (MCA) bus, an Enhanced ISA (EISA) bus, a Video Electronics Standards Association (VESA) local bus, and a Peripheral Component Interconnect (PCI) bus.

[0064] The computer system may include a variety of computer system readable media, which may be any available media that can be accessed by the computer system and may include both volatile and nonvolatile media, removable and non-removable media.

[0065] System memory 16 may include computer system-readable media in the form of volatile memory, such as random access memory (RAM) or cache memory, or both. The computer system may further include other removable / non-removable, volatile / non-volatile computer system storage media. By way of example only, storage system 18 may be provided for reading from and writing to non-removable, non-volatile magnetic media (e.g., a "hard drive"). Although not shown, a magnetic disk drive may be provided for reading from and writing to removable, non-volatile magnetic disks (e.g., a "floppy disk"), and an optical disk drive may be provided for reading from and writing to removable, non-volatile optical disks, such as CD-ROMs, DVD-ROMs, or other optical media. In such an example, each may be connected to bus 14 by one or more data media interfaces.

[0066] The computer system may also communicate with one or more external devices 26, such as a keyboard, pointing device, display 28, etc., one or more devices that allow a user to interact with the computer system or any device (e.g., network card, modem, etc.) that allows the computer system or both to communicate with one or more other computing devices. Such communication may occur via input / output (I / O) interface 20.

[0067] Furthermore, the computer system may communicate with one or more networks 24, such as a local area network (LAN), a general-purpose wide area network (WAN), or a public network (e.g., the Internet), or a combination thereof, via a network adapter 22. As shown, the network adapter 22 communicates with other components of the computer system via bus 14. Although not shown, it should be understood that other hardware and / or software components may be used with the computer system. Examples include, but are not limited to, microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data archive storage systems.

[0068] The present invention may be a system, method, or computer program product, or combination thereof, at any possible level of technical detail of integration. The computer program product may include computer-readable storage medium(s) having computer-readable program instructions thereon for causing a processor to perform aspects of the present invention.

[0069] A computer-readable storage medium may be any tangible device capable of holding and storing instructions for use by an instruction execution device. A computer-readable storage medium may be, for example, but not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the above. A non-exhaustive list of more specific examples of computer-readable storage media includes the following: portable computer diskettes, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disk read-only memory (CD-ROM), digital versatile disks (DVDs), memory sticks, floppy disks, mechanically encoded devices such as punch cards or ridge structures in grooves having instructions recorded thereon, and any suitable combination of the above. As used herein, computer-readable storage media is not to be construed as transitory signals per se, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through a waveguide or other transmission medium (e.g., light pulses through fiber optic cable), or electrical signals sent through wires.

[0070] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to each computing / processing device or to an external computer or storage device over a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network can include copper cables, optical fibers, wireless networks, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and transfers the computer-readable program instructions for storage in a computer-readable storage medium within the respective computing / processing device.

[0071] The computer-readable program instructions for carrying out the operations of the present invention may be source or object code written in any combination of one or more programming languages, including assembler instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state setting data, configuration data for integrated circuits, or object-oriented programming languages ​​such as Smalltalk, C++, and conventional procedural programming languages ​​such as the "C" programming language or similar programming languages. The computer-readable program instructions may run entirely on the user's computer, partially on the user's computer as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the latter scenario, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet Service Provider). In some embodiments, electronic circuitry including, for example, a programmable logic circuit, a field programmable gate array (FPGA), or a programmable logic array (PLA) can execute computer readable program instructions to individualize the electronic circuitry by utilizing state information in the computer readable program instructions to implement aspects of the present invention.

[0072] Aspects of the present invention will be described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems) and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.

[0073] These computer-readable program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus to produce a machine, whereby the instructions, executed by the processor of the computer or other programmable data processing apparatus, create means for performing the functions / acts specified in one or more blocks of the flowcharts or block diagrams, or both. These computer program instructions can also be stored in a computer-readable medium that can direct a computer, other programmable data processing apparatus, or other device, or combination thereof, to function in a particular manner, whereby the instructions stored in the computer-readable medium include an article of manufacture containing instructions that implement aspects of the functions / acts specified in one or more blocks of the flowcharts or block diagrams, or both.

[0074] The computer-readable program instructions may be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, whereby the instructions executing on the computer or other programmable apparatus provide a process for performing the functions / operations specified in one or more blocks of the flowchart or block diagram, or both.

[0075] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowcharts may represent a module, segment, or portion of code, including one or more executable instructions for implementing the specified logical function(s). In some alternative implementations, the functions shown in the blocks may occur in a different order than that shown in the figures. For example, two blocks shown in succession may in fact be executed substantially simultaneously, or the blocks may sometimes be executed in the reverse order, depending on the functionality involved. It should also be noted that each block of the block diagrams and / or flowchart diagrams, and combinations of blocks in the block diagrams and / or flowchart diagrams, may be implemented by a dedicated hardware-based system that performs the specified functions or operations, or a combination of dedicated hardware and computer instructions.

[0076] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting of the disclosure. As used herein, the singular forms "a," "an," and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. Furthermore, it will be understood that the terms "comprise" and / or "comprising," when used herein, specify the presence of stated features, integers, steps, operations, elements, or components, or combinations thereof, but do not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, or groups, or combinations thereof.

[0077] Corresponding structure, materials, acts, and equivalents of "means-or-step-plus-function" elements in the following claims are intended to include any structure, material, or acts for performing the functions in conjunction with other claimed elements as explicitly claimed. The description in the present disclosure has been presented only for purposes of illustration and description and is not intended to be exhaustive or to limit the invention to the precise form disclosed. Many modifications and variations will be apparent to those skilled in the art that do not depart from the scope of the invention. The embodiments were chosen and described to best explain the principles and practical applications of the invention and to enable those skilled in the art to understand the invention in terms of various embodiments with various modifications as suited to the particular applications contemplated.

[0078] Furthermore, while preferred embodiments of the present invention have been described using specific terminology, it should be understood that such description is for illustrative purposes only, and that changes and modifications can be made without departing from the spirit or scope of the following claims.

Claims

1. 1. A method for image capture of a target area by computer information processing, comprising: controlling a gimbal-mounted image sensor to capture images of the target area over time; reducing the 3D angular position of the image sensor to a one-dimensional proxy, controlling the gimbal to adjust the orientation of the image sensor along three degrees of freedom using an elevator algorithm, and constructing a 3D image model of the target area over time from the 3D angular position of the image sensor and distance information to the target area at the 3D angular position; A method comprising:

2. The method of claim 1 , wherein controlling the gimbal includes operating the gimbal to capture an image of a subregion within the target region.

3. 10. The method of claim 1, wherein controlling the gimbal includes actuating the gimbal to acquire images of an object moving within the target area, thereby capturing images of the object over time as the object moves within the target area.

4. 1. A method for image capture of a target area by computer information processing, comprising: controlling a gimbal-mounted image sensor to capture images of the target area over time; reducing the 3D angular position of the image sensor to a one-dimensional proxy, controlling the gimbal to adjust the orientation of the image sensor along three degrees of freedom using an elevator algorithm, and constructing a 3D image model of the target area over time from the 3D angular position of the image sensor and distance information to the target area at the 3D angular position; using a deep neural network (DNN) to increase the fidelity of the 3D image model of the target area; and actuating the gimbal based on a confidence level of the DNN-generated increased-fidelity 3D image model, whereby image capture of areas of lower confidence is prioritized over areas of higher confidence.

5. using the DNN to increase the fidelity of the 3D image model of the target area includes training the DNN using captured images of the target area; and actuating the gimbal based on the confidence of the DNN-generated increased-fidelity 3D image model includes capturing one or more additional images of the lower confidence areas, with rescanning of the lower confidence areas being prioritized over the higher confidence areas. The method of claim 4.

6. 2. The method of claim 1, wherein reducing the 3D angular positions of the image sensor to a one-dimensional proxy value comprises reducing the 3D angular positions of the image sensor to a one-dimensional proxy value using a Z-order curve.

7. The method according to any one of claims 1 to 6, wherein the software is provided as a service in a cloud environment.

8. 1. A system for image capture of a target area, comprising: one or more storage devices; one or more hardware processors coupled to the one or more storage devices; one or more hardware processors operable to control gimbal-mounted image sensors to capture images of the target area over time; one or more hardware processors operable to reduce the 3D angular positions of the image sensor to a one-dimensional proxy, control the gimbal to adjust the orientation of the image sensor along three degrees of freedom using an elevator algorithm, and build a 3D image model of the target area over time from the 3D angular positions of the image sensor and distance information to the target area at the 3D angular positions; Including, the system.

9. 9. The system of claim 8, wherein the one or more hardware processors controlling the gimbals are further configured to control the gimbals by actuating the gimbals to acquire images of subregions within the target region.

10. 10. The system of claim 8, wherein the one or more hardware processors controlling the gimbals are further configured to control the gimbals by operating the gimbals to acquire images of an object moving within the target area, thereby capturing images of the object over time as the object moves within the target area.

11. 1. A system for image capture of a target area, comprising: one or more storage devices; one or more hardware processors coupled to the one or more storage devices; one or more hardware processors operable to control gimbal-mounted image sensors to capture images of the target area over time; one or more hardware processors operable to reduce the 3D angular positions of the image sensor to a one-dimensional proxy, control the gimbals to adjust the orientation of the image sensor along three degrees of freedom using an elevator algorithm, and build a 3D image model of the target area over time from the 3D angular positions of the image sensor and distance information to the target area at the 3D angular positions; Including, further configured to use a deep neural network (DNN) to increase the fidelity of the 3D image model of the target area; the one or more hardware processors controlling the gimbals are further configured to operate the gimbals based on a confidence level of the DNN-based increased-fidelity 3D image model, whereby image capture of regions with lower confidence is prioritized over regions with higher confidence. system.

12. using the DNN to increase the fidelity of the 3D image model of the target area includes training the DNN using captured images of the target area; and wherein the one or more hardware processors controlling the gimbal actuating the gimbal based on the confidence of the DNN-generated increased-fidelity 3D image model includes capturing one or more additional images of the lower confidence areas, with rescanning of the lower confidence areas being prioritized over the higher confidence areas. The system of claim 11.

13. 9. The system of claim 8, wherein the one or more hardware processors that reduce the 3D angular positions of the image sensor to one-dimensional proxy values ​​are further configured to reduce the 3D angular positions of the image sensor to one-dimensional proxy values ​​using a Z-order curve.

14. A computer program comprising computer readable program instructions for causing a computer to carry out the method of any of claims 1 to 7.

15. A computer readable storage medium having stored thereon computer readable program instructions for causing a computer to perform the method of any of claims 1 to 7.

Citation Information

Patent Citations

  • Laser radar device and imaging target selection device using the same

    JP2013083510A

  • 3D geometry denoising method and apparatus using deep learning

    KR101853237B1

  • Variable field of view and directional sensors for mobile machine vision applications

    US20180227566A1