Generation of three-dimensional and four-dimensional visual representations

By using synchronized camera groups with precise alignment and low thermal expansion materials, the method addresses the limitations of existing technologies in generating detailed multi-dimensional representations, improving accuracy and realism while reducing processing time and enhancing AI training.

WO2026030287A1PCT designated stage Publication Date: 2026-02-05STEUART III LEONARD P
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/US2025/039615
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-07-30
Filing Date
2025-07-29
Publication Date
2026-02-05

AI Technical Summary

Technical Problem

Current methods for creating accurate and realistic multi-dimensional representations of real-world environments are limited by lack of detail, accuracy, and realism, requiring excessive processing power and time, and often rely on stationary image capturing devices that struggle with moving scenes, while AI models lack high-quality training data and fundamental understanding of the physical world.

Method used

Employing a combination of rigid camera mounting, low coefficient of thermal expansion materials, and precise alignment and calibration of camera groups to capture images with synchronized three-hundred-and-sixty degree horizontal fields of view, generating three- or four-dimensional representations.

Benefits of technology

Produces highly detailed, accurate, and flexible multidimensional representations with improved fidelity and realism, reducing processing time and enhancing AI model training with high-quality visual data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025039615_05022026_PF_FP_ABST
    Figure US2025039615_05022026_PF_FP_ABST
Patent Text Reader

Abstract

Systems and methods are provided for generating and using multidimensional visual representations. In some implementation, an apparatus for generating multidimensional visual representations may comprise a first camera group, a second camera group, a frame, and a base. The first camera group may include a first plurality of cameras that are configured to collectively have a three-hundred-and-sixty degree horizontal field of view and to capture images synchronously within a first time-synchronization threshold. The second camera group may include a second plurality of cameras that are configured to collectively have a three-hundred-and-sixty degree horizontal field of view and to capture images synchronously within a second time-synchronization threshold. The frame may be configured to support the first camera group and the second camera group. The base may be configured to support the frame.
Need to check novelty before this filing date? Find Prior Art

Description

GENERATION OF THREE-DIMENSIONAL AND FOUR-DIMENSIONAL VISUAL REPRESENTATIONSCross Reference to Related Application

[0001] This application claims the benefit of priority of United States Provisional Patent Application No. 63 / 677,163, filed on July 30, 2024. The foregoing application is incorporated herein by reference in its entirety.BACKGROUNDTechnical Field

[0002] The present disclosure relates generally to photogrammetry.Background Information

[0003] Many problems exist with creating accurate and realistic multi-dimensional representations of real-world environments. While some approximative representations can be created, they lack detail, accuracy, and realism. In situations where the representation of the environment is intended for user immersion into the representation, such shortcomings destroy the experience. In situations where the representation is intended for analysis by a user or machine, such shortcomings can lead to incorrect analysis, which can cause disruptions or malfunctions for other systems or processes.

[0004] Moreover, many existing techniques require days of processing to create representations, using up large amounts of processing power, energy, and time. And even when a representation is generated, it may lack sufficient quality to be useful for its intended purpose, such as dimensional accuracy or immersion.

[0005] Many current implementations also use stationary image capturing devices, which make it difficult to accurately capture representations of scenes that move away from the stationary image capturing devices.

[0006] Many artificial intelligence (Al) models also lack high-quality visual training data, and instead train on simple images or videos. This can lead to inaccurate and un-immersive outputs, including hallucinations. Performance of Al models may be noticeably enhanced with using high quality visual data as training data.

[0007] As many companies try to create Artificial General Intelligence, they have run out of good training data. “Hitting the data wall” in Al, a current problem experienced by many of these companies and others, refers to the point where further development of a model's performance plateaus due to a lack of sufficient, high-quality data to train it. This means that simply adding more data or increasing the model's size doesn't result in significant improvements, suggesting a limit to the current approaches. Some argue that Al development isslowing down due to this data limitation. This deficiency makes them data-hungry and prone to errors when encountering situations outside of their training data.

[0008] Some have also noted that current Al systems, including large language models (LLMs), lack a fundamental understanding of the physical world and common sense, something that is already available to toddlers. To achieve more human-like intelligence in Al, methods that allow machines to learn by observing and building internal “world models” would be incredibly beneficial, though creating high-quality models of this type may only be possible through comprehensive visual and / or auditory information.

[0009] Hence, solutions are needed that provide detailed, accurate, flexible, and realistic multidimensional representations of real-world environments. Embodiments disclosed below describe apparatuses, systems, methods, and computer-readable media for achieving these solutions and more.

[0010] The disclosed embodiments may address one or more of the problems set forth above.SUMMARY

[0011] Embodiments consistent with the present disclosure provide apparatuses, systems, and methods for generating multidimensional visual representations. In some embodiments, a combination of rigid camera mounting, use of low coefficient of thermal expansion (CTE) materials, and precise alignment and calibration of camara apparatuses may result in computerized multi-dimensional representations of real world environments with extremely high degrees of fidelity

[0012] Disclosed embodiments include an apparatus with camera groups for capturing images. For example, one embodiment includes two camera groups supported by a frame and a base. Disclosed embodiments also include methods for generating multi-dimensional representations of real-world environments. For example, some embodiments include generating four-dimensional representations based on images captured by multiple camera groups. As another example, disclosed embodiments also include a system with multiple camera apparatuses configured for synchronous image capturing.

[0013] In an embodiment, an apparatus for generating multidimensional visual representations may include a first camera group, a second camera group, a frame configured to support the first camera group and the second camera group, and a base configured to support the frame. The first camera group may include a first plurality of cameras that are configured to collectively have a three-hundred-and-sixty degree horizontal field of view and to capture images synchronously within a first time-synchronization threshold. The second camera group may include a second plurality of cameras that are configured to collectively have a three-hundred-and-sixty degree horizontal field of view and to capture images synchronously within a second time-synchronization threshold.

[0014] In an embodiment, a method for generating a multi-dimensional model of an environment using a plurality of camera groups may include causing cameras associated with the plurality of camera groups to capture first images at a first time; generating, based on the first images, a first three-dimensional (3-D) model; causing the cameras associated with the plurality of camera groups to capture second images at a second time; generating, based on the second images, a second 3-D model; and generating a four-dimensional (4-D) model based on the first 3-D model and the second 3-D model. The cameras of each camera group may collectively have a three-hundred-and-sixty degree horizontal stereoscopic field of view, and the cameras of each camera group may be configured to capture images within a time-synchronization threshold.

[0015] In an embodiment, a system for generating multidimensional visual representations may include a first movable camera apparatus and a second movable camera apparatus. The first movable camera apparatus may include a first camera group including a first plurality of cameras that are configured to collectively have a three-hundred-and-sixty degree horizontal field of view and to capture images synchronously within a first time-synchronization threshold, a second camera group including a second plurality of cameras that are configured to collectively have a three-hundred-and-sixty degree horizontal field of view and to capture images synchronously within a second time-synchronization threshold, a first frame connecting the first camera group and the second camera group, and a first base configured to support the first frame. The second movable camera apparatus may include a third camera group including a third plurality of cameras that are configured to collectively have a three-hundred-and-sixty degree horizontal field of view and to capture images synchronously within a third timesynchronization threshold, a fourth camera group including a fourth plurality of cameras that are configured to collectively have a three-hundred-and-sixty degree horizontal field of view and to capture images synchronously within a fourth time-synchronization threshold, a second frame connecting the third camera group and the fourth camera group, and a second base configured to support the second frame.

[0016] In an embodiment, a method for generating a multi-dimensional model of an environment using a plurality of independently movable camera apparatuses may include receiving first model data from the plurality of independently movable camera apparatuses; receiving second model data from the plurality of independently movable camera apparatuses; and generating a comprehensive model based on the first model data and the second model data. Each independently movable camera apparatus may include a first camera group including a first plurality of cameras that are configured to collectively have a three-hundred-and-sixty degreehorizontal field of view and to capture images synchronously within a first time-synchronization threshold and a second camera group including a second plurality of cameras that are configured to collectively have a three-hundred-and-sixty degree horizontal field of view and to capture images synchronously within a second time-synchronization threshold.

[0017] In an embodiment, a method for training an artificial intelligence (Al) model based on multi-dimensional model data may include generating Al model training data based on multi-dimensional model data, the multi-dimensional model data including at least one digital 3- D representation of an environment; training an Al model using the Al model training data to predict or generate visual information; and validating the trained Al modelBRIEF DESCRIPTION OF THE DRAWINGS

[0018] Fig. 1 depicts a digital camera apparatus, consistent with disclosed embodiments.

[0019] Fig. 2 depicts another apparatus, at a bird’s eye view, consistent with disclosed embodiments.

[0020] Fig. 3 depicts an attachment configuration for an apparatus, consistent with disclosed embodiments.

[0021] Fig. 4 depicts an attachment configuration for an apparatus, at a bird’s eye view, consistent with disclosed embodiments.

[0022] Fig. 5 depicts an apparatus with cameras having overlapping fields of view, consistent with disclosed embodiments.

[0023] Fig. 6 depicts a first capturing apparatus, consistent with disclosed embodiments.

[0024] Fig. 7 depicts a second capturing apparatus, consistent with disclosed embodiments.

[0025] Fig. 8 depicts a third capturing apparatus, consistent with disclosed embodiments.

[0026] Figs. 9A and 9B depict a synchronized system for generating multidimensional visual representations, consistent with disclosed embodiments.

[0027] Fig. 10 depicts a scene having an object for multi-dimensional representation, consistent with disclosed embodiments.

[0028] Fig. 11 shows an exemplary process for generating a multi-dimensional model of an environment using a plurality of camera groups, consistent with disclosed embodiments.

[0029] Fig. 12 shows an exemplary process for generating a multi-dimensional model of an environment using a plurality of independently movable camera apparatuses, consistent with disclosed embodiments.

[0030] Fig. 13 shows an exemplary process for training an artificial intelligence (Al) model based on multi-dimensional model data, consistent with disclosed embodiments.DETAILED DESCRIPTION

[0031] The present disclosure addresses systems, components, and techniques primarily for use to generate multidimensional representations of physical real-world environments. For example, the systems, components, and / or techniques described below may be used to generate three- or four-dimensional digital representations of physical real- world environments.

[0032] Some embodiments include an apparatus for capturing images. For example, some embodiments include multiple cameras included in multiple housings, which may be configured to capture images of an environment of the cameras. In some embodiments, two cameras may be included in and / or mounted to a housing.

[0033] As depicted in Fig. 1, a housing 102 may include two cameras 104, which may be placed at a distance from each other, but with a same or similar viewing angle (e.g., according to the design of the housing). For example, two cameras 104 of housing 102 may be angled in the same or similar direction relative to a plane (e.g., a face of housing 102), but may be placed a certain distance from each other (e.g., several inches apart, a foot apart, etc.), thereby resulting in overlapping but different fields of view. Housing 102 may be composed of plastic, metal, fibers, paint, glass, or any combination thereof, as discussed further herein. The two cameras 104 may be considered as a stereoscopic pair. In some embodiments, housing 102 and cameras 104 may be part of a digital camera apparatus 100, which may include any number of housings 102 and / or cameras 104. For example, digital camera apparatus 100 may include 16 housings 102 and 32 cameras 104 (e.g., with each housing including two cameras). Alternatively, digital camera apparatus 100 may include a single housing and multiple cameras (e.g., 32 cameras). In some embodiments, housing 102 may also be constructed to include one or more openings to allow light from the environment to reach components within or partially within housing 102, such as, for example, cameras 104. While not depicted, housing 102 may also include one or more audio capture devices (e.g., microphones or other transducers), as discussed further below. For example, one or more microphones may be included by housing 102, and may be arranged adjacent to cameras 104 on a same face of housing 102. The one or more audio capture devices may be configured to capture audio while cameras are capturing images.

[0034] Regardless of the geometry of the cameras 104, the cameras 104 may be rigidly mounted and / or information about their geometry may be known to and / or determined by a system for use in different image analysis techniques, such as those described further herein. For example, a digital camera apparatus 100 used inside a room or used to capture images of a small scene may be constructed with the cameras separated by only inches, while a digital camera apparatus 100 used to cover an airport or a city block may be constructed with the cameras separated by several feet. Cameras 104 may be arranged in stereoscopic pairs where each paircaptures a portion of a 360-degree field. In some embodiments, the lenses of the cameras may have a field of view of about 100° (e.g., 100° + / - 10°), along one or more axes. Alternatively, the lenses of the cameras may have a field of view of 50° (e.g., 50° + / - 10°), along one or more axes. In some embodiments, one or more of cameras 104 may be configured to capture images at movie speed, for example at 1-60 frames per second (fps) (inclusive).

[0035] In some embodiments, digital camera apparatus 100 may include at least one processor (e.g., at least one graphic processing unit, or GPU) and at least one non-transitory computer-readable medium storing instructions that are executable by the at least one processor. For example, the instructions, when executed by the at least one processor, may cause the at least one processor to perform operations related to image analysis (e.g., operations performed on images captured by cameras of digital camera apparatus 100), consistent with disclosed embodiments (e.g., at least a portion of process 1100, at least a portion of process 1200, and / or at least a portion of process 1300).

[0036] In some embodiments, the at least one non-transitory computer-readable medium may include information describing location information (e.g., placement, location relative to a reference, and / or angle) of one or more cameras 104 and / or one or more housings 102. For example, the at least one non-transitory computer-readable medium may store data indicating a location and / or angle of at least camera 104, relative to another camera 104, relative to a housing 102, relative to digital camera apparatus 100, and / or relative to a reference point (e.g., a reference point used by other cameras 104 in digital camera apparatus 100), consistent with disclosed embodiments. This location information may be usable by at least one processor to calibrate multiple cameras 104. For example, at least one processor may align fields of view and / or images of multiple cameras 104 to a singular reference frame (e.g., coordinate system) using the location information. In some embodiments, digital camera apparatus 100 may include at least one processor in an area separate from a housing 102. In some embodiments, each housing may include a separate processor, which may be configured to perform, collectively or individually, operations with respect to images captured by cameras of the same housing 102 and / or of at least one other housing 102. Alternatively or additionally, the instructions may cause the at least one processor to perform operations such as transmitting data (e.g., metadata) related to one or more captured images, transmit data (e.g., coordinates) related to a location of one or more cameras, or other functions, as discussed herein.

[0037] Fig. 2 depicts a bird’s eye view of an apparatus 200 of multiple housings 102, which may be the same as or include features of digital camera apparatus 100. As shown in Fig. 2, multiple housings 102 may be positioned in a circular arrangement, as viewed from at least one perspective.

[0038] Fig. 3 depicts an apparatus 300 and shows an attachment configuration usable for apparatus 100, apparatus 200, or any other camera apparatus described herein. As depicted in Fig. 3, multiple housings 102 may be attached (e.g., snapped, fastened, welded, screwed) to support pieces 3000. Additionally or alternatively, multiple housings 102 may be attached to each other (e.g., in a ring-like fashion). Support pieces 3000 may be composed of plastic, metal, fibers, paint, glass, or any combination thereof. Only two housings 102 are shown in Fig. 3 to demonstrate how they can be configured to attach to parts of support pieces 3000, which may be done in a repeating fashion, with each housing 102 being configured (e.g., shaped) to attach to one of multiple potential portions of support pieces 3000 (e.g., based on the respective geometries of housing 102 and a support piece 3000). As shown in Fig. 3, an apparatus 300 (or any other apparatus described herein, such as apparatus 100 or 500) may include an upper support piece 3000 and a lower support piece 3000.

[0039] Fig. 4 depicts an apparatus 400 and shows a bird’s eye view of the embodiment shown in Fig. 3, with two housings 102 attached to two support pieces 3000. As depicted in Fig 4, housing 102 may be configured to attach to (e.g., snap to) a flat edge of a support piece 3000 and / or an angled point of a support piece 3000. In some embodiments, housing 102 may be configured (e.g., shaped, having fastener area placements, etc.) to attach to an upper support piece 3000 differently than to a lower support piece 3000, which may depend on whether the housing 102 is configured to cause cameras 104 to point upward or downward when attached to support pieces 3000. For example, an upward-facing housing 102 may be configured to attach to a flat edge of an upper support piece 3000 and to attach to an angled point of a lower support piece 3000. As another example, a downward-facing housing 102 may be configured to attach to an angled point of an upper support piece 3000 and to attach to a flat edge of a lower support piece 3000. In some embodiments, housings 102 and / or support pieces 3000 may be configured to cause housings 102 to attach to support pieces 3000 such that adjacent housings 102 are angled in different directions (e.g., thereby achieving different angles for the fields of view of their respective cameras).

[0040] Fig. 5 depicts an apparatus 500, which may be the same as or include features of digital camera apparatus 100. For example, apparatus 500 may include multiple housings and cameras, some of which may be angled upward and others downward. As discussed above, a housing 102 may include two cameras 104, which may be spaced apart from each other, but which may be angled at a same or similar direction in relation to the housing 102 (e.g., normal to a face of the housing 102) and / or apparatus 500. In some embodiments, one camera 104 may have a field-of-view (FOV) 5002a, and another camera 104, which may be included in the same housing as the camera with FOV 5002a, may have an FOV 5002b. In some embodiments,FOV 5002a and FOV 5000b may at least partially overlap, with an example of this overlap depicted in Fig. 5. Additionally, cameras included in different housings may also have respective FOVs that overlap with those of nearby cameras (e.g., cameras in adjacent housings, cameras in a housing next to an adjacent housing), which may cause apparatus 500 (which may be an instance of digital camera apparatus 100) to have a 360° FOV for one or more axes.

[0041] Fig. 6. depicts a first capturing apparatus 600, which may include a digital camera apparatus 100, which may be connected to (e.g., attached to) a frame 602 (e.g., a vertical support, such as a rod, beam, or pole). Frame 602 (e.g., a vertical support) may be composed of plastic, metal, fibers, paint, glass, or any combination thereof. In some embodiments, digital camera apparatus 100 may be configured to attached to frame 602 based on support pieces 3000. For example, support pieces 3000 may include openings through which frame 602 support can pass. Digital camera apparatus 100 may be connected to frame 602 using one or more of a fastener (e.g., screws, bolts), a friction fit, a snap fit, an adhesive, or any other stable connector. In some embodiments, digital camera apparatus 100 may be detachable from and / or moveable along frame 602, such as to allow for different height placements of digital camera apparatus along frame 602. For example, frame 602 may be a rod or a pole, and a digital camera apparatus 100 may be attached to frame 602 at one of different potential places. Additionally or alternatively, frame 602 may be configured to be extended or shortened (e.g., through it being constructed of telescoping segments), which may improve transportation capabilities of first capturing apparatus 600 while still allowing a digital camera apparatus 100 to reach a desired height. In some embodiments, frame 602 may be connected to (e.g., attached to) a base 604, which may provide stability for frame 602 and / or digital camera apparatus 100, such as by reducing sway of frame 602. Base 604 may be composed of plastic, metal, fibers, paint, glass, or any combination thereof. Frame 602 may be detachable from base 604. In some embodiments, base 604 may include wheels, rollers, spheres, or any other moveable part configured to allow base 604 to move along a surface.

[0042] Fig. 7 depicts a second capturing apparatus 700, which includes the same characteristics of capturing apparatus 600, but is shown with two digital camera apparatuses 100 attached to the vertical support. Having multiple digital camera apparatuses 100 included with a single capturing apparatus may allow for improved visual information capturing, which may lead to higher fidelity of subsequently generated visual representations, as discussed further below. The digital camera apparatuses 100 may be spaced apart from each other along an axis (e.g., along a vertical axis defined by the vertical support). Digital cameras within the same digital camera apparatus 100, as well as digital cameras within different digital camera apparatuses 100, may have at least partially overlapping fields of view.

[0043] Fig. 8 depicts a third capturing apparatus 800, which includes the same characteristics of capturing apparatus 600, but has three digital camera apparatuses 100 attached to a frame. The three digital camera apparatuses 100 may be spaced apart from one another with an equal or unequal spacing (e.g., the middle digital camera apparatuses 100 may be equidistant from a top digital camera apparatus 100 and a bottom digital camera apparatus 100, or may be closer to one of the top digital camera apparatus 100 or the bottom digital camera apparatus 100). Digital cameras within the same digital camera apparatus 100, as well as digital cameras within different digital camera apparatuses 100, may have at least partially overlapping fields of view. The features described below may be applicable to any or all of a capturing apparatus 600, a capturing apparatus 700, or a capturing apparatus 800, consistent with the disclosed embodiments.

[0044] In some embodiments, an apparatus (e.g., capturing apparatus 700 and / or capturing apparatus 800) may include a first camera group including a first plurality of cameras (e.g., cameras included in a digital camera apparatus 100). The first plurality of cameras may be configured to collectively have a three -hundred-and-sixty degree horizontal field of view (e.g., as discussed with respect to Fig. 5) and / or to capture images synchronously within a first timesynchronization threshold. The apparatus may also include a second camera group including a second plurality of cameras (e.g., cameras included in another digital camera apparatus 100). The second plurality of cameras may be configured to collectively have a three-hundred- and- sixty degree horizontal field of view and / or to capture images synchronously within a second timesynchronization threshold. Capturing apparatus 700 may also include a frame (e.g., frame 602) configured to support the first camera group and the second camera group. Capturing apparatus 700 may also include a base (e.g., base 604) configured to support the frame.

[0045] As depicted in Fig. 7 (as well as Fig. 8, discussed further below), two camera groups (or more) may be positioned (e.g., connected to, attached to, fastened to) on a frame. In some embodiments, the first camera group may be positioned on the frame (e.g., a pole) at a height different from the second camera group. For example, the bottom of the first camera group (e.g., a lowest edge of the first camera group) is positioned at least 50 centimeters away from the top of the second camera group (e.g., a highest edge of the second camera group).

[0046] In some embodiments, camera groups may be connected along an axis. For example, the first camera group and the second camera group may be connected along an axis of the frame. In some embodiments, the axis may be a vertical axis.

[0047] In some embodiments, the first plurality of cameras may include thirty-two cameras and the second plurality of cameras may also thirty-two cameras. In some embodiments, the first plurality of cameras and the second plurality of cameras may include any number ofcameras (e.g., four cameras, eight cameras, 16 cameras, 20 cameras, 64 cameras). In some embodiments, the first plurality of cameras may include a different number of cameras than the second plurality of cameras.

[0048] In some embodiments, each camera of the first plurality of cameras may comprise an image sensor and a lens (e.g., one lens and at least one image sensor per camera). Additionally or alternatively, the second plurality of cameras, or any other cameras, may also comprise an image sensor and a lens. In some embodiments, at least one camera (e.g., of the first plurality, second plurality, or any plurality) may be configured to capture a certain number of megapixels (e.g., at least one megapixel, eight megapixels, etc.).

[0049] In some embodiments, at least one of the first plurality of cameras and / or the second plurality of cameras may be arranged in a circular configuration. For example, the first plurality of cameras or the second plurality of cameras may form a circular shape with respect to at least one cross-section or axis, an example of which is depicted in Figs. 1 and 2. In some embodiments, the first plurality of cameras and / or the second plurality of cameras may be arranged in a circular configuration based on a configuration of housings 102, consistent with disclosed embodiments.

[0050] In some embodiments, at least one of the first plurality of cameras or the second plurality of cameras may include at least two cameras having overlapping fields of view (e.g., the fields of view may at least partially overlap). For example, at least two cameras of the first plurality of cameras may be angled to both capture image data (e.g., pixels, intensity values, brightness values, hue values, saturation values, color values, or any light information) representing a same portion of a physical environment of capturing apparatus 700 and capture image data representing different portions of the physical environment. One example of cameras having overlapping fields of view is depicted in Fig. 5.

[0051] The time-synchronization threshold may include a duration of time during which a plurality of cameras is configured to capture at least one image (e.g., a single image). In some embodiments, a plurality of cameras may be configured to capture an image within a centisecond of each other, within a millisecond of each other, within a microsecond of each other, or within a nanosecond of each other. For example, at least one of the first time-synchronization threshold or the second time-synchronization synchronization threshold is 1 / 1, 000th of a second or less. In some embodiments, a plurality of cameras may be configured to capture an image while using a relatively long exposure time. For example, the first plurality of cameras may be configured to collectively capture images with an exposure of between one and five seconds.

[0052] In some embodiments, the first plurality of cameras may be included in a plurality of respective first housings (e.g., first housings 102) and the plurality of second cameras may beincluded in a plurality of respective second housings (e.g., second housings 102). The first housings and / or second housings may be made of materials with a low coefficient of thermal expansion (CTE) such that temperature changes in an environment of the apparatus (e.g., capturing apparatus 700 or 800) will not cause significant changes to the geometry of any housing, which could interfere with the quality of images captured by its associated cameras. For example, in some embodiments, the plurality of first housings and the plurality of second housings are at least partially made of a material with a coefficient of thermal expansion of 25 ppm / °C or less. Additionally or alternatively, the plurality of first housings and the plurality of second housings are at least partially made of carbon fiber. Additionally or alternatively, the plurality of first housings and the plurality of second housings are at least partially made of plastic carbon.

[0053] In some embodiments, the plurality of first housings (e.g., part of a digital camera apparatus 100) may be configured to physically connect with each other. For example, first housings 102 of a digital camera apparatus 100 (e.g., included in apparatus 700 and / or 800) may be shaped to connect to or interlock with each other. Additionally or alternatively, first housings 102 of a digital camera apparatus 100 may be shaped to fit adjacently to one another and may be configured to be connected (e.g., fastened, attached, snapped) to each other. Additionally or alternatively, the plurality of second housings, or any housings 102, may also be configured to physically connect with each other.

[0054] In some embodiments, the plurality of first housings may be configured to physically connect with each other to form a first circular arrangement. For example, first housings 102 of a digital camera apparatus 100 (e.g., included in apparatus 700 and / or 800) may be shaped such that when they are connected to each other, they form a circular shape with respect to at least one cross-section or axis, an example of which is depicted in Figs. 1 and 2. Additionally or alternatively, the plurality of second housings, or any housings 102, may also be configured to form a circular arrangement.

[0055] In some embodiments, the plurality of first housings may configured to, when physically connected in the first circular arrangement, cause the first plurality of cameras to alternate having upward angled and downward angled fields of view. For example, first housings 102 of a digital camera apparatus 100 (e.g., included in apparatus 700 and / or 800) may be shaped such that they can only be connected in one or more predetermined patterns (e.g., a circular arrangement), and in such a pattern, the first housings 102 connect with one another to have alternating angles (e.g., between upward and downward angling). By way of further example, when housings 102 of a digital camera apparatus 100 are connected, a first subset of the housings 102 may be angled upward at an angle of 10° to 80° with respect to a horizontalaxis, and a second subset of housings 102 (e.g., alternating with the first subset) may be angled downward at an angle of 10° to 80° with respect to a horizontal axis. Additionally or alternatively, the plurality of second housings, or any housings 102, may also be configured to, when physically connected in a circular arrangement, cause the a plurality of cameras to alternate having upward angled and downward angled fields of view.

[0056] In some embodiments, the apparatus (e.g., capturing apparatus 700 and / or capturing apparatus 800) may include a third camera group that includes a third plurality of cameras. The third plurality of cameras may be configured to collectively have a three-hundred- and-sixty degree horizontal field of view and to capture images synchronously within a third time-synchronization threshold. The third plurality of cameras may also include any feature of the first and / or second pluralities of cameras, consistent with disclosed embodiments. For example, the third plurality of cameras may be included in a set of third housings 102.

[0057] In some embodiments, the apparatus (e.g., capturing apparatus 700 and / or capturing apparatus 800) may include a first structure configured to support the first camera group (e.g., a digital camera apparatus 100) and a second structure configured to support the second camera group (e.g., another digital camera apparatus 100). The first structure and the second structure may each include a combination of pieces (e.g., rigid pieces) made of plastic, metal, fibers, paint, glass, or any combination thereof. For example, the first structure and / or the second structure may include a support piece 3000.

[0058] In some embodiments, the apparatus (e.g., capturing apparatus 700 and / or capturing apparatus 800) may include one or more wheels. For example, the apparatus may include a based that includes one or more wheels, consistent with disclosed embodiments.

[0059] In some embodiments, the apparatus (e.g., capturing apparatus 700 and / or capturing apparatus 800) may include at least one battery configured to power the one or more wheels (e.g., cause the one or more wheels to turn according to a command). In some embodiments, the one or more wheels may be configured to spin (e.g., around a horizontal axis) and / or pivot (e.g., around a vertical axis, to allow capturing apparatus 700 to adjust a direction of travel). In some embodiments, the one or more wheels may be configured to spin and / or pivot based on one or more received commands. In some embodiments, the apparatus may be configured to receive one or more commands from a remote control device (e.g., a computer, smartphone, tablet, smartwatch, or device distinct from capturing apparatus 700). In some embodiments, the one or more commands may be configured to cause the apparatus to move from a first location to a second location. For example, the one or more commands may be configured to cause the apparatus to move at least one centimeter, at least one foot, or at least one meter, within an environment. In some embodiments, the apparatus may be configured tomove while cameras associated with it capture images. For example, the apparatus may be configured to move while the first plurality of cameras and the second plurality of cameras capture images.

[0060] In some embodiments, the apparatus (e.g., capturing apparatus 700 and / or capturing apparatus 800) may be configured to move autonomously. For example, the apparatus may be configured to capture images of its environment (e.g., using one or more cameras), analyze the images to detect one or more objects in the environment (e.g., using at least one processor), and transmit commands based on the image analysis to cause the apparatus to move within the environment while avoiding the one or more objects.

[0061] In some embodiments, the apparatus (e.g., capturing apparatus 700 and / or capturing apparatus 800) may include a non-transitory computer-readable medium storing instructions and at least one processor configured to execute the instructions to perform operations. The operations may include receiving images from the first plurality of cameras and the second plurality of cameras. The images may have been captured by the first plurality of cameras and the second plurality of cameras, for example, within a time- synchronization threshold, as disclosed herein. In some embodiments, a sequence of images may be received. For example, the operations may include receiving first images from the first plurality of cameras and the second plurality of cameras that were captured at a first time and receiving second images from the first plurality of cameras and the second plurality of cameras that were captured at a second time (e.g., a time after the first time, within a millisecond after the first time, within a microsecond after the first time, within a centisecond after the first time). In some embodiments, the first images may have been captured by the first plurality of cameras and the second plurality of cameras when the first plurality of cameras and the second plurality of cameras were at a first location, and the second images may have been captured by the first plurality of cameras and the second plurality of cameras when the first plurality of cameras and the second plurality of cameras were at a second location. By way of further example, the operations may include receiving any number of groups of images, where each group may have been captured by the first plurality of cameras and the second plurality of cameras within a particular time synchronization threshold. In some embodiments, the first plurality of cameras and the second plurality of cameras may have captured the groups of images over several seconds, minutes, hours, or days. The operations may also include constructing (e.g., generating, configuring, synthesizing, and / or compressing) a three-dimensional (3-D) representation or a fourdimensional (4-D) representation of an environment (e.g., physical environment where the apparatus is present) of the apparatus (e.g., capturing apparatus 700 and / or 800) based on the received images. A 3-D representation may include a digital 3-D visualization and / or model ofthe environment, such as a 3-D image, a Gaussian splatting model, a continuous volumetric function, a virtualized version of the environment, and / or an extended reality version of the environment. In some embodiments, the 3-D representation may be interpretable, manipulable, and / or displayable by a computing device (e.g., computer, laptop, smartphone, server, or the like).

[0062] In some embodiments, the apparatus (e.g., capturing apparatus 700 and / or capturing apparatus 800) may include wiring configured to transfer data between the first camera group and the second camera group, or between any number of camera groups associated with (e.g., connected to) the apparatus. The wiring may allow cameras of different groups to share data (e.g., image data) for pre-processing, refinement, correction, enhancement, multidimensional representation generation, or any other operation discussed herein (e.g., performable by at least one processor). Additionally or alternatively, the apparatus may include wiring configured to transfer data between cameras of a same group.

[0063] In some embodiments, the apparatus (e.g., capturing apparatus 700 and / or capturing apparatus 800) may include a wireless transmission unit configured to wirelessly transmit image data based on images captured by the first plurality of cameras and the second plurality of cameras to at least one processing device. A wireless transmission unit may include an antenna, a Bluetooth transmitter, a Wi-Fi transmitter, a Zigbee transmitter, any electronic signal-producing device, or any combination thereof. The at least one processing device may include a computing device with one or more processors, which may be included in a camera apparatus (e.g., apparatus 700 and / or 800) or may be part of a separate device (e.g., server, separate system).

[0064] In some embodiments, the image data may include a 3-D representation or a 4-D representation of an environment of the apparatus, which are discussed further herein. Additionally or alternatively, the image data may include time information (e.g., timestamps, time information indicating an absolute or relative time when at least one image was captured), camera orientation information (e.g., information describing an angle of a camera when at least one image was captured), and / or camera location information (e.g., information describing a location of a camera, camera group, or camera apparatus, when at least one image was captured). For example, the image data may include metadata indicating a location of the apparatus and / or a location of at least one camera of the first plurality of cameras or the second plurality of cameras. Camera orientation information may include multiple values for a single camera, such as a roll value, pitch value, and yaw value (or any 3-axis-based combination of values).

[0065] Time information, camera orientation information, and / or camera location information may be expressed relative to a reference (e.g., a “0” time, a predetermined 0° anglealong or more axes, a predetermined point, a predetermined two-dimensional or three- dimensional grid). In some embodiments, a reference may be determined based on a single camera or single camera unit. For example, a single camera may define a reference for one or more other cameras. As another non-mutually exclusive example, a single device (e.g., processor of a camera apparatus or a separate processing device) may define a reference for one or more other cameras. As another non-mutually exclusive example, the location indicated by the metadata may be expressed relative to a reference point used by another apparatus or relative to another location of another apparatus or a camera of another apparatus.

[0066] In some embodiments, the image data may be compressed prior to transmission between cameras, between camera groups, between apparatuses, and / or to a separate device or system, for example by removing redundancies (e.g., duplicative pixel values for a same location in the environment) prior to transmission. Other operations, such as those discussed below with respect to processes 1100 and 1200 may also be applied to the image data.

[0067] In some embodiments, the apparatus (e.g., capturing apparatus 700 and / or capturing apparatus 800) may be configured to switch between operating in a first power mode and operating in a second power mode higher than the first power mode. The apparatus may use the first power mode to capture images and / or navigate autonomously, and may use the second power mode to perform image processing operations, such as those discussed herein. For example, the apparatus may be configured to, while in the first power mode, capture images using the first plurality of cameras and the second plurality of cameras. Additionally or alternatively, the apparatus may be configured to, while in the second power mode, construct multi-dimensional visual representations using the captured images. Additionally or alternatively, the apparatus may be configured to, while in the second power mode, reduce image data redundancies, enhance the image data, and / or format the image data for transmission to a device (e.g., another apparatus, a remote computing device). In some embodiments, the first power mode and / or second power mode may draw power from at least one battery of the apparatus, consistent with disclosed embodiments. In some embodiments, the apparatus may use (e.g., may only use) the second power mode when it is connected to a non- mobile power source (e.g., a power grid, being plugged into a wall socket).

[0068] In some embodiments, the apparatus (e.g., capturing apparatus 700 and / or capturing apparatus 800) may include at least one audio capture device. An audio capture device may include a microphone or any transducer. In some embodiments, the apparatus may cause the at least one audio capture device to record, as digital audio information, sound of the environment in which the apparatus is capturing images. The digital audio information may be associated with metadata (e.g., time information), which may allow it to be correlated (e.g.,synchronized, such as by at least one processor, a capturing apparatus, etc.) with image data captured at or near the same time (e.g., according to a same timestamp or timestamp within predetermined range). In some embodiments, such as where the apparatus includes multiple audio capture devices and / or where multiple audio capture devices are included across multiple apparatuses in a same environment, at least one apparatus may be configured to create a spatial audio representation of the environment for a certain period of time (e.g., when the apparatus also captured images, when the apparatus and other apparatuses captured images and / or audio). Digital audio information and / or a spatial audio representation may be included as part of, or may be associated with (e.g., using metadata), a multi-dimensional model.L0069J For example, a capturing apparatus (e.g., capturing apparatus 700 and / or capturing apparatus 800) may include one or more microphones, which may capture audio at a predetermined sampling rate, such as 44Khz. This information may be synchronized with visual information captured by one or more cameras of the capturing apparatus. For example, audio and visual information may be synchronized to within a sampling rate of an associated microphone and / or camera. In some embodiments, audio and visual information captured by a capturing apparatus may be synchronized to within 1 / 44, 000th of a second (e.g., about .0227 milliseconds). As mentioned above, one or more microphones may be included by housing 102, and may be arranged adjacent to cameras 104 on a same face of housing 102. In some embodiments, audio captured by microphones of a single capturing apparatus or housing may be used to generate (e.g., by at least one processing device) stereo audio information. Stereo audio information from multiple sources (e.g., multiple housings, multiple apparatuses) may be combined (e.g., by at least one processing device) to generate a spatial audio representation of an environment, which may be associated with images of the same environment.

[0070] In some embodiments, multiple apparatuses, such as multiple capturing apparatuses 800, may be positioned or otherwise located within a same environment (e.g., room, building, outdoor area, venue, event). For example, Fig. 9A depicts a synchronized system 900 for generating multidimensional visual representations, which includes multiple (four, in the example) capturing apparatuses 800, which each have at least some field of view towards a scene 902. While four capturing apparatuses 800 are shown in the example of Fig. 9, it is appreciated that more or fewer capturing apparatuses 800 could be used, such as two, three, 10, 12, 20, etc. Moreover, while the capturing apparatuses 800 are shown in particular locations, it is appreciated that they may move and capture images at different locations over time, consistent with disclosed embodiments. A scene may be at least a portion of a room, at least a portion of a building, at least a portion of an outdoor area, at least a portion of a venue, at least a portion of at least one object (e.g., a person, an animal, a piece of art, a soccer ball), and / or at least a portionof any 3-D space. In some embodiments, the multiple apparatuses may be configured to capture images of the scene 902, consistent with disclosed embodiments.

[0071] In some embodiments, the synchronized system 900 may include a first movable camera apparatus (e.g., capturing apparatus 700 and / or capturing apparatus 800). The first movable camera apparatus may include a first camera group including a first plurality of cameras that are configured to collectively have a three-hundred-and-sixty degree horizontal field of view and to capture images synchronously within a first time-synchronization threshold, consistent with disclosed embodiments (e.g., as discussed above). The first movable camera apparatus may also include a second camera group including a second plurality of cameras that are configured to collectively have a three-hundred-and-sixty degree horizontal field of view and to capture images synchronously within a second time-synchronization threshold, consistent with disclosed embodiments (e.g., as discussed above). The first movable camera apparatus may also include a first frame connecting the first camera group and the second camera group, consistent with disclosed embodiments (e.g., as discussed above). The first movable camera apparatus may also include a first base configured to support the first frame, consistent with disclosed embodiments (e.g., as discussed above).

[0072] In some embodiments, the synchronized system 900 may also include a second movable camera apparatus (e.g., a second capturing apparatus 700 and / or capturing apparatus 800). A second movable camera apparatus (or third, fourth, etc.) may include the same or similar features as the first movable camera apparatus. For example, the second movable camera apparatus may include a third camera group including a third plurality of cameras that are configured to collectively have a three-hundred-and-sixty degree horizontal field of view and to capture images synchronously within a third time- synchronization threshold. The second movable camera apparatus may also include a fourth camera group including a fourth plurality of cameras that are configured to collectively have a three-hundred-and-sixty degree horizontal field of view and to capture images synchronously within a fourth time-synchronization threshold. The second movable camera apparatus may also include a second frame connecting the third camera group and the fourth camera group. The second movable camera apparatus may also include a second base configured to support the second frame.

[0073] In some embodiments, the first time-synchronization threshold may equal the second time- synchronization threshold. For example, the first camera group and the second camera group may be configured to synchronously capture images within one millisecond of each other (or any other threshold discussed herein). In some embodiments, the third timesynchronization threshold may equal the fourth time-synchronization threshold. In some embodiments, a time-synchronization threshold for one camera group may be within a thresholdof a time- synchronization threshold for another camera group. For example, a single camera group may be configured to capture images within one microsecond of each other, and multiple camera groups may be configured to capture images within one millisecond of each other.

[0074] In some embodiments, the first and second movable camera apparatuses may be configured to wirelessly transmit image data to a processing device (e.g., remote computing device, another apparatus), consistent with disclosed embodiments (e.g., as discussed above). The image data may include at least one of images captured by the first and second movable camera apparatuses, location information associated with images captured by the first and second movable camera apparatuses, or multidimensional information based on images captured by the first and second movable camera apparatuses, consistent with disclosed embodiments (e.g., as discussed above). In some embodiments, the multidimensional information including at least one of a 3-D representation of an environment of the first and second movable camera apparatuses or a 4-D representation of an environment of the first and second movable camera apparatuses, consistent with disclosed embodiments (e.g., as discussed above).

[0075] In some embodiments, the location information (e.g., included in the image data) may indicate at least one of (a) locations of the first and second movable camera apparatuses when the images were captured or (b) locations of at least one of the first camera group, the second camera group, the third camera group, or the fourth camera group, when the images were captured. The location information for each independently movable camera apparatus may be expressed either: (a) relative to a reference point used by the first and second movable camera apparatuses, (b) relative to a location of the first movable camera apparatus or the second movable camera apparatus, or (c) relative to a location of a camera of the first movable camera apparatus or the second movable camera apparatus, consistent with disclosed embodiments (e.g., as discussed above). In some embodiments, the image data may include orientation information associated with the first and second movable camera apparatuses, consistent with disclosed embodiments (e.g., as discussed above).

[0076] In some embodiments, at least one processor may localize (e.g., as part of a calibration or initialization process) locations and / or orientations of cameras of the camera group to a particular reference, such as by establishing locations of cameras of the camera group relative to a reference frame, a reference point, and / or to at least one apparatus or other camera. Localizing the camera locations and / or orientations may include determining camera locations and / or orientations relative to at least one reference and / or storing data defining the camera locations and / or orientations, which may be included in metadata and / or usable by at least one processing device for image and / or model alignment. Localization information for cameras thatare part of a single camera apparatus may be referred to as camera group localization information.

[0077] In some embodiments, at least one processor may localize (e.g., as part of a calibration or initialization process) locations and / or orientations of apparatuses of cameras to a particular reference, such as by establishing locations of the apparatuses relative to a reference frame, a reference point, and / or to at least one other apparatus. In some embodiments, localizing the locations and / or orientations of apparatuses may be based on the camera group localization information for each apparatus. For example, at least one processor may apply a single reference frame (e.g., coordinate system, single reference point, and / or apparatus as a reference point) to camera group localization information from multiple camara apparatuses to establish camera system localization information, which may relate location and / or orientation information of cameras from different camera groups (e.g., different camera apparatuses) to each other. The at least one processor may generate and store the camera system localization information, which may be included in metadata and / or usable by at least one processing device for image and / or model alignment. Establishing camera group localization information prior to (e.g., and using it for) establishing camera system localization information may significantly reduce processing time necessary for determining the camera system localization information and / or for aligning image and / or model data.

[0078] Fig. 9B depicts synchronized system 900, but in a scenario where scene 902, or at least one subject of scene 902, is moving, as indicated by the dashed arrow in the center of the figure. As discussed above, capturing apparatuses 800 may move, either autonomously, semi- autonomously, or under complete human control, as depicted by the dashed arrows for each respective capturing apparatus 800. In some embodiments, capturing apparatuses 800 may move in the same direction, or a similar direction, as scene 902 (or a subject of the scene). For example, a remote control device (e.g., operated by a human) may instruct at least one capturing apparatus 800 to move a particular distance, in a particular direction, at a particular speed, and / or at a particular speed profile. In some embodiments, a computing device may automatically generate and send movement instructions to at least one capturing apparatus 800. For example, the computing device, which may include a processor included in, or separate from, the at least one capturing apparatus 800, may analyze images of an environment of the at least one capturing apparatus 800 and based on the analysis, may track motion of at least one object (e.g., an object within scene 902) in the environment of the at least one capturing apparatus 800, such as by using a tracking algorithm (e.g., a Kalman filter, DeepSORT, ByteTrack, SORT, MDNet, or any of the like). The computing device may generate and transmit movement instructions to the at least one capturing apparatus 800 to cause it to correlate with the movement of the at least onetracked object. For example, the computing device may generate and transmit movement instructions to the at least one capturing apparatus 800 to cause it to move in a similar direction as the at least one object, maintain a field of view with respect to the at least one object, maintain a distance from the at least one object, and / or maintain an angle of orientation with respect to the at least one object.

[0079] Fig. 10 depicts a scene 1000 where an object 1002 is present (e.g., a car). A higher fidelity 3-D or 4-D representation of the object 1002 may be achieved by capturing image data with at least one apparatus (e.g., using a process described herein) at multiple viewpoints 1004 with respect to the car. As depicted in Fig. 10, the viewpoints 1004 are shaped to represent FOVs that a camera may have with respect to object 1002 at one or more times.

[0080] Fig. 11 shows an exemplary process 1100 for generating a multi-dimensional model of an environment using a plurality of camera groups. In accordance with disclosed embodiments, process 1100 may be implemented with (e.g., using) at least one digital camera apparatus 100, at least one capturing apparatus 700, at least one capturing apparatus 800, and / or any type of computerized image processing environment. For example, process 1 100 may be performed by at least one processor, at least one memory component, at least one camera, and / or by any computing device. In some embodiments, at least a portion of process 1100 may be performed by at least one processing device that is separate from cameras and / or one or more capturing apparatuses that capture image data (e.g., images used in the process). All or part of process 1100 may be implemented in conjunction with all or part of other processes and / or apparatuses discussed herein (e.g., process 1200 and / or process 1300). For example, Process 1100 may use a capturing apparatus 700 or a capturing apparatus 800 to perform at least some of the steps of the process.

[0081] At step 1102, process 1 100 may cause cameras associated with a plurality of camera groups (e.g., camera groups of a capturing apparatus 700 or a capturing apparatus 800) to capture first images at a first time. The first images may be captured within a timesynchronization threshold, consistent with disclosed embodiments (e.g., as discussed above). Causing the cameras to capture images may include transmitting a command to the cameras instructing them to capture one or more images over a period of time. The command may include information regarding technical camera settings for the capturing, such as an exposure time, consistent with disclosed embodiments. In some embodiments, the first images may include visual data of a scene (e.g., the first images are of a same scene), consistent with disclosed embodiments.

[0082] At step 1104, process 1100 may generate, based on the first images, a first three- dimensional (3-D) model. A 3-D model may include a 3-D image, a virtualized environment, avirtualized object, or any computerized visualization of an environment, consistent with disclosed embodiments. In some embodiments, generating the first 3-D model may be based on first location information associated with the first images. For example, process 1100 may access location information (e.g., coordinates expressed relative to a reference, discussed above) included in metadata of the first images to determine two-dimensional (2-D) and / or 3-D points in the environment that correspond to one or more pixels in the first images. In some embodiments, generating the first 3-D model may be based on first time information associated with the first images. For example, process 1100 may access timestamp or other time information associated with the first images to determine that the first images represent at least a portion of an environment at a particular time (e.g. “time 0” + 1 millisecond).

[0083] In some embodiments, process 1100 may (e.g., prior to generating a 3-D model, such as between step 1102 and step 1104) pre-process or enhance image data (e.g., image data generated from the first images). For example, process 1100 may remove aberrations and / or distortions (e.g., barrel distortions, pincushion distortions) within one or more images, which may thereby generate one or more corresponding enhanced images. For example, process 1100 may apply one or more filters to pixels of the one or more images. In some embodiments, process 1100 may eliminate distortions from an image to generate an undistorted image. Additionally or alternatively, process 1100 may combine depth maps of pixels (e.g., image data), average noise information associated with image data (e.g., noise within the first images), and / or use a high-dynamic range processing technique to produce image data that more accurately represents an environment (e.g., by reducing effects of shadows or poor lighting conditions). Any images discussed herein may be pre-processed or enhanced using any of these techniques (e.g., to create enhanced images prior to model generation and / or Al training).

[0084] At step 1106, process 1 100 may cause the cameras associated with the plurality of camera groups to capture second images at a second time. The second time may occur after the first time. The second images may be captured within a time-synchronization threshold, consistent with disclosed embodiments (e.g., as discussed above), which may be the same as, or different from, the time-synchronization threshold according to which the first images are captured. Causing the cameras to capture the second images may include transmitting a command (e.g., as discussed above) to the cameras instructing them to capture one or more images over a period of time. In some embodiments, the second images may include visual data of the scene (e.g., the same scene as the first images), consistent with disclosed embodiments.

[0085] At step 1108, process 1 100 may generate, based on the second images, a second 3-D model. Generating a second 3-D model may include any of the same aspects discussed above (e.g., with respect to generating the first 3-D model, expect that the second images areused). In some embodiments, generating the second 3-D model may be based on first location information associated with the first images. For example, process 1 100 may access location information (e.g., coordinates expressed relative to a reference, discussed above) included in metadata of the second images to determine two-dimensional (2-D) and / or 3-D points in the environment that correspond to one or more pixels in the second images. In some embodiments, generating the second 3-D model may be based on second time information associated with the second images. For example, process 1100 may access timestamp or other time information associated with the second images to determine that the second images represent at least a portion of an environment at a particular time (e.g. “time 0” + 2 milliseconds).

[0086] At step 1110, process 1100 may generate a four-dimensional (4-D) model based on the first 3-D model and the second 3-D model. A 4-D model may include a combination of at least one 3-D model and time and / or motion information. For example, a 4-D model may include multiple 3-D models sequenced according to when images associated with the models were captured. Additionally or alternatively, a 4-D model may model movement of one or more objects in an environment (e.g., one or more objects represented in the first and second images). For example, the 4-D model may represent motion of at least one object over time (e.g., milliseconds, seconds, minutes, hours, or days).

[0087] In some embodiments, the cameras of each camera group (e.g., the camera groups capturing images as discussed in process 1100) may collectively have a three-hundred-and-sixty degree horizontal stereoscopic field of view, consistent with disclosed embodiments (e.g., as discussed above). Additionally or alternatively, the cameras of each camera group may be configured to capture images within a time-synchronization threshold, consistent with disclosed embodiments (e.g., as discussed above).

[0088] In some embodiments, the plurality of camera groups may comprise a first camera group including a first plurality of cameras and a second camera group including a second plurality of cameras. For example, the plurality of camera groups may comprise a first digital camera apparatus 100 and a second digital camera apparatus 100. In some embodiments, the first camera group may be supported by a first structure, such as one or more physical connecting pieces configured to hold the cameras (and any associated housings 102) together, such as support pieces 3000. Additionally or alternatively, the second camera group may be supported by a second structure, such as one or more physical connecting pieces configured to hold the cameras (and any associated housings 102) together, such as support pieces 3000. In some embodiments, the first structure and the second structure are supported by a frame (e.g., a frame 602), consistent with disclosed embodiments (e.g., a discussed above). In some embodiments, a base (e.g., a base 604) may be configured to support the frame, consistent withdisclosed embodiments (e.g., a discussed above). In some embodiments, the base may include one or more wheels, consistent with disclosed embodiments (e.g., a discussed above).

[0089] In some embodiments, process 1100 may be performed by a single apparatus, which may generate a 3-D model (e.g., the first 3-D model and / or the second 3-D model) based on image data received other apparatuses. For example, a capturing apparatus 800 may define image data based on images captured by its cameras as reference images, and may receive additional supporting image data from at least one other capturing apparatus 800, which it may use to adjust and / or enhance the reference images, as well as for generating a 3-D or 4-D model (e.g., an updated 3-D or 4-D model). This process may be performed by a plurality of apparatuses, such that each apparatus has its own set of reference images, supplements those images with images from other apparatuses, and generates image data (e.g., a model) using both. Additionally or alternatively, a processing device (e.g., a computer, laptop, server) distinct from a camera apparatus may synthesize data from multiple apparatuses (e.g., 3-D models from the perspective of multiple apparatuses) to generate a comprehensive model (e.g., according to process 1200).

[0090] Fig. 12 shows an exemplary process 1200 for generating a multi-dimensional model of an environment using a plurality of independently movable camera apparatuses. In accordance with disclosed embodiments, process 1200 may be implemented with at least one a digital camera apparatus 100, at least one capturing apparatus 700, at least one capturing apparatus 800, and / or any type of computerized image processing environment. For example, process 1200 may be performed by at least one processor, at least one memory component, at least one camera, and / or by any computing device. In some embodiments, at least a portion of process 1200 may be performed by at least one processing device that is separate from cameras and / or one or more capturing apparatuses that capture image data (e.g., images used in the process). All or part of process 1200 may be implemented in conjunction with all or part of other processes and / or apparatuses discussed herein (e.g., process 1100 and / or process 1300). For example, Process 1200 may use at least one capturing apparatus 700 and / or at least one capturing apparatus 800 to perform at least some of the steps of the process.

[0091] At step 1202, process 1200 may receive first model data from a plurality of independently movable camera apparatuses. In some embodiments, the first model data may include a plurality of images and / or image data, which may be associated with a scene, as discussed herein. For example, the first model data may include one or more depth maps, which may include 2-D images and depth information (e.g., a depth value for each pixel in an image) associated with the 2-D images. Additionally or alternatively, the first model data may include apoint cloud, which may be generated by fusing multiple depth maps (e.g., depth maps for multiple images).

[0092] In some embodiments, the first model data may include first 3-D representations of the environment. For example, the first model data may include one or more of images (e.g., at least one 3-D image), metadata, depth values, time information, orientation information, or location information, consistent with disclosed embodiments. In some embodiments, each first 3- D representation may represent a different perspective of the environment (e.g., according to a field of view of at least one camera sourcing data upon which the first 3-D model is based) and is based on at least a portion of the first images captured by a respective one of the independently movable camera apparatuses. In some embodiments, each first 3-D representation may be generated based on location information associated with the at least a portion of the first images, consistent with disclosed embodiments (e.g., as discussed above).

[0093] In some embodiments, the first model data from the plurality of independently movable camera apparatuses may be based on first images captured by the plurality of independently movable camera apparatuses at a first time, consistent with disclosed embodiments. For example, the first images may be captured at a first time (e.g., “time 0” + 1 millisecond), which may precede a second time (e.g., “time 0” + 2 milliseconds).

[0094] In some embodiments, each independently movable camera apparatus may include multiple camera groups. For example, each independently movable camera apparatus may include at least a first camera group including a first plurality of cameras and a second camera group including a second plurality of cameras. The first plurality of cameras may be configured to collectively have a three-hundred-and-sixty degree horizontal field of view and to capture images synchronously within a first time-synchronization threshold, consistent with disclosed embodiments (e.g., as discussed above). The second plurality of cameras may be configured to collectively have a three-hundred-and-sixty degree horizontal field of view and to capture images synchronously within a second time-synchronization threshold. In some embodiments, the first time-synchronization threshold and the second time-synchronization threshold may each be 1 / 102th of a second or less, consistent with disclosed embodiments. Additionally, at least one of the independent movable camera apparatuses (e.g., each of the independent movable camera apparatuses) may include three camera groups, consistent with disclosed embodiments (e.g., as discussed above).

[0095] At step 1204, process 1200 may receive second model data from the plurality of independently movable camera apparatuses. In some embodiments, the second model data may include second 3-D representations of the environment. For example, the second model data may include one or more of images (e.g., at least one 3-D image), metadata, depth values, timeinformation, orientation information, or location information, consistent with disclosed embodiments. In some embodiments, each second 3-D representation may represent a different perspective of the environment (e.g., according to a FOV of at least one camera sourcing data upon which the second 3-D model is based) and is based on at least a portion of the second images captured by a respective one of the independently movable camera apparatuses. In some embodiments, each second 3-D representation may be generated based on location information associated with the at least a portion of the second images, consistent with disclosed embodiments (e.g., as discussed above). In some embodiments, at least one of the first model data or the second model data may include a point cloud.

[0096] In some embodiments, the second model data from the plurality of independently movable camera apparatuses may be based on second images captured by the plurality of independently movable camera apparatuses at a second time, consistent with disclosed embodiments. For example, the second images may be captured at a second time (e.g., “time 0” + 2 milliseconds), which may follow a first time (e.g., “time 0” + 1 millisecond).

[0097] At step 1206, process 1200 may generate a comprehensive model based on the first model data and the second model data. A comprehensive model may include a computerized and / or digital model that represents (e.g., visually) a particular physical environment at one or more times. For example, a comprehensive model may include a 3-D model or a 4-D model, described above. In some embodiments, the comprehensive model may be a 4-D model representing motion of at least one object over time, consistent with disclosed embodiments (e.g., as discussed above).

[0098] Generating the comprehensive model may include generating a first-time comprehensive 3-D representation by synthesizing the first 3-D representations, generating a second-time comprehensive 3-D representation by synthesizing the second 3-D representations, and sequencing the first-time comprehensive 3-D representation with the second-time comprehensive 3-D representation. Synthesizing 3-D representations may include enhancing image data, as discussed above, and / or correlating image data (e.g., associating pixels or other image data with points or areas in a 3-D physical or virtualized space). Sequencing 3-D representations may include merging, collating, combining, or organizing the 3-D representations into a particular format, which may model movement of one or more objects in an environment.

[0099] Fig. 13 shows an exemplary process 1300 for training an artificial intelligence (Al) model. In accordance with disclosed embodiments, process 1300 may be implemented with (e.g., using) at least one digital camera apparatus 100, at least one capturing apparatus 700, at least one capturing apparatus 800, and / or any type of computerized image processingenvironment. For example, process 1300 may be performed by at least one processor, at least one memory component, at least one camera, and / or by any computing device. In some embodiments, at least a portion of process 1300 may be performed by at least one processing device that is separate from cameras and / or one or more capturing apparatuses that capture image data (e.g., images used in the process). Additionally or alternatively, at least a portion of process 1100 may be performed by at least one processing device that is separate from another processing device that generated multi-dimensional model data (e.g., from which Al model training data is derived). All or part of process 1300 may be implemented in conjunction with all or part of other processes and / or apparatuses discussed herein (e.g., process 1100 and / or process 1200). For example, process 1300 may use a capturing apparatus 700 or a capturing apparatus 800 to perform at least some of the steps of the process. In some embodiments, process 1300 may be performed by at least one processing device (e.g., computer, multiple processors, multiple GPUs, etc.) in direct or indirect connection with one or more cameras or at least one processing device controlling the one or more cameras. In some embodiments, process 1300 may be performed after, and / or may use information from, steps of other processes, such as process 1100 and process 1200.

[0100] At step 1302, process 1300 may receive image data from a plurality of cameras. Receiving image data from a plurality of cameras may include receiving images captured by different camera apparatuses or groups and / or metadata associated with the images. Step 1302 may include the same or similar operations as steps 1102, 1106, 1202, and 1204. In some embodiments, the image data may be associated with a scene (e.g., a same scene), which may include a particular area, background, individual, group of persons, animal(s), object, group of objects, time period, and / or the like.

[0101] At step 1304, process 1300 may generate multi-dimensional model data based on (e.g., using) the received image data. The multi-dimensional model data may include at least one 3-D model and / or at least one 4-D model, each of which is discussed further above. Generating the multi-dimensional model data may include, for example, performing step 1108, step 1110, and / or step 1206. As another non-exclusive example, process 1300 may receive image data from a plurality of cameras and / or a plurality of capturing apparatuses (e.g., at least one capturing apparatus 600, at least one capturing apparatus 700, and / or at least one capturing apparatus 800).

[0102] In some embodiments, the multi-dimensional model data may include one or more models based on data captured from a same scene (e.g., same object, same group of objects, and / or same background, etc., captured within a threshold time period, such as within a minute, within 10 minutes, within an hour, within a day, etc.). For example, the multi-dimensional model data may include one or more models based on data captured from a scene that includes a same person (e.g., object).

[0103] Additionally or alternatively, the multi-dimensional model data may include one or more models based on data captured from different scenes (e.g., a different object, different combinations of objects, and / or different backgrounds, etc., captured within a threshold time period, or at different times not within a threshold time period, such as days or weeks apart). For example, the multi-dimensional model data may include one or more models based on data captured from a first scene that includes a person against a first background (e.g., in a garage) and one or more models based on data captured from a second scene that includes the person against a second background (e.g., in a park). In some embodiments, the multi-dimensional model data may include one or more models based on data captured from different scenes, but for a same type of object. For example, the multi-dimensional model data may include one or more models based on data captured from scenes that include a same type of animal (e.g., human, dog, cat, elephant, etc.), a same type of inanimate object (e.g., vehicle, house, building), or a same type of area (e.g., outdoor area, beach, park, landscape, etc.).

[0104] In some embodiments, steps 1302 and 1304 may be performed by at least one processing device designated or configured for model data generation, and steps 1306, 1308, and 1310 may be performed by at least one processing device designated or configured for Al model training.

[0105] At step 1306, process 1300 may generate Al model training data based on multidimensional model data (e.g., the multi-dimensional model data generated at step 1304). The multi-dimensional model data may include, or have been generated based upon, time- synchronized image data, consistent with disclosed embodiments. Generating training data may include filtering, compressing, trimming, resizing, reformatting, converting, and / or performing any pre-processing operation to the generated multi-dimensional model data to make it interpretable for a device or system configured to train the Al model. For example, generating training data may include downsampling the generated multi-dimensional model data or data derived from it, such as pixel data. As another example, generating training data may include normalizing the generated multi-dimensional model data or data derived from it.

[0106] In some embodiments, generating the Al model training data may include removing at least a portion of received data, such as by removing frames, images, values, or other information from multi-dimensional model data, thereby resulting in a remainder of multidimensional model data. For example, for a sequence of 3-D models (e.g., associated with different points in time), generating the Al model training data may include removing one or more of the 3-D models from the sequence (e.g., and designating the removed models as part ofa validation dataset). Removing information from multi-dimensional model data may be part of a data augmentation operation, a random erasing operation, cross validation operation, holdout validation operation, or any operation or process for generating training data and / or validation data for an Al model. In some embodiments, one or more portions of the multi-dimensional model data may be removed randomly or semi-randomly. Generating the Al model training data may include using (e.g., placing, including) the non-removed multi-dimensional model data (which may be termed a “remainder” of multi-dimensional model data) (e.g., after performance of one or more additional pre-processing operations) within the Al model training data. In some embodiments, a threshold amount of removal may be used to limit the amount of data that can be removed from the multi-dimensional model data prior to its inclusion within the Al model training data.

[0107] In some embodiments, the multi-dimensional model data upon which the generation of the Al model training data is based on is accessed or received by a device or system that will perform the training (e.g., step 1308). For example, an Al model training system may receive the multi-dimensional model data from a plurality of cameras and / or one or more devices configured to control the cameras.

[0108] In some embodiments, generating Al model training data may be further based on digital audio information, which may include audio data recorded by one or more microphones or other audio recording devices, which may be associated with (e.g., physically and / or communicably connected with) one or more cameras and / or capturing apparatuses, consistent with disclosed embodiments. Digital audio information may include, or be used to generate, a spatial audio representation, discussed above. The digital audio information may be correlated (e.g., synchronized, associated with the same metadata as, etc.) with image data, consistent with disclosed embodiments. Additionally or alternatively, the digital audio information may be part of, or separate from but associated with, a multi-dimensional model (e.g., a multi-dimensional visual model).

[0109] At step 1308, process 1300 may train an Al model using the Al model training data. The Al model may include an artificial general intelligence (AGI) model, a computer vision model, a pattern recognition model, a neural network (such as a recurrent neural network or a convolutional neural network), a long short-term memory network (LSTM network), an encoder-decoder model, a deep learning model, a transformer, a generative Al model, a video foundation model (ViFM), any computerized model, or any combination thereof. A device or system that trains the Al model may be the same as, or different from, an entity that generated the training data.

[0110] In some embodiments, training an Al model may include initializing an untrained, partially trained, or fully trained Al model, which may include setting at least one of one or more model parameters (e.g., variables, coefficients), one or more model hyperparameters (e.g., model architectural parameters, batch size, learning rate, etc.), one or more weights, one or more model node settings, loss function variables, one or more seed values, or the like. In some embodiments, initializing the Al model may be based on a type or amount of training data to be used. For example, if training data includes only 3-D model data, the Al model may be initialized according to a first configuration (e.g., combination of parameters and hyperparameters). As another example, if training data includes only 4-D model data, the Al model may be initialized according to a second configuration. Additionally or alternatively, an initialization of the Al model may be based on a frame rate, resolution, lighting condition, and / or the like, associated with image data from which the model training data was generated.

[0111] Training an Al model may also include inputting the training data to the Al model, such as at a first layer of the Al model. During training, the Al model may learn one or more visual and / or behavioral attributes of at least one subject (e.g., object, scene, etc.). For example, the Al model may learn and digitally represent the behavior of a moving object (e.g., an object included in multi-dimensional model data from which training data was generated).

[0112] Training an Al model may include training the Al model to predict or generate visual information. Generating visual information (e.g., using generative Al) may include generating images, videos, 3-D environments or representations, or 4-D environments or representations.

[0113] At step 1310, process 1300 may validate the trained Al model. Validating the trained Al may include determining whether the Al model satisfies a performance metric, such as an accuracy threshold and / or speed threshold. Validating the trained Al model may include applying one or more validation datasets to the trained Al model and analyzing output that the Al model generates based on the applied one or more validation datasets.

[0114] The features and advantages of the disclosure are apparent from the detailed specification, and thus, it is intended that the appended claims cover all systems and methods falling within the true spirit and scope of the disclosure. It is to be understood that the disclosed embodiments are not necessarily limited in their application to the details of construction and the arrangement of the components and / or methods set forth in the following description and / or illustrated in the drawings and / or the examples.

[0115] As used herein, the indefinite articles “a” and “an” mean “one or more.” Similarly, the use of a plural term does not necessarily denote a plurality unless it is unambiguous in the given context. Words such as “and” or “or” mean “and / or” unlessspecifically directed otherwise. As used herein, unless specifically stated otherwise, being “based on” may include being dependent on, being interdependent with, being associated with, being defined at least in part by, being derived from, being influenced by, or being responsive to. As used herein, “related to” may include being inclusive of, being expressed by, being indicated by, or being based on. Further, since numerous modifications and variations will readily occur from studying the present disclosure, it is not desired to limit the disclosure to the exact construction and operation illustrated and described, and accordingly, all suitable modifications and equivalents may be resorted to, falling within the scope of the disclosure.

[0116] The disclosed embodiments are capable of variations, or of being practiced or carried out in various ways. For example, while some embodiments may be described with respect to a “first” or “second” plurality of cameras, images, image groups, locations, times, time-synchronization thresholds, and / or apparatuses, any number of camera pluralities, image groups, locations, times, time-synchronization thresholds, and / or apparatuses may also be implemented. Moreover, it is appreciated that while some aspects may be discussed with respect to a particular apparatus (e.g., capturing apparatus 700), such aspects may also be implemented with other apparatuses, (e.g., digital camera apparatus 100, capturing apparatus 800, or any other apparatus or component discussed herein).

[0117] For example, while some embodiments are discussed in a context involving particular support or connection pieces, these elements need not be present in each embodiment. Such variations are fully within the scope and spirit of the described embodiments. The disclosed embodiments may be implemented in a system, an apparatus, a method, and / or a computer program product. The computer program product may include a computer-readable storage medium (or media) having computer readable program instructions thereon for causing a processor to carry out aspects of the present disclosure.

[0118] The computer-readable storage medium can be a tangible and non-transitory device that can retain and store instructions for use by an instruction execution device. The computer-readable storage medium may be, for example, but is not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. A non- exhaustive list of more specific examples of the computer-readable storage medium includes the following: a portable computer diskette, a hard disk, a random access memory (RAM), a readonly memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CDROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanically encoded device such as punch-cards or raised structures in a groove having instructions recorded thereon, andany suitable combination of the foregoing. A computer-readable storage medium, as used herein, is not to be construed as being transitory signals per se, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through a waveguide or other transmission media (e.g., light pulses passing through a fiber-optic cable), or electrical signals transmitted through a wire.

[0119] Computer-readable program instructions for carrying out operations of the present disclosure may be assembler instructions, instruction-set-architecture (ISA) instructions, machine instructions, machine dependent instructions, microcode, firmware instructions, statesetting data, or either source code or object code written in any combination of one or more programming languages, including an object oriented programming language such as Smalltalk, C++ or the like, and conventional procedural programming languages. In some embodiments, electronic circuitry including, for example, programmable logic circuitry, field-programmable gate arrays (FPGA), or programmable logic arrays (PLA) may execute the computer-readable program instructions by utilizing state information of the computer-readable program instructions to personalize the electronic circuitry, in order to perform aspects of the present disclosure.

[0120] Aspects of the present disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the disclosure. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.

[0121] These computer-readable program instructions may be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks. These computer- readable program instructions may also be stored in a computer-readable storage medium that can direct a computer, a programmable data processing apparatus, and / or other devices to function in a particular manner, such that the computer-readable storage medium having instructions stored therein comprises an article of manufacture including instructions which implement aspects of the function / act specified in the flowchart and / or block diagram block or blocks.

[0122] The computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operationalsteps to be performed on the computer, other programmable apparatus or other device to produce a computer implemented process, such that the instructions which execute on the computer, other programmable apparatus, or other device implement the functions / acts specified in the flowchart and / or block diagram block or blocks.

[0123] The flowcharts and block diagrams in the Figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowcharts or block diagrams may represent a software program, segment, or portion of code, which comprises one or more executable instructions for implementing the specified logical function(s). It should also be noted that, in some alternative implementations, the functions noted in the block may occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently, or the blocks may sometimes be executed in the reverse order, depending upon the functionality involved. Moreover, some blocks may be executed iteratively, and some blocks may not be executed at all. It will also be noted that each block of the block diagrams and / or flowchart illustration, and combinations of blocks in the block diagrams and / or flowchart illustration, can be implemented by special purpose hardware-based systems that perform the specified functions or acts, or combinations of special purpose hardware and computer instructions.

[0124] The descriptions of the various embodiments of the present disclosure have been presented for purposes of illustration, but are not intended to be exhaustive or limited to the embodiments disclosed. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The terminology used herein was chosen to best explain the principles of the embodiments, the practical application or technical improvement over technologies found in the marketplace, or to enable others of ordinary skill in the art to understand the embodiments disclosed herein.

[0125] It is appreciated that certain features of the disclosure, which are, for clarity, described in the context of separate embodiments, may also be provided in combination in a single embodiment. Conversely, various features of the disclosure, which are, for brevity, described in the context of a single embodiment, may also be provided separately or in any suitable sub-combination or as suitable in any other described embodiment of the disclosure. Certain features described in the context of various embodiments are not to be considered essential features of those embodiments, unless the embodiment is inoperative without those elements.

[0126] Although the disclosure has been described in conjunction with specific embodiments thereof, it is evident that many alternatives, modifications and variations will beapparent to those skilled in the art. Accordingly, it is intended to embrace all such alternatives, modifications and variations that fall within the spirit and broad scope of the appended claims.

Claims

CLAIMSWhat is claimed is:

1. An apparatus for generating multidimensional visual representations, the apparatus comprising: a first camera group including a first plurality of cameras that are configured to collectively have a three-hundred-and-sixty degree horizontal field of view and to capture images synchronously within a first time- synchronization threshold; a second camera group including a second plurality of cameras that are configured to collectively have a three-hundred-and-sixty degree horizontal field of view and to capture images synchronously within a second time-synchronization threshold; a frame configured to support the first camera group and the second camera group; and a base configured to support the frame.

2. The apparatus of claim 1, further comprising: a first structure configured to support the first camera group; and a second structure configured to support the second camera group.

3. The apparatus of claim 1, wherein the first plurality of cameras are included in a plurality of respective first housings and the second plurality of cameras are included in a plurality of respective second housings.

4. The apparatus of claim 3, wherein the plurality of first housings and the plurality of second housings are at least partially made of a material with a coefficient of thermal expansion of 25 ppm / °C or less.

5. The apparatus of claim 3, wherein the plurality of first housings and the plurality of second housings are at least partially made of carbon fiber.

6. The apparatus of claim 3, wherein the plurality of first housings and the plurality of second housings are at least partially made of plastic carbon.

7. The apparatus of claim 3, wherein the plurality of first housings are configured to physically connect with each other.

8. The apparatus of claim 7, the plurality of first housings are configured to physically connect with each other to form a first circular arrangement.

9. The apparatus of claim 8, wherein the plurality of first housings are configured to, when physically connected in the first circular arrangement, cause the first plurality of cameras alternate to have upward angled and downward angled fields of view.

10. The apparatus of claim 1, wherein the first plurality of cameras includes thirty-two cameras and the second plurality of cameras includes thirty-two cameras.

11. The apparatus of claim 1 , wherein each camera of the first plurality of cameras comprises an image sensor and a lens.

12. The apparatus of claim 1, wherein the first time-synchronization threshold equals the second time-synchronization threshold.

13. The apparatus of claim 1, wherein the apparatus further comprises: a non- transitory computer-readable medium storing instructions; and at least one processor configured to execute the instructions to perform operations comprising: receiving images from the first plurality of cameras and the second plurality of cameras; and constructing a three-dimensional (3-D) representation or a fourdimensional (4-D) representation of an environment of the apparatus based on the received images.

14. The apparatus of claim 1, wherein the apparatus further comprises a third camera group including a third plurality of cameras that are configured to collectively have a three-hundred- and-sixty degree horizontal field of view and to capture images synchronously within a third time-synchronization threshold.

15. The apparatus of claim 1, wherein at least one of the first plurality of cameras or the second plurality of cameras is arranged in a circular configuration.

16. The apparatus of claim 1, wherein at least one of the first plurality of cameras or the second plurality of cameras includes at least two cameras having overlapping fields of view.

17. The apparatus of claim 1, wherein the base includes one or more wheels.

18. The apparatus of claim 17, further comprising at least one battery configured to power the one or more wheels.

19. The apparatus of claim 17, wherein the apparatus is configured to move autonomously.

20. The apparatus of claim 1, wherein the apparatus is configured to receive one or more commands from a remote control device.

21. The apparatus of claim 20, wherein the one or more commands are configured to cause the apparatus to move from a first location to a second location.

22. The apparatus of claim 1 , wherein the apparatus is configured to move while the first plurality of cameras and the second plurality of cameras capture images.

23. The apparatus of claim 1, wherein at least one of the first time-synchronization threshold or the second time-synchronization synchronization threshold is 1 / 1, 000th of a second or less.

24. The apparatus of claim 1 , wherein the first camera group is positioned on the frame at a height different from the second camera group.

25. The apparatus of claim 24, wherein the bottom of the first camera group is positioned at least 50 centimeters away from the top of the second camera group.

26. The apparatus of claim 1 , wherein the first camera group and the second camera group are connected along an axis of the frame.

27. The apparatus of claim 26, wherein the axis is a vertical axis.

28. The apparatus of claim 1, further comprising wiring configured to transfer data between the first camera group and the second camera group.

29. The apparatus of claim 1 , further comprising a wireless transmission unit configured to wirelessly transmit image data based on images captured by the first plurality of cameras and the second plurality of cameras to at least one processing device.

30. The apparatus of claim 29, wherein the image data comprises a 3-D representation or a 4-D representation of an environment of the apparatus.

31. The apparatus of claim 29, wherein the image data comprises metadata indicating a location of the apparatus or a location of at least one camera of the first plurality of cameras or the second plurality of cameras.

32. The apparatus of claim 31, wherein the location is expressed either: relative to a reference point used by another apparatus; or relative to another location of another apparatus or a camera of another apparatus.

33. The apparatus of claim 1, wherein the apparatus is configured to switch between operating in a first power mode and operating in a second power mode higher than the first power mode.

34. The apparatus of claim 33, wherein the apparatus is configured to: while in the first power mode, capture images using the first plurality of cameras and the second plurality of cameras; and while in the second power mode, construct multi-dimensional visual representations using the captured images.

35. The apparatus of claim 1, further comprising at least one audio capture device.

36. The apparatus of claim 1, wherein the first plurality of cameras are configured to collectively capture images with an exposure of between one and five seconds.

37. A method for generating a multi-dimensional model of an environment using a plurality of camera groups, the method comprising: causing cameras associated with the plurality of camera groups to capture first images at a first time; generating, based on the first images, a first three-dimensional (3-D) model;causing the cameras associated with the plurality of camera groups to capture second images at a second time; generating, based on the second images, a second 3-D model; and generating a four-dimensional (4-D) model based on the first 3-D model and the second 3-D model, wherein: the cameras of each camera group collectively have a three-hundred- and- sixty degree horizontal stereoscopic field of view; and the cameras of each camera group are configured to capture images within a time-synchronization threshold.

38. The method of claim 37, wherein the plurality of camera groups comprise: a first camera group including a first plurality of cameras; and a second camera group including a second plurality of cameras.

39. The method of claim 38, wherein: the first camera group is supported by a first structure; the second camera group is supported by a second structure; and the first structure and the second structure are supported by a frame.

40. The method of claim 39, wherein a base is configured to support the frame.

41. The method of claim 40, wherein the base includes one or more wheels.

42. The method of claim 37, wherein the 4-D model represents motion of at least one object over time.

43. The method of claim 37, wherein: generating the first 3-D model is further based on first location information associated with the first images; and generating the second 3-D model is further based on second location information associated with the second images.

44. A synchronized system for generating multidimensional visual representations, comprising: a first movable camera apparatus comprising:a first camera group including a first plurality of cameras that are configured to collectively have a three-hundred-and-sixty degree horizontal field of view and to capture images synchronously within a first time- synchronization threshold; a second camera group including a second plurality of cameras that are configured to collectively have a three-hundred-and-sixty degree horizontal field of view and to capture images synchronously within a second time-synchronization threshold; a first frame connecting the first camera group and the second camera group; and a first base configured to support the first frame; and a second movable camera apparatus, independently movable from the first movable camera apparatus, comprising: a third camera group including a third plurality of cameras that are configured to collectively have a three-hundred-and-sixty degree horizontal field of view and to capture images synchronously within a third time-synchronization threshold; a fourth camera group including a fourth plurality of cameras that are configured to collectively have a three-hundred-and-sixty degree horizontal field of view and to capture images synchronously within a fourth time-synchronization threshold; a second frame connecting the third camera group and the fourth camera group; and a second base configured to support the second frame.

45. The synchronized system of claim 44, wherein the first time-synchronization threshold equals the second time-synchronization threshold.

46. The synchronized system of claim 44, wherein the first and second movable camera apparatuses are configured to wirelessly transmit image data to a processing device, the image data comprising at least one of: images captured by the first and second movable camera apparatuses; location information associated with images captured by the first and second movable camera apparatuses; ormultidimensional information based on images captured by the first and second movable camera apparatuses, the multidimensional information including at least one of: a three-dimensional (3-D) representation of an environment of the first and second movable camera apparatuses; or a four-dimensional (4-D) representation of an environment of the first and second movable camera apparatuses.

47. The synchronized system of claim 46, wherein the location information indicates at least one of locations of the first and second movable camera apparatuses when the images were captured or locations of at least one of the first camera group, the second camera group, the third camera group, or the fourth camera group, when the images were captured.

48. The synchronized system of claim 47, wherein the location information for each independently movable camera apparatus is expressed either: relative to a reference point used by the first and second movable camera apparatuses; relative to a location of the first movable camera apparatus or the second movable camera apparatus; or relative to a location of a camera of the first movable camera apparatus or the second movable camera apparatus.

49. The synchronized system of claim 46, wherein the image data further comprises orientation information associated with the first and second movable camera apparatuses.

50. A method for generating a multi-dimensional model of an environment using a plurality of independently movable camera apparatuses, the method comprising: receiving first model data from the plurality of independently movable camera apparatuses; receiving second model data from the plurality of independently movable camera apparatuses; and generating a comprehensive model based on the first model data and the second model data, wherein each independently movable camera apparatus includes: a first camera group including a first plurality of cameras that are configured to collectively have a three-hundred-and-sixty degree horizontal field of view and to capture images synchronously within a first time- synchronization threshold; anda second camera group including a second plurality of cameras that are configured to collectively have a three-hundred-and-sixty degree horizontal field of view and to capture images synchronously within a second time-synchronization threshold.

51. The method of claim 50, wherein at least one of the first model data or the second model data includes a point cloud.

52. The method of claim 51, wherein the comprehensive model is a 4-D model representing motion of at least one object over time.

53. The method of claim 51, wherein the first time-synchronization threshold and the second time-synchronization threshold are each 1 / 102th of a second or less.

54. The method of claim 51, wherein: the first model data comprises first 3-D representations of the environment; each first 3-D representation represents a different perspective of the environment and is based on at least a portion of the first images captured by a respective one of the independently movable camera apparatuses; the second model data comprises second 3-D representations of the environment; each second 3-D representation represents a different perspective of the environment and is based on at least a portion of second images captured by a respective one of the independently movable camera apparatuses; and generating the comprehensive model comprises: generating a first-time comprehensive 3-D representation by synthesizing the first 3-D representations; generating a second- time comprehensive 3-D representation by synthesizing the second 3-D representations; and sequencing the first- time comprehensive 3-D representation with the second-time comprehensive 3-D representation.

55. The method of claim 54, wherein: each first 3-D representation is generated based on location information associated with the at least a portion of the first images; andeach second 3-D representation is generated based on location information associated with the at least a portion of the second images.

56. The method of claim 51, wherein: the first model data from the plurality of independently movable camera apparatuses is based on first images captured by the plurality of independently movable camera apparatuses at a first time; and the second model data from the plurality of independently movable camera apparatuses is based on second images captured by the plurality of independently movable camera apparatuses at a second time.

57. A method for training an artificial intelligence (Al) model based on multi-dimensional model data, the method comprising: generating Al model training data based on multi-dimensional model data, the multidimensional model data including at least one digital 3-D representation of an environment; training an Al model using the Al model training data to predict or generate visual information; and validating the trained Al model.

58. The method of claim 57, further comprising generating the multi-dimensional model data.

59. The method of claim 57, wherein the multi-dimensional model data is generated based on image data received from multiple camera apparatuses.

60. The method of claim 59, wherein the image data received from the multiple camera apparatuses is associated with a same scene.

61. The method of claim 57, wherein the multi-dimensional model data includes at least one 4-D model.

62. The method of claim 57, wherein generating Al model training data comprises removing a portion of the multi-dimensional model data, thereby resulting in a remainder of multidimensional model data.

63. The method of claim 62, wherein generating Al model training data comprises including the remainder of multi-dimensional model data in the Al model training data.

64. The method of claim 57, wherein generating Al model training data is further based on digital audio information.