Artificial intelligence modeling techniques for vision-based occupancy determination.

A trained AI model processes camera data to predict voxel occupancy, improving navigation safety by generating real-time occupancy maps for autonomous vehicles and robots.

JP2025533398APending Publication Date: 2025-10-07TESLA INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025513616
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-09-30
Filing Date
2023-09-07
Publication Date
2025-10-07

AI Technical Summary

Technical Problem

Existing autonomous vehicles and robots lack effective methods to analyze and predict the occupancy of objects in their surroundings, which is crucial for safe navigation in complex environments.

Method used

A system utilizing a trained artificial intelligence model processes image data from cameras to predict occupancy attributes of voxels in the ego's surroundings, generating a dataset and graphical indicators for display, and can be integrated with autonomous driving protocols.

Benefits of technology

Enables accurate prediction of object occupancy in the vehicle's environment, enhancing safety and navigation capabilities by providing real-time occupancy maps for autonomous decision-making.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025533398000001_ABST
    Figure 2025533398000001_ABST
Patent Text Reader

Abstract

Disclosed herein are methods and systems for using artificial intelligence modeling techniques to train and run an artificial intelligence model to analyze camera feeds received from an ego and generate occupancy data indicating whether various voxels within the ego's periphery are occupied by objects having mass. The method includes inputting image data of the space surrounding the ego object into an artificial intelligence model using the ego object's camera, running the artificial intelligence model to predict occupancy attributes of a plurality of voxels, and generating a dataset based on the plurality of voxels and their corresponding occupancy attributes.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] [CROSS-REFERENCE TO RELATED APPLICATIONS] This application claims priority to U.S. Provisional Patent Application No. 63 / 375,199, filed September 9, 2022, and U.S. Provisional Patent Application No. 63 / 377,954, filed September 30, 2022, each of which is incorporated by reference herein in its entirety.

[0002] The present disclosure generally relates to artificial intelligence-based modeling techniques that analyze image data and predict occupancy attributes of ego surroundings. [Background technology]

[0003] Due to dramatic advances in computer technology, autonomous navigation technology used by autonomous vehicles and robots (collectively, ego) has become ubiquitous. These advances enable safer and more reliable autonomous navigation for egos. Egos often must navigate complex, dynamic environments and terrain that may include vehicles, traffic, pedestrians, cyclists, and a variety of other static and dynamic obstacles. Understanding an ego's surroundings is essential for making informed decisions to avoid collisions. Summary of the Invention

[0004] For these reasons, there is a need for a method and system that can analyze an ego's surroundings and predict the objects having mass that are present within the ego's surroundings. Specifically, a trained artificial intelligence (AI) model used within a particular AI architecture can predict occupancy data associated with the space surrounding the ego. As used herein, occupancy data, or occupancy attributes, may refer to whether a defined space is occupied by objects having mass (e.g., occupied or unoccupied).

[0005] In one embodiment, the method includes inputting image data of a space surrounding the ego object into an artificial intelligence model by a processor using a camera of the ego object; predicting occupancy attributes of a plurality of voxels by the processor executing the artificial intelligence model; and generating, by the processor, a dataset based on the plurality of voxels and their corresponding occupancy attributes.

[0006] The method may further include generating, by the processor, an output representing the environment of the ego object and showing a plurality of voxels and their corresponding occupancy attributes, the output including a graphical indicator of the occupancy attributes for at least some of the plurality of voxels.

[0007] The graphical indicator may correspond to a detected object associated with at least a portion of the plurality of voxels.

[0008] The method may further include displaying, by the processor, the output on a screen associated with the ego object.

[0009] The dataset may be a queryable dataset configured to transmit occupancy attributes of a plurality of voxels to an autonomous driving protocol of the ego object.

[0010] An artificial intelligence model can be trained using sensor attributes of multiple voxels.

[0011] The ego object may be an autonomous vehicle that executes a driving protocol based on a dataset.

[0012] The method may further include characterizing, by the processor, the image data prior to executing the artificial intelligence model.

[0013] The image data may include multiple camera feeds from multiple cameras of the ego object, and the method may further include temporally aligning, by the processor, the multiple camera feeds.

[0014] In another embodiment, the ego object includes a camera, a first processor, a second processor, and a non-transitory computer-readable medium storing an artificial intelligence model configured to be executed by the first processor, wherein the first processor is configured to input image data of a space around the ego object into the artificial intelligence model using the ego object's camera, execute the artificial intelligence model to predict occupancy attributes of a plurality of voxels, and generate a dataset based on the plurality of voxels and their corresponding occupancy attributes, and wherein the second processor is configured to use the dataset to autonomously navigate the ego object.

[0015] The first processor may be further configured to generate an output representing the environment of the ego object and showing a plurality of voxels and their corresponding occupancy attributes, the output including a graphical indicator of the occupancy attributes for at least some of the plurality of voxels.

[0016] The graphical indicator may correspond to a detected object associated with at least a portion of the plurality of voxels.

[0017] The first processor may be further configured to display the output on a screen associated with the ego object.

[0018] An artificial intelligence model can be trained using sensor attributes of multiple voxels.

[0019] The ego object may be an autonomous vehicle that executes a driving protocol based on a dataset.

[0020] In another embodiment, a method includes training, by a processor, an artificial intelligence model using a training dataset including data received from a camera of the ego object, the training dataset having a set of data points, each data point in the set of data points corresponding to a location and image attribute of at least one voxel in a space surrounding the ego object, whereby the artificial intelligence model correlates each data point in the first set of data points with a corresponding data point in the second set of data points using each data point's respective location, whereby once the artificial intelligence model is trained, the artificial intelligence model is configured to receive a camera feed from a second ego object and predict a third set of data points, each data point in the third set of data points corresponding to an occupancy attribute indicating whether at least one voxel in the space surrounding the second ego object is occupied by any object having mass.

[0021] The artificial intelligence model may further be configured to generate an output that represents the ego object's environment and indicates at least one voxel and its corresponding occupancy attribute.

[0022] The training data set may further include a second set of data points, each data point in the second set of data points corresponding to a location of at least one voxel in space surrounding the ego object and a sensor attribute.

[0023] The graphical indicator may correspond to a detected object associated with at least a portion of the at least one voxel.

[0024] The artificial intelligence model uses a 3D multi-view reconstruction protocol to generate the output. [Brief explanation of the drawings]

[0025] Non-limiting embodiments of the present disclosure are described, by way of example, with reference to the accompanying figures, which are schematic and are not intended to be drawn to scale. Unless otherwise indicated as representing background art, the figures represent aspects of the present disclosure.

[0026] [Figure 1A] FIG. 1 illustrates components of an AI-enabled visual data analysis system, according to one embodiment.

[0027] [Figure 1B] FIG. 1 illustrates various sensors associated with an ego, according to one embodiment.

[0028] [Figure 1C] FIG. 1 illustrates components of a vehicle, according to one embodiment.

[0029] [Figure 2] FIG. 1 is a flow diagram of a process performed in an AI-enabled visual data analysis system, according to one embodiment.

[0030] [Figure 3A] 1A-1C illustrate various occupancy maps generated by an AI-enabled visual data analysis system, according to one embodiment. [Figure 3B] 1A-1C illustrate various occupancy maps generated by an AI-enabled visual data analysis system, according to one embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0031] Reference will now be made to the exemplary embodiments illustrated in the drawings and specific language will be used to describe them herein. It will be understood, however, that no limitation of the claims or the scope of the present disclosure is intended in this regard. Alterations and further modifications of the features of the invention illustrated herein, and further applications of the principles of the subject matter illustrated herein, which will occur to those skilled in the relevant arts and in possession of this disclosure, are deemed to be within the scope of the subject matter disclosed herein. Other embodiments may be utilized and / or other changes may be made without departing from the spirit or scope of the disclosure. The exemplary embodiments described in the detailed description do not limit the presented subject matter.

[0032] By implementing the methods described herein, a system can use a trained AI model to determine the occupancy state of various voxels in an image (or video) of an ego's surroundings. The ego can be an autonomous vehicle (e.g., a car, truck, bus, motorcycle, all-terrain vehicle, or cart), a robot, or other automated device. The ego can be configured to operate on a production line, within a building, a residence, or a medical center, or to transport people, carry cargo, or perform military functions. Within these environments, the ego can navigate known or unknown paths to perform specific tasks or travel to specific destinations. Because of a desire to avoid collisions during operation, the ego attempts to understand its environment. For example, in the context of an autonomous vehicle or robot, the system can use a camera (or other visual sensor) to receive real-time or near-real-time images of the ego's surroundings. The system can then execute the trained AI model to determine the occupancy state of the ego's surroundings. After the AI ​​model divides the ego's surroundings into various voxels, it can then determine the occupancy state for each voxel. Thus, using the methods discussed herein, the system can generate a map of the ego's surroundings. Using the voxel data (e.g., the coordinates of each voxel) and the corresponding occupancy states, the AI ​​model (or, in some cases, another model using data predicted by the AI ​​model) can generate a map of the ego's surroundings.

[0033] FIG. 1A is a non-limiting example of system components that may implement the methods and systems discussed herein. For example, an analytics server may train an AI model and use the trained AI model to generate one or more ego occupancy datasets and / or maps. FIG. 1A illustrates components of an AI-enabled visual data analysis system 100. System 100 may include analytics server 110a, system database 110b, administrator computing device 120, egos 140a-b (collectively, egos 140), ego computing devices 141a-c (collectively, ego computing devices 141), and server 160. System 100 is not limited to the components described herein and may include additional or other components not shown for simplicity that are considered within the scope of the embodiments described herein.

[0034] The above components may be connected through a network 130. Examples of the network 130 may include, but are not limited to, a private or public LAN, a WLAN, a MAN, a WAN, and the Internet. The network 130 may include wired and / or wireless communications conforming to one or more standards and / or over one or more transport media.

[0035] Communications over network 130 may be performed according to various communications protocols, such as Transmission Control Protocol / Internet Protocol (TCP / IP), User Datagram Protocol (UDP), and IEEE communications protocols. In one embodiment, network 130 may include wireless communications conforming to the Bluetooth set of specifications or other standard or proprietary wireless communications protocols. In another embodiment, network 130 may also include communications over cellular networks, including, for example, Global System for Mobile Communications (GSM), Code Division Multiple Access (CDMA), or Enhanced Data for Global Evolution (EDGE) networks.

[0036] System 100 illustrates one example system architecture and components that can be used to train and run one or more AI models, such as AI model(s) 110c. Specifically, as depicted in FIG. 1A and described herein, analytics server 110a can train AI model(s) 110c with data obtained from ego 140 (e.g., using data streams 172 and 174) using methods discussed herein. Once AI model(s) 110c are trained, each of ego 140 can access and run trained AI model(s) 110c. For example, vehicle 140a with ego computing device 141a can send the vehicle's camera feed to trained AI model(s) 110c, which can determine its surrounding occupancy conditions (e.g., data stream 174). Furthermore, data captured and / or predicted by AI model(s) 110c with respect to ego 140 (during inference) can also be used to improve AI model(s) 110c. Thus, system 100 depicts a continuous loop that can periodically improve the accuracy of AI model(s) 110c. Furthermore, system 100 depicts a loop that allows ego 140 to use received data in a training phase in addition to the inference phase.

[0037] The analytics server 110a may be configured to collect, process, and analyze navigation data (e.g., images captured during navigation) and various sensor data collected from the ego 140. The collected data may then be processed and prepared as a training dataset. The training dataset may then be used to train one or more AI models, such as AI model 110c. Additionally, the analytics server 110a may be configured to collect visual data from the ego 140. Using the AI ​​model 110c (trained by the methods and systems discussed herein), the analytics server 110a may generate a dataset and / or an occupancy map for the ego 140. The analytics server 110a may display the occupancy map on the ego 140 and / or transmit the occupancy map / dataset to the ego computing device 141, the administrator computing device 120, and / or the server 160.

[0038] In FIG. 1A, the AI ​​model 110c is illustrated as a component of the system database 110b, but the AI ​​model 110c may be stored in a different or separate component, such as cloud storage or any other data repository accessible to the analytics server 110a.

[0039] Additionally, the analytic server 110a may be configured to display an electronic platform showing various training attributes for training the AI ​​model 110c. The electronic platform may be displayed on the administrator computing device 120 to allow an analyst to monitor the training of the AI ​​model 110c. One exemplary electronic platform generated and hosted by the analytic server 110a may be a web-based application or website configured to display the training dataset collected from the ego 140 and / or the training status / metrics of the AI ​​model 110c.

[0040] The analytic server 110a can be any computing device including a processor and non-transitory machine-readable storage capable of performing the various tasks and processes described herein. Non-limiting examples of such computing devices can include a workstation computer, a laptop computer, a server computer, etc. Although the system 100 includes a single analytic server 110a, the system 100 can include any number of computing devices operating in a distributed computing environment, such as a cloud environment.

[0041] Ego 140 may represent various electronic data sources that transmit data related to previous or current navigation sessions to analytics server 110a. Ego 140 may be any device configured for navigation, such as vehicle 140a and / or truck 140c. Ego 140 is not limited to vehicles and may include robotic devices as well. For example, ego 140 may include robot 140b, which may be a versatile, bipedal, autonomous humanoid robot capable of navigating various terrains. Robot 140b may be equipped with software that enables balance, navigation, perception, or interaction with the physical world. Additionally, robot 140b may include various cameras configured to transmit visual data to analytics server 110a.

[0042] Although referred to herein as an “ego,” ego 140 may or may not be an autonomous device configured for autonomous navigation. For example, in some embodiments, ego 140 may be controlled by a human operator or a remote processor. Ego 140 may include various sensors, such as those depicted in FIG. 1B . The sensors may be configured to collect data as ego 140 navigates various terrains (e.g., roads). Analytics server 110a may collect data provided by ego 140. For example, analytics server 110a may obtain navigation session and / or road / terrain data (e.g., images of ego 140 navigating roads) from various sensors, such that the collected data is ultimately used by AI model 110c for training purposes.

[0043] As used herein, a navigation session corresponds to a trip in which ego 140 travels a route, regardless of whether the trip is autonomous or human-controlled. In some embodiments, the navigation session may be for data collection and model training purposes. However, in some other embodiments, ego 140 may refer to a vehicle purchased by a consumer, and the purpose of the trip is classified as daily use. A navigation session may begin when ego 140 travels more than a threshold distance (e.g., 0.1 miles, 100 feet) from a non-moving location or exceeds a threshold speed (e.g., 0 miles per hour or greater, 1 mile per hour or greater, 5 miles per hour or greater). A navigation session may end when ego 140 returns to a non-moving location and / or is turned off (e.g., the driver exits the vehicle).

[0044] The ego 140 may represent a collection of egos monitored by the analytics server 110a to train the AI ​​model(s) 110c. For example, the driver of the vehicle 140a may authorize the analytics server 110a to monitor data associated with their respective vehicle. As a result, the analytics server 110a can collect sensor / camera data using various methods discussed herein and, accordingly, generate a training dataset for training the AI ​​model(s) 110c. The analytics server 110a can then apply the trained AI model(s) 110c to analyze the data associated with the ego 140 and predict an occupancy map for the ego 140. Furthermore, additional / ongoing data associated with the ego 140 can also be processed and added to the training dataset, and the analytics server 110a can recalibrate the AI ​​model(s) 110c accordingly. Thus, the system 100 depicts a loop in which the AI ​​model(s) 110c can be trained using navigation data received from the ego 140. Ego 140 can include a processor that executes trained AI models 110c for navigation purposes. During navigation, ego 140 can collect additional data about these navigation sessions and use this additional data to calibrate AI models 110c. That is, ego 140 represents an ego that can be used to train, run / use, and recalibrate AI models 110c. As a non-limiting example, ego 140 represents a vehicle purchased by a customer that uses AI models 110c to navigate autonomously while simultaneously improving AI models 110c.

[0045] Ego 140 may be equipped with various technologies that enable it to gather data from its surroundings and (potentially) navigate autonomously. For example, ego 140 may be equipped with an inference chip for running self-driving software.

[0046] Various sensors in each ego 140 can monitor and transmit collected data related to various navigation sessions to analytics server 110a. FIGS. 1B-1C illustrate block diagrams of sensors integrated within ego 140, according to one embodiment. The number and location of each sensor discussed with respect to FIGS. 1B-1C may depend on the type of ego discussed in FIG. 1A. For example, robot 140b may include different sensors than vehicle 140a or truck 140c. For example, robot 140b may not include airbag activation sensor 170q. Additionally, sensors in vehicle 140a and truck 140c may be located in different locations than those shown in FIG. 1C.

[0047] As discussed herein, various sensors integrated within each ego 140 may be configured to measure various data relevant to each navigation session. Analytics server 110a may periodically collect data monitored and collected by these sensors, which is processed according to methods described herein and used to train AI model 110c and / or run AI model 110c to generate occupancy maps.

[0048] Ego 140 may include user interface 170a. User interface 170a may refer to the user interface of an ego computing device (e.g., ego computing device 141 of FIG. 1A). User interface 170a may be implemented as a display screen integrated with or coupled to the vehicle interior, a head-up display, a touchscreen, or the like. User interface 170a may include input devices such as a touchscreen, knobs, buttons, a keyboard, a mouse, a gesture sensor, a steering wheel, or the like. In various embodiments, user interface 170a may be adapted to provide user input (e.g., as certain types of signals and / or sensor information) to other devices of ego 140, such as controller 170c, or to sensors (e.g., the sensors illustrated in FIG. 1B).

[0049] Furthermore, user interface 170a may be implemented by one or more logic devices that may be adapted to execute instructions, such as software instructions, that implement any of the various processes and / or methods described herein. For example, user interface 170a may be adapted to form a communication link, send and / or receive communications (e.g., sensor signals, control signals, sensor information, user input, and / or other information), or perform various other processes and / or methods. In another example, a driver may use user interface 170a to control the temperature of ego 140 or activate its functions (e.g., autonomous driving or steering system 170o). Thus, user interface 170a may monitor and collect driving session data in conjunction with other sensors described herein. Furthermore, user interface 170a may be configured to display various data generated / predicted by analytics server 110a and / or AI model 110c.

[0050] Orientation sensor 170b may be implemented as one or more of a compass, float, accelerometer, and / or other digital or analog device capable of measuring the orientation of ego 140 (e.g., the magnitude and direction of roll, pitch, and / or yaw relative to one or more reference orientations, such as gravity and / or magnetic north). Orientation sensor 170b may be adapted to provide orientation measurements of ego 140. In other embodiments, orientation sensor 170b may be adapted to provide roll, pitch, and / or yaw rate of ego 140 using a time series of orientation measurements. Orientation sensor 170b may be positioned and / or adapted to provide orientation measurements relative to a particular coordinate frame of ego 140.

[0051] Controller 170c may be implemented as any suitable logic device (e.g., a processing device, microcontroller, processor, application specific integrated circuit (ASIC), field programmable gate array (FPGA), memory storage device, memory reader, or other device or combination of devices) that may be adapted to execute, store, and / or receive appropriate instructions, such as software instructions that implement control loops for controlling various operations of ego 140. Such software instructions may also implement methods for processing sensor signals, determining sensor information, providing user feedback (e.g., through user interface 170a), interrogating a device for operating parameters, selecting operating parameters for a device, or performing any of the various operations described herein.

[0052] Communications module 170e may be implemented as any wired and / or wireless interface configured to communicate sensor data, configuration data, parameters, and / or other data and / or signals to any of the functions shown in FIG. 1A (e.g., analytics server 110a). As described herein, in some embodiments, communications module 170e may be implemented in a distributed manner, such that portions of communications module 170e are implemented within one or more elements and sensors shown in FIG. 1B. In some embodiments, communications module 170e may delay communication of sensor data. For example, if ego 140 does not have network connectivity, communications module 170e may store sensor data in temporary data storage and transmit the sensor data when ego 140 is identified as having adequate network connectivity.

[0053] Speed ​​sensor 170d may be implemented as an electronic pitot tube, a metering gear or wheel, a water speed sensor, a wind speed sensor, a wind speed sensor (e.g., direction and magnitude), and / or other device capable of measuring or determining the linear speed of ego 140 (e.g., in the surrounding medium and / or aligned with the longitudinal axis of ego 140) and further providing such measurements as sensor signals that can be communicated to various devices.

[0054] Gyroscope / accelerometer 170f may be implemented as one or more electronic sextants, semiconductor devices, integrated chips, accelerometer sensors, or other systems or devices capable of measuring angular velocity / acceleration and / or linear acceleration (e.g., direction and magnitude) of ego 140 and providing such measurements as sensor signals that can be communicated to other devices, such as analytics server 110a. Gyroscope / accelerometer 170f may be positioned and / or adapted to make such measurements relative to a particular coordinate frame of ego 140. In various embodiments, gyroscope / accelerometer 170f may be implemented within a common housing and / or module with other elements depicted in FIG. 1B to ensure a common frame of reference or known transformations between frames of reference.

[0055] Global navigation satellite system (GNSS) 170h may be implemented as a global positioning satellite receiver and / or another device capable of determining the absolute and / or relative position of ego 140, for example, based on radio signals received from spacecraft-based and / or terrestrial sources, and further capable of providing such measurements as sensor signals that may be communicated to various devices. In some embodiments, GNSS 170h may be adapted to determine the velocity, speed, and / or yaw rate of ego 140 (e.g., using a time series of position measurements), such as the absolute velocity and / or the yaw component of angular velocity of ego 140.

[0056] Temperature sensor 170i may be implemented as a thermistor, an electrical sensor, an electrical thermometer, and / or other device capable of measuring a temperature associated with ego 140 and providing such a measurement as a sensor signal. Temperature sensor 170i may be configured to measure an environmental temperature associated with ego 140, such as the temperature of a cockpit or dash, which may be used to estimate the temperature of one or more elements of ego 140.

[0057] Humidity sensor 170j may be implemented as a relative humidity sensor, an electrical sensor, an electrical relative humidity sensor, and / or other device capable of measuring the relative humidity associated with ego 140 and providing such measurement as a sensor signal.

[0058] Steering sensor 170g may be adapted to physically adjust the heading of ego 140 according to one or more control signals provided by a logic device, such as controller 170c, and / or user input. Steering sensor 170g may include one or more actuators and control surfaces (e.g., rudder or other type of steering or trim mechanism) of ego 140 and may be adapted to physically adjust the control surfaces to various positive and / or negative steering angles / positions. Further, steering sensor 170g may be adapted to sense the current steering angle / position of such steering mechanism and provide such measurements.

[0059] Propulsion system 170k may be implemented as a propeller, turbine, or other thrust-based propulsion system, a mechanical wheel and / or tracked propulsion system, a wind / sail-based propulsion system, and / or any other type of propulsion system usable to provide motive power to ego 140. Propulsion system 170k may further monitor the motive force and / or direction of thrust of ego 140 relative to a coordinate frame of reference of ego 140. In some embodiments, propulsion system 170k may be coupled to and / or integrated with steering sensor 170g.

[0060] Occupant restraint sensor 170l can monitor the seat belt detection and lock / unlock assembly, as well as other occupant restraint subsystems. Occupant restraint sensor 170l can include various environmental and / or status sensors, actuators, and / or other devices that assist in the operation of safety mechanisms associated with the operation of ego 140. For example, occupant restraint sensor 170l can be configured to receive motion and / or status data from other sensors depicted in FIG. 1B. Occupant restraint sensor 170l can determine whether a safety measure (e.g., a seat belt) is engaged.

[0061] As depicted in FIG. 1C , camera 170m may refer to one or more cameras integrated into ego 140 and may include multiple cameras integrated into (or retrofitted to) ego 140. Camera 170m may be an interior-facing or exterior-facing camera of ego 140. For example, as depicted in FIG. 1C , ego 140 may include one or more interior-facing cameras that may monitor and collect video of the occupants of ego 140. Ego 140 may include eight exterior-facing cameras. For example, ego 140 may include a front camera 170m-1, a forward-facing side camera 170m-2, a forward-facing side camera 170m-3, a rearward-facing side camera 170m-4 on each front fender, a camera 170m-5 on each side (e.g., integrated into a B-pillar), and a rear camera 170m-6.

[0062] 1B, radar 170n and ultrasonic sensor 170p may be configured to monitor the distance of ego 140 relative to other objects, such as other vehicles or immovable objects (e.g., trees or garage doors). Additionally, ego 140 may include an autonomous driving system, or steering system 170o, configured to autonomously navigate ego 140 using data collected via various sensors (e.g., radar 170n, speed sensor 170d, and / or ultrasonic sensor 170p).

[0063] Thus, the autonomous driving or steering system 170o can analyze various data collected by one or more sensors described herein to identify driving data. For example, the autonomous driving or steering system 170o can calculate the risk of a forward collision based on the speed of the ego 140 and the distance to other vehicles on the road. Additionally, the autonomous driving or steering system 170o can determine whether the driver is touching the steering wheel. The autonomous driving or steering system 170o can transmit the analyzed data to various functions discussed herein, such as an analytics server.

[0064] Airbag activation sensor 170q may predict or detect a crash and cause the activation or deployment of one or more airbags. Airbag activation sensor 170q may transmit data regarding the deployment of the airbags, including data related to the event that caused the deployment.

[0065] 1A , the administrator computing device 120 may represent a computing device operated by a system administrator. The administrator computing device 120 may be configured to display data obtained or generated by the analytics server 110a (e.g., various analytics metrics and risk scores), allowing the system administrator to monitor various models utilized by the analytics server 110a, review feedback, and / or assist in training the AI ​​models 110c maintained by the analytics server 110a.

[0066] Ego(s) 140 may be any device configured to navigate various routes, such as vehicle 140a or robot 140b. As discussed with respect to FIGS. 1B-1C, ego 140 may include various telemetry sensors. Additionally, ego 140 may include ego computing device 141. Specifically, each ego may have its own ego computing device 141. For example, truck 140c may have ego computing device 141c. For ease of explanation, ego computing devices are collectively referred to as ego computing device(s) 141. Ego computing device 141 may control the presentation of content on the infotainment system of ego 140, process commands related to the infotainment system, aggregate sensor data, manage communication of data to electronic data sources, receive updates, and / or send messages. In one configuration, ego computing device 141 communicates with an electronic control unit. In another configuration, ego computing device 141 is an electronic control unit. Ego computing device 141 may include a processor and non-transitory machine-readable storage media that enable it to perform various tasks and processes described herein. For example, AI model(s) 110c described herein may be stored and executed (or directly accessed) by ego computing device 141. Non-limiting examples of ego computing device 141 may include vehicle multimedia and / or display systems.

[0067] In one example of how AI models 110c may be trained, analytic server 110a may collect data from ego 140 to train AI models 110c. Before running AI models 110c to generate / predict occupancy data sets, analytic server 110a may train AI models 110c using various methods. This training enables AI models 110c to capture data from one or more cameras of one or more ego 140 (without having to receive radar data) and predict occupancy data around the ego. The operations described in this example may be performed by any number of computing devices (e.g., processors of ego 140) operating in the distributed computing system described in FIGS. 1A and 1B.

[0068] To train the AI ​​model(s) 110c, the analytics server 110a may first employ one or more egos 140 to drive a particular route. While driving, the ego 140 may generate navigation session data using one or more sensors (including one or more cameras) of the ego. For example, one or more of the egos 140 equipped with various sensors may navigate a specified route. As the one or more egos 140 traverse the terrain, these sensors may capture continuous (or periodic) data of their surroundings. The sensors may indicate the occupancy status of the one or more egos' surroundings 140. For example, the sensor data may indicate various objects having mass around the one or more egos 140 as they navigate the route.

[0069] The analytics server 110a can generate a training dataset using data collected from the ego 140 (e.g., a camera feed received from the ego 140). The training dataset can indicate the occupancy state of different voxels within the periphery of one or more egos 140. As used herein, in some embodiments, a voxel is a three-dimensional pixel that forms the building block of the periphery of one or more egos 140. Within the training dataset, each voxel can encapsulate sensor data that indicates whether a mass has been identified for that particular voxel. As used herein, mass can indicate or represent any object identified using a sensor. For example, in some embodiments, the ego 140 can be equipped with a sensor that can identify masses in the vicinity of the ego 140.

[0070] In some embodiments, the training data set may include data received from cameras of ego 140. The data received from the cameras may have a set of data points, each data point corresponding to the location and image attributes of at least one voxel in the space surrounding ego 140. Additionally, the training data set may include three-dimensional geometry data that indicates whether one or more voxels surrounding ego 140 are occupied by an object having mass.

[0071] In operation, as one or more egos 140 navigate, these sensors collect data and transmit the data, as depicted by data stream 172, to analysis server 110a.

[0072] In some embodiments, one or more egos 140 may include one or more high-resolution cameras that capture a continuous stream of visual data from the surroundings of the one or more egos 140 as the one or more egos 140 navigate through a route. The analysis server 110a can then use the camera feeds to generate a second data set, with visual elements / depictions of various voxels of the surroundings 140 of the one or more egos included in the second data set.

[0073] In operation, as one or more egos 140 navigate, these cameras collect data and transmit the data to analysis server 110a, as depicted by data stream 172. For example, ego computing device 141 can use data stream 172 to transmit image data to analysis server 110a.

[0074] The analysis server 110a can use the first and second data sets to train an AI model such that the AI ​​model 110c correlates each data point in the first set of data points with a corresponding data point in the second set of data points using the respective positions of each data point to train itself, and after training, the AI ​​model 110c is configured to receive a camera feed from the new ego 140 and predict the occupancy state of at least one voxel in the camera feed.

[0075] Using the first and second data sets, analytics server 110a can train AI model(s) 110c so that AI model(s) 110c correlate various visual attributes of a voxel (in a camera feed in the second data set) with the occupancy state of that voxel (in the first data set). In this way, after training, AI model(s) 110c can receive a camera feed (e.g., from new ego 140) without receiving sensor data and then determine the occupancy state of each voxel of new ego 140.

[0076] The analysis server 110a can generate a training dataset including a first and a second dataset. The analysis server 110a can use the first dataset as ground truth. For example, the first dataset can indicate different locations of voxels and their occupancy states. The second dataset can include a visual (e.g., camera feed) representation of the same voxels. The analysis server 110a can use the first dataset to label the data such that data records associated with each voxel corresponding to an object are indicated as having a positive occupancy state.

[0077] Labeling the occupancy states of various voxels can be performed automatically and / or manually. For example, in some embodiments, the analytics server 110a can use a human reviewer to label the data. For example, as discussed herein, camera feeds from one or more cameras on a vehicle can be displayed and labeled to a human reviewer on an electronic platform. Additionally or alternatively, the data can be ingested in its entirety into AI model(s) 110c, which identify corresponding voxels, analyze the first digital map, and further associate the images of each voxel with a respective occupancy state.

[0078] The AI ​​model(s) 110c can be trained using ground truth so that the visual elements of each voxel are analyzed and correlated to whether the voxel is occupied by mass. Thus, the AI ​​model 110c can obtain the occupancy state of each voxel (using a first data set) and use that information as ground truth. Additionally, the AI ​​model(s) 110c can obtain the visual attributes of the same voxels using a second data set.

[0079] In some embodiments, analytics server 110a may use supervised training methods. For example, AI models 110c may use ground truth and received visual data to train themselves to predict the occupancy state of a voxel based solely on images of that voxel. As a result, during training, AI models 110c may receive camera feeds, analyze the camera feeds, and determine the occupancy state of each voxel in the camera feeds (without the need to use radar).

[0080] The analytic server 110a can send a set of training data sets to the AI ​​model(s) 110c and obtain a set of predicted outputs (e.g., predicted occupancy states). The analytic server 110a can then train the AI ​​model(s) 110c by comparing the predicted data with ground truth data to determine differences and adjusting the internal weights and parameters of the AI ​​model(s) 110c in proportion to the determined differences according to a loss function. The analytic server 110a can train the AI ​​model(s) 110c in a similar manner until the predictions of the trained AI model(s) 110c are accurate up to a certain threshold (e.g., recall or precision).

[0081] Additionally or alternatively, the analytic server 110a can use unsupervised methods in which the training dataset is unlabeled. Because labeling the data in the training dataset can be time-consuming and require excessive computational power, the analytic server 110a can use unsupervised training techniques to train the AI ​​model 110c.

[0082] After AI model 110c is trained, it can be used by ego 140 to predict occupancy data around one or more egos 140. For example, AI model(s) 110c can divide the ego's surroundings into various voxels and predict an occupancy state for each voxel. In some embodiments, AI model(s) 110c (or analytic server 110a using data predicted using AI model 110c) can generate an occupancy map, or occupancy network, representing the surroundings of one or more egos 140 at any given time.

[0083] In another example of how AI models 110c may be used, after training AI models 110c, analytic server 110a (or a local chip on ego 140) may collect data from egos (e.g., one or more egos 140) and predict occupancy data sets for one or more egos 140. This example describes how AI models 110c may be used to predict occupancy data in real time or near real time for one or more egos 140. In this configuration, a processor, such as analytic server 110a, may execute the AI ​​models. However, one or more actions may be executed locally, for example, via a chip located within one or more egos 140. In operation, an ego 140 may execute AI models 110c locally and use the results to navigate itself autonomously.

[0084] The processor can input image data of the space around the ego object 140 to the AI ​​model 110c using the cameras of the ego object 140. The processor can collect and / or analyze data received from various cameras (e.g., outward-facing cameras) of one or more of the ego 140. In another embodiment, the processor can collect and aggregate video footage recorded by one or more cameras of the ego 140. The processor can then send the video footage to the AI ​​model(s) 110c trained using the methods discussed herein.

[0085] The processor can predict occupancy attributes of multiple voxels by executing AI model 110c. AI model(s) 110c can predict occupancy states of various voxels surrounding one or more egos 140 using received image data using methods discussed herein.

[0086] The processor can generate a dataset based on the plurality of voxels and their corresponding occupancy attributes. The analysis server 110a can generate a dataset including occupancy states of various voxels according to their respective coordinate values. The dataset can be a queryable dataset available for sending predicted occupancy states to various software modules.

[0087] In operation, one or more ego 140 may collect image data from their cameras and transmit the image data to a processor (located locally on one or more ego 140) and / or to analytics server 110a, as depicted by data stream 172. The processor may then execute AI model(s) 110c to predict occupancy data for one or more ego 140. If the prediction is performed by analytics server 110a, the occupancy data may be transmitted to one or more ego 140 using data stream 174. If a processor is located locally within one or more ego 140, the occupancy data is transmitted to ego computing device 141 (not shown in FIG. 1A ).

[0088] Training of AI model(s) 110c may be performed using methods discussed herein, such that execution of AI model(s) 110c may be performed locally (during inference) on any of ego 140. Collected data (e.g., navigation data collected during navigation of ego 140, such as image data of the trip) may then be fed back to AI model(s) 110c, allowing the additional data to improve AI model(s) 110c.

[0089] FIG. 2 illustrates a flow diagram of a method 200 performed in an AI-enabled visual data analysis system, according to one embodiment. Method 200 may include steps 210-270. However, other embodiments may include additional or alternative steps, or omit one or more steps. Method 200 is performed by an analysis server (e.g., a computer similar to analysis server 110a). However, one or more steps of method 200 may be performed by any number of computing devices (e.g., processors of ego 140 and / or ego computing device 141) operating in the distributed computing system described in FIGS. 1A-C. For example, one or more computing devices of an ego may locally perform some or all of the steps described in FIG. 2.

[0090] 2 illustrates a model architecture for how image input is captured from the ego (step 210) and can be analyzed (step 270) to predict queryable output. Using the methods and systems discussed herein, the analysis server can capture only image data (e.g., camera feeds from the ego's surroundings) to generate queryable output. Thus, the methods and systems discussed herein can operate even in the absence of data received from radar, LiDAR, and the like.

[0091] The queryable output (generated in step 270) can be used for a variety of purposes. In one embodiment, the queryable output may be available to an autonomous driving module, where various navigation decisions may be made based on whether voxels in the space surrounding the ego are predicted to be occupied. In another embodiment, the analytics server may use the queryable output to generate a digital map showing the occupancy status of the ego's surroundings. For example, the analytics server may generate a three-dimensional (3D) geometric representation of the ego's surroundings. The digital map may be displayed, for example, on the ego's computing device.

[0092] As used herein, a voxel may refer to a volumetric pixel, which may be a three-dimensional pixel equivalent to a two-dimensional pixel. Thus, a voxel may represent a volumetric space around (e.g., surrounding) the ego, or a point defined in a three-dimensional grid within the environment. In some embodiments, the space surrounding the ego may be referred to as a voxel grid and may be divided into various voxels. As used herein, a voxel grid may refer to a set of cubes stacked (or arranged) together to represent objects in the space surrounding the ego. Each voxel may store information about a specific location within the ego's surrounding space. The methods and systems discussed herein can be used to evaluate the occupancy of each voxel. For example, an analytics server can determine (using the AI ​​models discussed herein) whether an object with mass occupies each voxel. The voxel predictions may be aggregated into a dataset, referred to herein as a queryable result. Using the queryable results, the voxel information can be queried by the processor or a downstream software module (e.g., autonomous driving software / processor) to identify occupancy data around the ego.

[0093] In some embodiments, a voxel may be designated as occupied if any portion of it is occupied. Thus, in some embodiments, each voxel may include a binary designation of 0 (unoccupied) or 1 (occupied). Alternatively, in some embodiments, the AI ​​model may also predict detailed occupancy data inside / within a particular voxel. For example, a voxel with a binary value of 1 (occupied) may be further analyzed at a more granular level, such that the occupancy of each point within the voxel is also determined. For example, an object may be curved. Some voxels (associated with the object) may be fully occupied, while some other voxels may be partially occupied. These voxels may be divided into smaller voxels such that some of the smaller voxels are unoccupied. As described herein, this method can be used to identify the shape of an object.

[0094] Method 200 begins at step 210, where image data is received from one or more cameras of the ego. Method 200 visually illustrates how an AI model (trained using the methods discussed herein) can ingest the image data and generate a queryable output that can indicate the volume occupancy of various voxels within the ego's perimeter. Image data may refer to any data received from one or more images of the ego.

[0095] The captured image data can then be characterized (step 220). Image characterizers, or various characterization algorithms, can be used to extract relevant and meaningful features from the received image data. Image characterizers can be used to convert the image data into a data representation that captures important information about the content of the image, allowing the image data to be analyzed more efficiently.

[0096] In some embodiments, an AI model can perform the characterizations discussed herein. In some other embodiments, a convolutional neural network can be used to characterize the image data. As one non-limiting example, a RegNet (Regularized Neural Network) can be used to transform the data into a BiFPN (Bidirectional Feature Pyramid Network) as shown. However, other protocols can also be used. In some other embodiments, a transformer can be used to characterize the image data.

[0097] After the image data has been encoded / characterized, a transformer can be used to change the image data from a two-dimensional image to a three-dimensional image (step 230). As discussed herein, in one example configuration, there may be eight separate cameras in communication with the ego. As a result, the image data may include eight separate camera feeds (one feed corresponding to each camera or other sensor) and may include overlapping views. The transformer can aggregate these individual camera feeds and generate one or more three-dimensional representations using the received camera feeds.

[0098] The Transformer can take in three separate inputs: an image key, an image value, and a three-dimensional query. The image key and image value may refer to attributes associated with the two-dimensional image data received from the ego. For example, these values ​​may be output by image characterization (step 220). The Transformer can also use an image query from three-dimensional space. The rendered spatial attention module can analyze the two-dimensional image key and image value using the three-dimensional query. As shown, the BiFPN generated in step 220 can be aggregated into a multi-camera query embedding and used to perform three-dimensional spatial queries. In some embodiments, each voxel can have its own query. Using the three-dimensional spatial query, the analysis server can identify regions in the two-dimensional characterized image that correspond to specific portions of the three-dimensional representation. The identified regions in the characterized image can then be analyzed to transform the multi-camera image data into three-dimensional representations of each voxel, which can generate a three-dimensional representation of the ego's surroundings. Thus, the rendered spatial attention module can output a single three-dimensional vector space representing the ego's surroundings. This effectively moves all of the image data generated by all of the camera feeds into the entire space around the ego, or into a three-dimensional spatial representation.

[0099] Steps 210-230 may be performed for each video frame received from each camera on the ego. For example, at each timestamp, steps 210-230 may be performed for eight separate images received from eight different cameras on the ego. As a result, at each timestamp, method 200 may generate one three-dimensional spatial representation of the eight images. At step 240, method 200 may fuse the three-dimensional images (for the various timestamps). This fusion may be performed based on the timestamps of each set of images. For example, the three-dimensional spatial representations may be fused (e.g., in a sequential manner) based on their respective timestamps.

[0100] As shown, the 3D spatial representation at timestamp t can be merged with the 3D spatial representation of the ego's surroundings at t-1, t-2, and t-3. As a result, the output can have both spatial and temporal information. This concept is depicted in Figure 2 as spatial-temporal features.

[0101] Next, deconvolution can be used to convert the spatial-temporal features into various voxels (step 250). As discussed herein, various data points are characterized and fused. At this step 250, method 200 can perform various mathematical operations to reverse this process so that the fused data can be converted back into various voxels. As used herein, deconvolution can refer to a mathematical operation used to reverse the effect of convolution.

[0102] After applying deconvolution to the (characterized, transformed, and fused) image data, method 200 can then apply various trained AI modeling techniques discussed herein (e.g., FIGS. 3-4) to generate volumetric output (step 260). The volumetric output can include binary data for various voxels indicating whether a particular voxel is occupied by an object having mass. Specifically, the volumetric output can include occupancy data, including binary data indicating whether the voxel is occupied, and / or occupancy flow data, indicating how fast the voxel is moving (if present, where velocity is calculated using temporal alignment).

[0103] The volume output may also include shape information (the shape of the mass occupying the voxel). In some embodiments, the size of each voxel may be predetermined, but the size can be varied to obtain more granular results. For example, the default size of various voxels may be 33 centimeters (at each vertex). While this size is generally acceptable for voxels, reducing the size of the voxels can improve results. For example, if a voxel is detected to be outside the ego's driving surface, a 33 cm voxel may be appropriate. However, the analysis server may reduce the size of voxels that are occupied within a threshold distance from the ego and / or the ego's driving surface (e.g., to 10 cm). Once voxel occupancy data is identified, a regression model can be run to identify the shape of the voxels. For example, a 33 cm voxel (belonging to a curb) may only be half occupied (e.g., only 16 cm occupied). The analysis server can use regression to determine the degree to which a voxel is occupied.

[0104] Additionally or alternatively, the analysis server can decode the subvoxel values ​​to identify the shape of the subvoxel (inside the occupied voxel). For example, if a voxel is half-occupied, the analysis server can define a set of subvoxels and identify the volumetric output of the subvoxel using the methods discussed herein. Once the subvoxels are aggregated (back to the original voxel), the analysis server can determine the shape of the voxel. For example, each voxel may have eight vertices. In some embodiments, each vertex may be analyzed individually and have its embedding. As a result, any point within each vertex of a voxel can be queried individually. Therefore, in this "continuous resolution" approach, the analysis server may not define the size of the subvoxel. In some embodiments, the analysis server may use a multivariate interpolation (e.g., trilinear interpolation) protocol to estimate the occupancy state of each subvoxel and / or any point within each vertex.

[0105] The volumetric output may also include three-dimensional semantic data indicating the object occupying the voxel (or group of voxels). The three-dimensional semantic may indicate whether the voxel and / or neighboring voxels are occupied by a car, a road curb, a building, or other object. The three-dimensional semantic may also indicate whether the voxel is occupied by a stationary mass or a moving mass. The three-dimensional semantic data may be identified using various temporal attributes of the voxels. For example, if a group of voxels is identified as being occupied by a mass, the collective shape of the voxels may indicate that the voxels belong to a vehicle. If, at a previous timestamp, the identified group of voxels (now known to be a vehicle) was identified as moving, the group of voxels may have a three-dimensional semantic indicating that the group of voxels belongs to a moving vehicle. In another example, if a group of voxels is identified as having a shape corresponding to a curb and is not identified as having motion, then that group of voxels may have a three-dimensional semantic that indicates a static curb.

[0106] In some embodiments, certain shapes or three-dimensional semantics may be prioritized. For example, certain objects, such as other vehicles on the road or objects related to the driving surface (e.g., curbs indicating the outer limits of the road), may be analyzed thoroughly. In contrast, the details of static objects, such as nearby buildings far from the ego's driving surface, may not be analyzed as thoroughly as vehicles moving near the ego. In some embodiments, certain objects with certain sizes or shapes may be ignored. For example, road debris may not be analyzed as thoroughly as vehicles moving near the ego.

[0107] In some embodiments, object-level detection may not need to be performed by method 200. For example, the ego needs to navigate to avoid voxels in front of it that are identified as static and occupied, regardless of whether the voxels belong to other vehicles, pedestrians, or traffic signs. Thus, occupancy information may not be object dependent. In some embodiments, object detection models may be run independently (e.g., in parallel) to be able to detect objects corresponding to different groups of voxels.

[0108] At step 270, method 200 can generate a queryable dataset so that other software modules can query the occupancy state of various voxels. For example, a software module can transmit coordinate values ​​(X, Y, Z axes) around the ego and receive any of the four categories of occupancy data (e.g., volumetric output) generated using method 200. The queryable dataset can be used to generate an occupancy map (e.g., FIGS. 3A-B) or to make autonomous navigation decisions for the ego.

[0109] Additionally or alternatively, the analysis server can generate a map corresponding to the predicted occupancy state of various voxels. As a non-limiting example, the analysis server can visualize each voxel and its occupancy state using a multi-view 3D reconstruction protocol. A non-limiting example of a map, or occupancy map, is presented in FIGS. 3A-B (e.g., simulation 350). In some embodiments, simulation 350 can be displayed on the ego's user interface. Simulation 350 can exemplify the camera feed 300 depicted in FIG. 3A. Camera feed 300 represents image data received (whether in real time or near real time) from eight different cameras on ego. Specifically, camera feed 300 may include camera feeds 310a-c received from three different front-facing cameras on the ego, camera feeds 320a-b received from two different right-facing cameras on the ego, camera feeds 330a-b received from two different left-facing cameras on the ego, and camera feed 340 received from a rear-facing camera on the ego.

[0110] Using the methods discussed herein, the analysis server can analyze the camera feed 300, divide the space surrounding the ego into voxels, and generate a graphical representation of the ego's surroundings, called simulation 350 (depicted in FIG. 3B). Simulation 350 can include a simulated ego (360) and its surrounding voxels. For example, simulation 350 can include graphical indicators for the various masses occupying the various voxels surrounding simulated ego 360. For example, simulation 350 can include simulated masses 370a-c.

[0111] Each simulated mass 370a-c may represent an object depicted in camera feed 300. For example, simulated mass 370a may correspond to mass 380a (a vehicle), simulated mass 370b may correspond to mass 380b (a vehicle), and simulated mass 370c may correspond to mass 380c (a building near a road). As shown, all simulated masses include various voxels. Furthermore, voxels depicted in simulation 350 may have distinct graphical / visual characteristics corresponding to their volumetric output (e.g., occupancy data). For example, simulated mass 370c (e.g., a building) may have a first color indicating that the mass is identified as static. Similarly, simulated mass 370b (e.g., a vehicle) may have a second color indicating that it is a parked or stationary vehicle. Meanwhile, simulated mass 370a (eg, other vehicles) may have a third color and / or other visual characteristics that indicate that it is predicted to be moving.

[0112] Additionally or alternatively, the analytics server can send the generated map to a downstream software application or another server. The predicted results can be further analyzed and used with various models and / or algorithms to perform various actions. For example, a software model or processor associated with Ego's autonomous navigation system can receive occupancy data predicted by a trained AI model and make navigation decisions accordingly.

[0113] The various illustrative logical blocks, modules, circuits, and algorithm steps described in connection with the embodiments disclosed herein can be implemented as electronic hardware, computer software, or a combination of both. To clearly illustrate this interchangeability of hardware and software, the various illustrative components, blocks, modules, circuits, and steps have been described generally in terms of their functionality. Whether such functionality is implemented as hardware or software depends on the particular application and design constraints imposed on the overall system. Those skilled in the art may implement the described functionality in various ways for each particular application, and such implementation decisions should not be interpreted as causing a departure from the scope of the present disclosure or the claims.

[0114] Computer software-implemented embodiments may be implemented in software, firmware, middleware, microcode, hardware description languages, or any combination thereof. A code segment, or machine-executable instructions, may represent a procedure, a function, a subprogram, a program, a routine, a subroutine, a module, a software package, a class, or any combination of instructions, data structures, or program statements. A code segment may be coupled to another code segment or a hardware circuit by passing and / or receiving information, data, arguments, parameters, or memory contents. Information, arguments, parameters, data, etc. may be passed, forwarded, or transmitted via any suitable means including memory sharing, message passing, token passing, network transmission, etc.

[0115] The actual software code or specialized control hardware used to implement these systems and methods does not limit the claimed features or the present disclosure. Thus, although the operation and behavior of the systems and methods have been described without reference to specific software code, it will be understood that software and control hardware can be designed to implement the systems and methods based on the description herein.

[0116] If implemented as software, the functions may be stored as one or more instructions or code on a non-transitory computer-readable or processor-readable storage medium. The steps of a method or algorithm disclosed herein may be embodied in a processor-executable software module, which may reside in a computer-readable or processor-readable storage medium. Non-transitory computer-readable or processor-readable media includes both computer storage media and tangible storage media that facilitate transfer of a computer program from one place to another. A non-transitory processor-readable storage medium may be any available medium that can be accessed by a computer. By way of example, and not limitation, such non-transitory processor-readable media may include RAM, ROM, EEPROM, CD-ROM, or other optical disk storage, magnetic disk storage, or other magnetic storage devices, or any other tangible storage medium that can be used to store desired program code in the form of instructions or data structures and that can be accessed by a computer or processor. As used herein, disk and disc include compact discs (CDs), laser discs, optical discs, digital versatile discs (DVDs), Blu-ray discs, and floppy disks, where a "disk" typically reproduces data magnetically and a "disk" typically reproduces data optically with a laser. Additionally, combinations of the above are also intended to be included within the scope of computer-readable media. Additionally, the operations of a method or algorithm may reside as one or any combination or set of code and / or instructions on a non-transitory processor-readable medium and / or computer-readable medium, which may be incorporated into a computer program product.

[0117] The previous description of the disclosed embodiments is provided to enable any person skilled in the art to make or use the embodiments described herein, and variations thereof. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the principles defined herein may be applied to other embodiments without departing from the spirit or scope of the subject matter disclosed herein. Thus, the present disclosure is not intended to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the following claims and the principles and novel features disclosed herein.

[0118] While various aspects and embodiments have been disclosed, other aspects and embodiments are contemplated. The various disclosed aspects and embodiments are for purposes of illustration and not limitation, with the actual scope and spirit being indicated by the following claims.

Claims

1. inputting image data of the space around the ego object into an artificial intelligence model by a processor using a camera of the ego object; predicting occupancy attributes of a plurality of voxels by executing the artificial intelligence model with the processor; generating, by the processor, a dataset based on the plurality of voxels and their corresponding occupancy attributes; A method comprising:

2. 2. The method of claim 1, further comprising generating, by the processor, an output representing the ego object's environment and showing the plurality of voxels and their corresponding occupancy attributes, the output including a graphical indicator of occupancy attributes for at least some of the plurality of voxels.

3. The method of claim 2 , wherein the graphical indicator corresponds to a detected object associated with the at least some of the plurality of voxels.

4. The method of claim 2 further comprising the step of displaying, by said processor, said output on a screen associated with said ego object.

5. The method of claim 1 , wherein the dataset is a queryable dataset configured to transmit the occupancy attributes of the plurality of voxels to an autonomous driving protocol of the ego object.

6. The method of claim 1 , wherein the artificial intelligence model is trained using sensor attributes of the plurality of voxels.

7. The method of claim 1 , wherein the ego object is an autonomous vehicle that executes a driving protocol based on the dataset.

8. The method of claim 1 , further comprising the step of characterizing the image data by the processor prior to executing the artificial intelligence model.

9. The image data includes multiple camera feeds from multiple cameras of the ego object, and the method further comprises: The method of claim 1 , further comprising the step of temporally aligning, by the processor, the multiple camera feeds.

10. An ego object, A camera and a first processor; a second processor; and a non-transitory computer-readable medium storing an artificial intelligence model configured to be executed by the first processor; The first processor inputting image data of the space around the ego object into the artificial intelligence model using the camera of the ego object; running the artificial intelligence model to predict occupancy attributes of a plurality of voxels; generating a dataset based on the plurality of voxels and their corresponding occupancy attributes; The ego object, wherein the second processor is configured to use the data set to autonomously navigate the ego object.

11. The first processor further configured to generate an output representing the ego object's environment and indicative of the plurality of voxels and their corresponding occupancy attributes; The ego object of claim 10 , wherein the output includes a graphical indicator of the occupancy attribute for at least a portion of the plurality of voxels.

12. The ego object of claim 11 , wherein the graphical indicator corresponds to a detected object associated with the at least some of the plurality of voxels.

13. The first processor The ego object of claim 11 , further configured to display the output on a screen associated with the ego object.

14. The ego object of claim 10 , wherein the artificial intelligence model is trained using sensor attributes of the plurality of voxels.

15. The ego object of claim 10 , wherein the ego object is an autonomous vehicle that executes a driving protocol based on the data set.

16. training, by a processor, an artificial intelligence model using a training dataset including data received from a camera of the ego object; the training data set has a set of data points, each data point in the set of data points corresponding to at least one voxel location in space surrounding the ego object and an image attribute; whereby the artificial intelligence model correlates each data point in the first set of data points with a corresponding data point in the second set of data points using the respective location of each data point; Thus, once the artificial intelligence model has been trained, the artificial intelligence model is configured to receive a camera feed from a second ego object and predict a third set of data points, each data point in the third set of data points corresponding to an occupancy attribute indicating whether at least one voxel of space surrounding the second ego object is occupied by any object having mass.

17. The method of claim 16 , wherein the artificial intelligence model is further configured to generate an output representing the ego object's environment and indicative of the at least one voxel and their corresponding occupancy attributes.

18. 17. The method of claim 16, wherein the training data set further comprises a second set of data points, each data point in the second set of data points corresponding to the location and sensor attributes of at least one voxel in the space surrounding the ego object.

19. The method of claim 17 , wherein the graphical indicator corresponds to a detected object associated with at least a portion of the at least one voxel.

20. The method of claim 17 , wherein the artificial intelligence model uses a three-dimensional multi-view reconstruction protocol to generate the output.