Modeling techniques for vision-based path determination
Through the AI model, analyzing the voxel occupancy attributes of the autonomous surrounding environment is generated, and a 3D model is generated, which solves the problem of lack of position data in the indoor environment of autonomous navigation and realizes the ability to navigate to the destination independently.
Patent Information
- Application Number
- CN202380082611.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-09-30
- Filing Date
- 2023-09-29
- Publication Date
- 2025-07-25
AI Technical Summary
Existing autonomous navigation technologies rely on location data in complex and dynamic environments, resulting in the problem of location data being unavailable when navigating indoors.
Analyzing the surroundings of the autologous using AI-based occupancy networks and surface networks, predicting the occupancy properties of voxels through image data captured by the camera, generating 3D models, positioning the autologous body and generating paths without the need for position tracking sensors.
The autologous body can understand its environment without relying on location data, navigating to the destination autonomously, improving navigation capabilities in complex environments.
Smart Images

Figure CN120380302A_ABST
Abstract
Description
[0001] Cross - Reference to Related Applications
[0002] This application claims the benefit of priority of U.S. Provisional Application No. 63 / 377,993, filed Sep. 30, 2022; U.S. Provisional Application No. 63 / 377,996, filed Sep. 30, 2022; U.S. Provisional Application No. 63 / 378,034, filed Sep. 30, 2022; and U.S. Provisional Application No. 63 / 377,919, filed Sep. 30, 2022, the entire contents of each of which are incorporated herein by reference. TECHNICAL FIELD
[0003] The present disclosure generally relates to artificial intelligence-based modeling techniques to analyze the surrounding environment of an agent and identify a path for the agent. BACKGROUND ART
[0004] Due to the rapid development of computer technology, autonomous navigation technologies used in autonomous vehicles and robots (collectively referred to as agents) have become ubiquitous. These advancements allow for safer and more reliable autonomous navigation of agents. Agents typically need to navigate in complex and dynamic environments and terrains, which may include vehicles, traffic, pedestrians, cyclists, and various other static or dynamic obstacles. Most agent path planning protocols use location data to determine a suitable path for the agent. However, for many agents designed for indoor navigation, location data may not be available. SUMMARY OF THE INVENTION
[0005] For the above reasons, there is a desire for methods and systems that can analyze the surrounding environment of an agent and predict the objects with mass present in the surrounding environment of the agent in order to identify a suitable path for the agent. It is necessary to determine the position of the agent without relying on location data so that a suitable path for the agent can be determined.
[0006] Using the methods and systems discussed herein allows for agent navigation without the need to use location data (such as GPS-enabled sensors) for localization. This is (at least in part) because the (multiple) AI models discussed herein can use images captured by the (multiple) cameras of the agent to predict the environment around the agent in real time, even if the image data has never been ingested to train the AI model. This concept is described herein as an occupancy network or (multiple) occupancy detection models. Additionally, the properties of various surfaces can be analyzed such that (when combined with occupancy network data) the agent can understand the environment in which it is navigating. This concept is referred to herein as surface detection or surface network.
[0007] The methods and systems discussed herein use various AI models (e.g., occupancy networks and surface networks) to analyze the surrounding environment of an ego, and identify a suitable path towards its destination for the ego. Using the methods and systems discussed herein, the ego can periodically (e.g., throughout its path) localize itself using only image data. Thus, the ego can understand its surrounding environment, localize itself, and navigate autonomously.
[0008] In one embodiment, a method includes: retrieving, by a processor, image data of a space surrounding an ego, the image data captured by a camera of the ego; predicting, by the processor, occupancy attributes of a plurality of voxels corresponding to the space surrounding the ego by executing an artificial intelligence model; generating, by the processor, a 3D model corresponding to the space surrounding the ego and the occupancy attributes of each voxel; after receiving a destination, localizing the ego by using key image features within the image data corresponding to the 3D model to identify the current position of the ego, without receiving the position of the ego from a position tracking sensor; and generating, by the processor, a path for the ego to proceed from the current position to the destination.
[0009] The method may further include localizing the ego periodically during the path by the processor.
[0010] Localizing the ego may include tracking key image features in consecutive image data.
[0011] Key image points correspond to unique points within the image data.
[0012] Generating the path may include generating at least one of a trajectory, yaw rate, forward speed, or lateral speed for the ego.
[0013] The path may be generated using an Iterative Linear Quadratic Regulator (ILQR) protocol.
[0014] The 3D model may further correspond to surface attributes of at least one object within the space surrounding the ego.
[0015] In another embodiment, a computer system includes a computer-readable medium having a set of instructions that, when executed, cause a processor to: retrieve image data of a space surrounding an ego, the image data captured by a camera of the ego; predict occupancy attributes of a plurality of voxels corresponding to the space surrounding the ego by executing an artificial intelligence model; generate a 3D model corresponding to the space surrounding the ego and the occupancy attributes of each voxel; after receiving a destination, localize the ego by using key image features within the image data corresponding to the 3D model to identify the current position of the ego, without receiving the position of the ego from a position tracking sensor; and generate a path for the ego to proceed from the current position to the destination.
[0016] The set of instructions also causes the processor to periodically localize itself during the path.
[0017] Localizing itself can include tracking key image features in successive image data.
[0018] The key image points can correspond to unique points within the image data.
[0019] Generating the path can include generating at least one of a trajectory, a yaw rate, a forward speed, or a lateral speed for the self.
[0020] The path can be generated using an Iterative Linear Quadratic Regulator (ILQR) protocol.
[0021] The 3D model can also correspond to surface properties of at least one object within the space surrounding the self.
[0022] In another embodiment, a self includes a processor configured to: retrieve image data of the space surrounding the self, captured by a camera of the self; predict occupancy properties of a plurality of voxels corresponding to the space surrounding the self by executing an artificial intelligence model; generate a 3D model corresponding to the space surrounding the self and the occupancy properties of each voxel; localize the self upon receiving a destination by identifying a current position of the self using key image features within the image data corresponding to the 3D model, without receiving the position of the self from a position tracking sensor; and generate a path for the self to proceed from the current position to the destination.
[0023] The processor can also be configured to periodically localize the self during the path.
[0024] Localizing the self can include tracking key image features in successive image data.
[0025] The key image points can correspond to unique points within the image data.
[0026] Generating the path can include generating at least one of a trajectory, a yaw rate, a forward speed, or a lateral speed for the self.
[0027] The path can be generated using an Iterative Linear Quadratic Regulator (ILQR) protocol. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] Non-limiting embodiments of the present disclosure are described by way of examples in relation to the accompanying drawings, which are schematic and not intended to be drawn to scale. Unless indicated as representing the background art, the drawings represent aspects of the present disclosure.
[0029] Figure 1A Components of an AI-enabled visual data analysis system according to an embodiment are shown.
[0030] Figure 1B Shows various sensors associated with the ego according to an embodiment.
[0031] Figure 1C Shows components of a vehicle according to an embodiment.
[0032] Figures 2A to 2B Shows a flowchart of different processes performed in an AI-enabled visual data analysis system according to an embodiment.
[0033] Figures 3A to 3B Shows different occupancy maps generated in an AI-enabled visual data analysis system according to an embodiment.
[0034] Figures 4A to 4C Shows different views of a surface map generated in an AI-supported visual data analysis system according to an embodiment.
[0035] Figure 5 Shows a flowchart of a process for executing an AI model to generate a surface map according to an embodiment.
[0036] Figure 6A Shows a flowchart of a process for executing an AI model to generate an ego path according to an embodiment.
[0037] Figure 6B Shows a diagram of a process for tuning an AI model according to an embodiment.
[0038] Figures 7A to 7B Shows a three-dimensional (3D) model representing the environment / space around the ego according to an embodiment.
[0039] Figures 8 to 10 Shows different 3D models according to different embodiments.
[0040] Figure 11 Shows the path taken by the ego according to an embodiment. DETAILED DESCRIPTION
[0041] Reference will now be made to the illustrative embodiments depicted in the accompanying drawings, and specific language will be used to describe them. However, it is understood that no limitation of the claims or the scope of the disclosure is thereby intended. Changes and further modifications of the features of the invention shown herein, as well as additional applications of the principles of the subject matter shown herein, which would occur to one of ordinary skill in the relevant art to which the disclosure pertains, are to be considered within the scope of the subject matter disclosed herein. Other embodiments may be used and / or other changes may be made without departing from the spirit or scope of the disclosure. The illustrative embodiments described in the detailed description do not represent a limitation of the subject matter presented.
[0042] By implementing the methods described herein, a system can use a trained AI model to determine the occupancy status of different voxels of an image (or video) of the surrounding environment of an ego. The ego can be an autonomous vehicle (e.g., a car, truck, bus, motorcycle, all-terrain vehicle, cart), a robot, or other automated device. The ego can be configured to operate within a production line, building, home, or medical center, or to transport people, deliver goods, perform military functions, etc. In these environments, the ego can navigate between known or unknown paths to complete a specific task or travel to a specific destination. During operation, it is desired to avoid collisions, so the ego seeks to understand the environment. For example, in the context of an autonomous vehicle or a robot, the system can use a camera (or other visual sensor) to receive real-time or near-real-time images of the surrounding environment of the ego. Then, the system can execute the trained AI model to determine the occupancy status of the surrounding environment of the ego. The AI model can divide the surrounding environment of the ego into different voxels and then determine the occupancy status of each voxel. Thus, using the methods discussed herein, the system can generate a map of the surrounding environment of the ego. Using voxel data (e.g., the coordinates of each voxel) and the corresponding occupancy status, the AI model (or sometimes another model using data predicted by the AI model) can generate a map of the surrounding environment of the ego.
[0043] Figure 1A is a non-limiting example of a component of a system in which the methods and systems discussed herein can be implemented. For example, an analysis server can train an AI model and use the trained AI model to generate an occupancy dataset and / or a map for one or more egos. Figure 1A Illustrates the components of an AI-enabled visual data analysis system 100. The system 100 can include an analysis server 110a, a system database 110b, an administrator computing device 120, egos 140a to 140b (collectively referred to as (multiple) egos 140), ego computing devices 141a to 141c (collectively referred to as ego computing devices 141), and a server 160. The system 100 is not limited to the components described herein and can include additional or other components not shown for the sake of brevity, which will be considered within the scope of the embodiments described herein.
[0044] The components mentioned above can be connected via a network 130. Examples of the network 130 can include, but are not limited to, a private or public LAN, WLAN, MAN, WAN, and the Internet. The network 130 can include wired and / or wireless communication according to one or more standards and / or via one or more transmission media.
[0045] Communication on network 130 can be performed according to various communication protocols, such as Transmission Control Protocol and Internet Protocol (TCP / IP), User Datagram Protocol (UDP), and IEEE communication protocols. In one example, network 130 can include wireless communication according to a Bluetooth specification set or another standard or proprietary wireless communication protocol. In another example, network 130 can also include communication over a cellular network, including, for example, GSM (Global System for Mobile Communications), CDMA (Code Division Multiple Access), or EDGE (Enhanced Data rates for Global Evolution) network.
[0046] System 100 illustrates an example of a system architecture and components that can be used to train and execute one or more AI models, such as (multiple) AI models 110c. Specifically, as Figure 1A shown and described herein, the analysis server 110a can use the methods discussed herein to train (multiple) AI models 110c using data retrieved from the self-bodies 140 (e.g., by using data streams 172 and 176). When the (multiple) AI models 110c have been trained, each self-body 140 can access and execute the (multiple) trained AI models 110c. For example, a vehicle 140a having a self-body computing device 141a can transmit its camera feed to the (multiple) trained AI models 110c and can determine the occupancy status of its surrounding environment (e.g., data stream 174). Additionally, data ingested and / or predicted by the (multiple) AI models 110c with respect to the self-bodies 140 (at inference time) can also be used to improve the (multiple) AI models 110c. Thus, system 100 depicts a continuous loop that can periodically improve the accuracy of the (multiple) AI models 110c. Additionally, system 100 depicts a loop in which data received by the self-bodies 140 can also be used during the training phase in addition to the inference phase.
[0047] The analysis server 110a can be configured to collect, process, and analyze navigation data (e.g., images captured during navigation) and various sensor data collected from the self-bodies 140. The collected data can then be processed and prepared into a training dataset. The training dataset can then be used to train one or more AI models, such as AI model 110c. The analysis server 110a can also be configured to collect visual data from the self-bodies 140. Using the AI model 110c (trained using the methods and systems discussed herein), the analysis server 110a can generate a dataset and / or an occupancy map for the self-bodies 140. The analysis server 110a can display the occupancy map on the self-bodies 140 and / or transmit the occupancy map / dataset to the self-body computing device 141, the administrator computing device 120, and / or the server 160.
[0048] In Figure 1AIn [the figure], the AI model 110c is shown as a component of the system database 110b, but the AI model 110c can be stored in different or separate components, such as a cloud storage device or any other data repository accessible to the analysis server 110a.
[0049] The analysis server 110a can also be configured to display an electronic platform that shows various training attributes for training the AI model 110c. The electronic platform can be displayed on the administrator computing device 120 so that an analyst can monitor the training of the AI model 110c. Examples of electronic platforms generated and hosted by the analysis server 110a can be web-based applications or websites that are configured to display the training data set collected from the vehicle 140 and / or the training status / metrics of the AI model 110c.
[0050] The analysis server 110a can be any computing device that includes a processor and a non-transitory machine-readable storage device capable of performing the various tasks and processes described herein. Non-limiting examples of such computing devices can include workstation computers, laptop computers, server computers, etc. Although the system 100 includes a single analysis server 110a, the system 100 can include any number of computing devices operating in a distributed computing environment (such as a cloud environment).
[0051] The vehicle 140 can represent various electronic data sources that transmit data associated with its previous or current navigation session to the analysis server 110a. The vehicle 140 can be any device configured for navigation, such as a car 140a and / or a truck 140c. The vehicle 140 is not limited to vehicles and can also include robotic devices. For example, the vehicle 140 can include a robot 140b, which can represent a general-purpose, bipedal, autonomous humanoid robot capable of navigating various terrains. The robot 140b can be equipped with software capable of achieving balance, navigation, perception, or interaction with the physical world. The robot 140b can also include various cameras configured to transmit visual data to the analysis server 110a.
[0052] Although referred to herein as a "vehicle", the vehicle 140 may or may not be an autonomous device configured for autonomous navigation. For example, in some embodiments, the vehicle 140 can be controlled by a human operator or a remote processor. The vehicle 140 can include various sensors, such as Figure 1BThe sensors shown. The sensors can be configured to collect data while the vehicle 140 navigates various terrains (e.g., roads). The analysis server 110a can collect the data provided by the vehicle 140. For example, the analysis server 110a can obtain navigation sessions and / or road / terrain data (e.g., images of the vehicle 140 navigating on a road) from various sensors, such that the collected data can ultimately be used by the AI model 110c for training purposes.
[0053] As used herein, a navigation session corresponds to a journey of the vehicle 140's travel route, regardless of whether the travel is autonomous or human-controlled. In some embodiments, the navigation session can be used for data collection and model training purposes. However, in some other embodiments, the vehicle 140 can refer to a vehicle purchased by a consumer, and the purpose of the journey can be classified as daily use. The navigation session can start when the vehicle 140 moves more than a threshold distance (e.g., 0.1 mile, 100 feet) or more than a threshold rate (e.g., more than 0 mph, more than 1 mph, more than 5 mph) from a non-moving position. The navigation session can end when the vehicle 140 returns to the non-moving position and / or shuts down (e.g., when the driver exits the vehicle).
[0054] The vehicle 140 can represent a collection of vehicles monitored by the analysis server 110a to train the AI model(s) 110c. For example, the driver of the vehicle 140a can authorize the analysis server 110a to monitor data associated with their respective vehicle. As a result, the analysis server 110a can utilize the various methods discussed herein to collect sensor / camera data and generate a training dataset to train the AI model(s) 110c accordingly. The analysis server 110a can then apply the trained AI model(s) 110c to analyze data associated with the vehicle 140 and predict the occupancy map of the vehicle 140. Additionally, the additional / ongoing data associated with the vehicle 140 can also be processed and added to the training dataset so that the analysis server 110a can recalibrate the AI model(s) 110c accordingly. Thus, the system 100 depicts a cycle where the navigation data received from the vehicle 140 can be used to train the AI model(s) 110c. The vehicle 140 can include a processor that executes the trained AI model(s) 110c for navigation purposes. During navigation, the vehicle 140 can collect additional data about its navigation session, and this additional data can be used to calibrate the AI model(s) 110c. That is, the vehicle 140 represents a vehicle that can be used to train, execute / use, and recalibrate the AI model(s) 110c. In a non-limiting example, the vehicle 140 represents vehicles purchased by customers that can use the AI model(s) 110c for autonomous navigation while improving the AI model(s) 110c.
[0055] The self - body 140 can be equipped with various technologies to allow the self - body to collect data from the surrounding environment and (possibly) navigate autonomously. For example, the self - body 140 can be equipped with an inference chip to run autonomous driving software.
[0056] The various sensors of each self - body 140 can monitor the data collected related to different navigation sessions and transmit it to the analysis server 110a. Figures 1B to 1C A block diagram of the sensors integrated within the self - body 140 according to an embodiment is shown. The number and location of each sensor discussed with respect to Figures 1B to 1C can depend on the type of self - body discussed in Figure 1A For example, the robot 140b can include different sensors from the vehicle 140a or the truck 140c. For example, the robot 140b may not include the airbag activation sensor 170q. Additionally, the sensors of the vehicle 140a and the truck 140c can be at different locations from those Figure 1C shown.
[0057] As described herein, the various sensors integrated within each self - body 140 can be configured to measure various data associated with each navigation session. The analysis server 110a can periodically collect the data monitored and collected by these sensors, where the data is processed according to the methods described herein and used to train the AI model 110c and / or execute the AI model 110c to generate an occupancy map.
[0058] The self - body 140 can include a user interface 170a. The user interface 170a can refer to the user interface of the self - body computing device (e.g., Figure 1A the self - body computing device 141 in Figure 1B ). The user interface 170a can be implemented as a display screen, a head - up display, a touch screen, etc. integrated with or coupled to the interior of the vehicle. The user interface 170a can include input devices such as a touch screen, a knob, a button, a keyboard, a mouse, a gesture sensor, a steering wheel, etc. In various embodiments, the user interface 170a can be adapted to provide user input (e.g., as a kind of signal and / or sensor information) to other devices or sensors of the self - body 140 (e.g.,
[0059] The user interface 170a can also be implemented with one or more logic devices that can be adapted to execute instructions, such as software instructions, to implement any of the various processes and / or methods described herein. For example, the user interface 170a can be adapted to form a communication link, transmit and / or receive communications (e.g., sensor signals, control signals, sensor information, user input, and / or other information), or perform various other processes and / or methods. In another example, the driver can use the user interface 170a to control the temperature of the vehicle body 140 or activate its features (e.g., the autonomous driving or steering system 170o). Thus, the user interface 170a can be combined with other sensors described herein to monitor and collect driving session data. The user interface 170a can also be configured to display various data generated / predicted by the analysis server 110a and / or the AI model 110c.
[0060] The orientation sensor 170b can be implemented as a compass, a float, an accelerometer, and / or one or more of any other digital or analog devices capable of measuring the orientation of the vehicle body 140 (e.g., the magnitude and direction of roll, pitch, and / or yaw relative to one or more reference orientations such as gravity and / or magnetic north). The orientation sensor 170b can be adapted to provide a heading measurement for the vehicle body 140. In other embodiments, the orientation sensor 170b can be adapted to provide roll, pitch, and / or yaw rate for the vehicle body 140 using a time series of orientation measurements. The orientation sensor 170b can be positioned and / or adapted to make orientation measurements relative to a specific coordinate system of the vehicle body 140.
[0061] The controller 170c can be implemented as any suitable logic device (e.g., a processing device, a microcontroller, a processor, an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), a memory storage device, a memory reader, or other device or combination of devices) that can be adapted to execute, store, and / or receive appropriate instructions, such as software instructions that implement control loops for controlling various operations of the vehicle body 140. Such software instructions can also implement methods for processing sensor signals, determining sensor information, providing user feedback (e.g., via the user interface 170a), querying devices for operating parameters, selecting operating parameters for devices, or performing any of the various operations described herein.
[0062] The communication module 170e can be implemented as any wired and / or wireless interface configured to transfer sensor data, configuration data, parameters, and / or other data and / or signals to Figure 1A any of the features shown (e.g., the analysis server 110a). As described herein, in some embodiments, the communication module 170e can be implemented in a distributed manner such that portions of the communication module 170e are inFigure 1B implemented within one or more of the illustrated components and sensors. In some embodiments, communication module 170e may delay the transmission of sensor data. For example, when the vehicle 140 does not have a network connection, communication module 170e may store the sensor data in a temporary data storage device and transmit the sensor data when the vehicle 140 is identified as having an appropriate network connection.
[0063] The rate sensor 170d may be implemented as an electronic pitot tube, a metering gear or wheel, a water speed sensor, a wind speed sensor, a wind velocity sensor (e.g., direction and magnitude), and / or other devices capable of measuring or determining the linear rate of the vehicle 140 (e.g., in the surrounding medium and / or aligned with the longitudinal axis of the vehicle 140) and providing such a measurement as a sensor signal, which may be transmitted to various devices.
[0064] The gyroscope / accelerometer 170f may be implemented as one or more electronic sextants, semiconductor devices, integrated chips, accelerometer sensors, or other systems or devices capable of measuring the angular velocity / acceleration and / or linear acceleration of the vehicle 140 (e.g., direction and magnitude) and providing such a measurement as a sensor signal, which may be transmitted to various devices, such as the analytics server 110a. The gyroscope / accelerometer 170f may be positioned and / or adapted to make such measurements with respect to a particular coordinate system of the vehicle 140. In various embodiments, the gyroscope / accelerometer 170f may be implemented in a common housing and / or module with Figure 1B the other illustrated components to ensure a common reference frame or known transformation between reference frames.
[0065] The Global Navigation Satellite System (GNSS) 170h may be implemented as a global positioning satellite receiver and / or other device capable of determining the absolute and / or relative position of the vehicle 140 based on wireless signals received, for example, from space and / or ground sources and capable of providing measurements such as sensor signals, which may be transmitted to various devices. In some embodiments, GNSS 170h may be adapted to determine the speed, rate, and / or yaw rate of the vehicle 140 (e.g., using a time series of position measurements), such as the yaw component of the absolute speed and / or angular velocity of the vehicle 140.
[0066] The temperature sensor 170i may be implemented as a thermistor, an electrical sensor, an electrical thermometer, and / or other devices capable of measuring the temperature associated with the vehicle 140 and providing such a measurement as a sensor signal. The temperature sensor 170i may be configured to measure the ambient temperature associated with the vehicle 140, such as the cockpit or dashboard temperature, which may be used to estimate the temperature of one or more components of the vehicle 140.
[0067] The humidity sensor 170j may be implemented as a relative humidity sensor, an electrical sensor, an electrical relative humidity sensor, and / or another device capable of measuring the relative humidity associated with the body 140 and providing such a measurement as a sensor signal.
[0068] The steering sensor 170g may be adapted to physically adjust the heading of the body 140 based on one or more control signals provided by a logic device (such as the controller 170c) and / or user input. The steering sensor 170g may include one or more actuators and control surfaces (e.g., a rudder or other type of steering or trimming mechanism) of the body 140, and may be adapted to physically adjust the control surface to various positive and / or negative steering angles / positions. The steering sensor 170g may also be adapted to sense the current steering angle / position of such a steering mechanism and provide such a measurement.
[0069] The propulsion system 170k may be implemented as a propeller, a turbine, or other thrust-based propulsion system, a mechanical wheeled and / or tracked propulsion system, a wind / sail-based propulsion system, and / or other types of propulsion systems that may be used to power the body 140. The propulsion system 170k may also monitor the prime power and / or thrust of the body 140 with respect to the coordinate system of the body 140. In some embodiments, the propulsion system 170k may be coupled and / or integrated with the steering sensor 170g.
[0070] The passenger restraint sensor 170l may monitor seat belt detection and locking / unlocking components and other passenger restraint subsystems. The passenger restraint sensor 170l may include various environmental and / or status sensors, actuators, and / or other devices that facilitate the operation of safety mechanisms associated with the operation of the body 140. For example, the passenger restraint sensor 170l may be configured to receive motion and / or status data from Figure 1B the other sensors shown. The passenger restraint sensor 170l may determine whether safety measures (e.g., seat belts) are being used.
[0071] As Figure 1C shown, the camera 170m may refer to one or more cameras integrated within the body 140, and may include multiple cameras integrated (or retrofitted) into the body 140. The camera 170m may be an internal or external camera of the body 140. For example, as Figure 1C shown, the body 140 may include one or more internal-facing cameras 170m-1. These cameras may monitor and collect footage of the occupants of the body 140. The body 140 may also include a front-facing side camera 170m-2, a camera 170m-3 (e.g., integrated within a door frame), and a rear-facing side camera 170m-4.
[0072] Refer toFigure 1B , the radar 170n and the ultrasonic sensor 170p can be configured to monitor the distance from the vehicle body 140 to other objects, such as other vehicles or immovable objects (e.g., trees or garage doors). As Figure 1C shown, the radar 170n and the ultrasonic sensor 170p can be integrated into the vehicle body 140. The vehicle body 140 can also include an autonomous driving or steering system 170o configured to autonomously navigate the vehicle body 140 using data collected via various sensors (e.g., the radar 170n, the rate sensor 170d, and / or the ultrasonic sensor 170p).
[0073] Accordingly, the autonomous driving or steering system 170o can analyze various data collected by one or more of the sensors described herein to identify driving data. For example, the autonomous driving or steering system 170o can calculate the risk of a forward collision based on the speed of the vehicle body 140 and its distance to another vehicle on the road. The autonomous driving or steering system 170o can also determine whether the driver is touching the steering wheel. The autonomous driving or steering system 170o can transmit the analyzed data to various features discussed herein, such as the analysis server.
[0074] The airbag activation sensor 170q can predict or detect a collision and cause the activation or deployment of one or more airbags. The airbag activation sensor 170q can transmit data regarding the airbag deployment, including data associated with the event that caused the deployment.
[0075] Referring again to Figure 1A , the administrator computing device 120 can represent a computing device operated by a system administrator. The administrator computing device 120 can be configured to display data retrieved or generated by the analysis server 110a (e.g., various analysis metrics and risk scores), where the system administrator can monitor the various models used by the analysis server 110a, view feedback, and / or facilitate the training of the AI model(s) 110c maintained by the analysis server 110a.
[0076] The vehicle body(ies) 140 can be any device configured to navigate various routes, such as the vehicle 140a or the robot 140b. As referenced in Figures 1B to 1CAs discussed, the vehicle 140 may include various telemetry sensors. The vehicle 140 may also include a vehicle computing device 141. Specifically, each vehicle may have its own vehicle computing device 141. For example, truck 140c may have a vehicle computing device 141c. For simplicity, the vehicle computing devices are collectively referred to as the vehicle computing device(s) 141. The vehicle computing device 141 may control content presentation on the infotainment system of the vehicle 140, process commands associated with the infotainment system, aggregate sensor data, manage communication of data to electronic data sources, receive updates and / or transmit messages. In one configuration, the vehicle computing device 141 communicates with an electronic control unit. In another configuration, the vehicle computing device 141 is an electronic control unit. The vehicle computing device 141 may include a processor and a non-transitory machine-readable storage medium capable of performing the various tasks and processes described herein. For example, the AI model(s) 110c described herein may be stored and executed (or directly accessed) by the vehicle computing device 141. Non-limiting examples of the vehicle computing device 141 may include vehicle multimedia and / or display systems.
[0077] In one example of how to train the AI model(s) 110c, the analytics server 110a may collect data from the vehicle 140 to train the AI model(s) 110c. Before executing the AI model(s) 110c to generate / predict an occupancy dataset, the analytics server 110a may use various methods to train the AI model(s) 110c. Training allows the AI model(s) 110c to ingest data from one or more cameras of one or more vehicles 140 (without receiving radar data) and predict occupancy data of the surrounding environment of the vehicle. The operations described in this example may be performed by any number of computing devices (e.g., the processors of the vehicle 140) operating in the distributed computing system described in Figure 1A and Figure 1B .
[0078] The analytics server 110a may use the sensors of the vehicle 140 to generate a first dataset having a first set of data points, where each data point within the first set of data points corresponds to the position and sensor attributes of at least one voxel of the space around the vehicle 140, and the sensor attributes indicate whether at least one voxel is occupied by an object having mass.
[0079] To train the AI model(s) 110c, the analysis server 110a can first use one or more vehicles 140 to drive a specific route. While driving, the vehicle 140 can use one or more of its sensors (including one or more cameras) to generate navigation session data. For example, one or more vehicles 140 equipped with various sensors can navigate a designated route. As one or more vehicles 140 traverse the terrain, their sensors can capture continuous (or periodic) data of their surrounding environment. The sensors can indicate the occupancy status of the surrounding environment of one or more vehicles 140. For example, the sensor data can indicate various objects with mass around one or more vehicles 140 as they navigate their route.
[0080] The analysis server 110a can use the sensor data received from one or more vehicles 140 to generate a first data set. The first data set can indicate the occupancy status of different voxels of the surrounding environment of one or more vehicles 140. As used herein in some embodiments, a voxel is a three-dimensional pixel that forms the building blocks of the surrounding environment of one or more vehicles 140. In the first data set, each voxel can encapsulate sensor data that indicates whether mass is identified for that particular voxel. As used herein, mass can indicate or represent any object identified using the sensors. For example, in some embodiments, the vehicle 140 can be equipped with LiDAR, which identifies mass by emitting laser pulses and measuring the time required for these pulses to reach an object (with mass) and return. The LiDAR sensor system can operate based on the principle of measuring the distance between the LiDAR sensor and an object in its field of view. This information combined with other sensor data can be analyzed to identify and characterize different masses or objects around one or more vehicles 140.
[0081] Various additional data can be used to indicate whether a voxel in the surrounding environment of one or more vehicles 140 is occupied by an object with mass. For example, in some embodiments, a digital map of the surrounding environment of one or more vehicles 140 (e.g., a digital map of the route the vehicle is traversing) can be used to determine the occupancy status of each voxel.
[0082] In operation, as one or more vehicles 140 navigate, their sensors collect data and transmit the data to the analysis server 110a, as shown by the data stream 176. For example, the vehicle 140 computing device 141 can use the data stream 176 to transmit the sensor data to the analysis server 110a.
[0083] The analysis server 110a can use the cameras of the vehicle 140 to generate a second data set with a second set of data points, where each data point within the second set of data points corresponds to the position and image attributes of at least one voxel of the space around the vehicle 140.
[0084] The analysis server 110a can receive camera feeds of one or more self-bodies 140 navigating the same route as in the first step. In some embodiments, the analysis server 110a can perform the first step and the second step simultaneously (or concurrently). Alternatively, two (or more) different self-bodies 140 can navigate the same route, where one self-body transmits its sensor data and the second self-body 140 transmits its camera feed.
[0085] In some embodiments, one or more self-bodies 140 can include one or more high-resolution cameras that capture a continuous visual data stream from the surrounding environment of the one or more self-bodies 140 as the one or more self-bodies 140 navigate the route. The analysis server 110a can then use the camera feed to generate a second data set that includes visual elements / depictions of different voxels of the surrounding environment of the one or more self-bodies 140.
[0086] In operation, as the one or more self-bodies 140 navigate, their cameras collect data and transmit the data to the analysis server 110a, as shown by the data stream 172. For example, the self-body computing device 141 can use the data stream 172 to transmit image data to the analysis server 110a.
[0087] The analysis server 110a can use the first data set and the second data set to train an AI model, where the AI model 110c trains itself using the corresponding positions of each data point, associating each data point in the first data point set with the corresponding data point in the second data point set. After being trained, the AI model 110c is configured to receive a camera feed from a new self-body 140 and predict the occupancy state of at least one voxel of the camera feed.
[0088] Using the first data set and the second data set, the analysis server 110a can train the AI model(s) 110c such that the AI model(s) 110c can associate different visual attributes of a voxel (within the camera feed in the second data set) with the occupancy state of that voxel (within the first data set). In this way, after being trained, the AI model(s) 110c can receive a camera feed (e.g., from a new self-body 140) without receiving sensor data and then determine the occupancy state of each voxel of the new self-body 140.
[0089] The analysis server 110a can generate a training dataset including a first dataset and a second dataset. The analysis server 110a can use the first dataset as the ground truth. For example, the first dataset can indicate the different positions of the voxels and their occupancy status. The second dataset can include visual (e.g., camera feed) shows of the same voxels. Using the first dataset, the analysis server 110a can label the data such that the (multiple) data records associated with each voxel corresponding to an object are indicated as having a positive occupancy status.
[0090] The labeling of the occupancy status of different voxels can be performed automatically and / or manually. For example, in some embodiments, the analysis server 110a can use human reviewers to label the data. For example, as described herein, the camera feed from one or more cameras of a vehicle can be displayed to a human reviewer on an electronic platform for labeling. Additionally or alternatively, the (multiple) AI models 110c can ingest the entire data, where the (multiple) AI models 110c identify the corresponding voxels, analyze the first digital map, and associate the (multiple) images of each voxel with its corresponding occupancy status.
[0091] Using the ground truth, the (multiple) AI models 110c can be trained to analyze the visual elements of each voxel and associate it with whether the voxel is occupied by a mass. Thus, the AI model 110c can retrieve the occupancy status of each voxel (using the first dataset) and use this information as the ground truth. The (multiple) AI models 110c can also retrieve the visual attributes of the same voxels using the second dataset.
[0092] In some embodiments, the analysis server 110a can use a supervised training method. For example, using the ground truth and the received visual data, the (multiple) AI models 110c can train themselves such that it can predict the occupancy status of a voxel using only the image of the voxel. As a result, when trained, the (multiple) AI models 110c can receive the camera feed, analyze the camera feed, and determine the occupancy status of each voxel within the camera feed (without using radar).
[0093] The analysis server 110a can feed a series of training datasets into the (multiple) AI models 110c and obtain a set of predicted outputs (e.g., predicted occupancy status). The analysis server 110a can then compare the predicted data with the ground truth data to determine the differences and train the (multiple) AI models 110c by adjusting the internal weights and parameters of the AI models 110c proportional to the determined differences according to a loss function. The analysis server 110a can train the (multiple) AI models 110c in a similar manner until the trained AI models 110c predict accurately to a certain threshold (e.g., recall or precision).
[0094] Additionally or alternatively, the analytics server 110a may use an unsupervised method in which the training dataset is not labeled. Since labeling the data in the training dataset can be time-consuming and may require excessive computing power, the analytics server 110a may utilize unsupervised training techniques to train the AI model 110c.
[0095] After the AI model 110c is trained, the vehicle 140 may use it to predict occupancy data of the surroundings of one or more vehicles 140. For example, the AI model(s) 110c may divide the surroundings of the vehicle into different voxels and predict the occupancy state of each voxel. In some embodiments, the AI model(s) 110c (or the analytics server 110a that uses the data predicted using the AI model 110c) may generate an occupancy map or occupancy network representing the surroundings of one or more vehicles 140 at any given time.
[0096] In another example of how the AI model(s) 110c may be used, after the AI model(s) 110c are trained, the analytics server 110a (or a local chip of the vehicle 140) may collect data from the vehicle (e.g., one or more vehicles 140) to predict an occupancy dataset of one or more vehicles 140. This example describes how the AI model(s) 110c may be used to predict occupancy data of one or more vehicles 140 in real-time or near real-time. This configuration may have a processor that executes the AI model, such as the analytics server 110a. However, one or more actions may be performed locally, e.g., via a chip located within one or more vehicles 140. In operation, the AI model(s) 110c may be executed locally via the vehicle 140 such that the results may be used for autonomous navigation.
[0097] The processor may input image data of the space around the vehicle object 140 into the AI model 110c using a camera of the vehicle object 140. The processor may collect and / or analyze data received from various cameras of one or more vehicles 140 (e.g., outward-facing cameras). In another example, the processor may collect and aggregate footage recorded by one or more cameras of the vehicle 140. The processor may then transmit the footage to the AI model(s) 110c trained using the methods discussed herein.
[0098] The processor may predict the occupancy attributes of multiple voxels by executing the AI model 110c. The AI model(s) 110c may use the methods discussed herein to predict the occupancy state of different voxels surrounding one or more vehicles 140 using the received image data.
[0099] The processor can generate a data set based on multiple voxels and their corresponding occupancy attributes. The analysis server 110a can generate a data set including its occupancy status according to the corresponding coordinate values of different voxels. This data set can be a queryable data set available for transmitting the predicted occupancy status to different software modules.
[0100] In operation, one or more aut bodies 140 can collect image data from their cameras and transmit the image data to the processor (locally placed on one or more aut bodies 140) and / or the analysis server 110a, as shown by the data stream 172. The processor can then execute the (multiple) AI models 110c to predict the occupancy data of one or more aut bodies 140. If the prediction is executed by the analysis server 110a, the occupancy data can be transmitted to one or more aut bodies 140 using the data stream 174. If the processor is locally placed within one or more aut bodies 140, the occupancy data is transmitted to the aut body computing device 141 ( Figure 1A not shown in).
[0101] Using the methods discussed herein, the training of the (multiple) AI models 110c can be performed such that the execution of the (multiple) AI models 110c can be locally executed on any aut body 140 (at inference time). The collected data (e.g., navigation data collected during the navigation of the aut body 140, such as image data of the journey) can then be fed back into the (multiple) AI models 110c such that additional data can improve the (multiple) AI models 110c.
[0102] FIG. 2 shows a flowchart of a method 200 executed in an AI-enabled visual data analysis system according to an embodiment. The method 200 can include step 210 to step 270. However, other embodiments can include additional or alternative steps, or one or more steps can be omitted. The method 200 is executed by an analysis server (e.g., a computer similar to the analysis server 110a). However, one or more steps of the method 200 can be executed by any number of computing devices (e.g., the processors of the aut body 140 and / or the aut body computing device 141) operating in the Figures 1A to 1C distributed computing system described in. For example, one or more computing devices of the aut body can locally execute some or all of the steps described in FIG. 2.
[0103] FIG. 2 shows a model architecture of how to ingest image input from an aut body (step 210) and analyze it to predict a queryable output (step 270). Using the methods and systems discussed herein, the analysis server can ingest only image data (e.g., a camera feed from the surrounding environment of the aut body) to generate a queryable output. Thus, the methods and systems discussed herein can operate without receiving any data from radar, LiDAR, etc.
[0104] The queryable output (generated in step 270) can be used for various purposes. In one example, the queryable output can be used in an autonomous driving module, where various navigation decisions can be made based on whether the voxels in the space around the ego are predicted to be occupied. In another example, using the queryable output, an analysis server can generate a digital map showing the occupancy status of the surrounding environment of the ego. For example, the analysis server can generate a three-dimensional (3D) geometric representation of the surrounding environment of the ego. For example, the digital map can be displayed on the ego's computing device.
[0105] As used herein, a voxel can refer to a volume pixel or the 3D equivalent of a pixel in 2D. Thus, a voxel can represent a defined point in a 3D grid within the volumetric space or environment around the ego (e.g., surrounding). In some embodiments, the space around the ego can be divided into different voxels, referred to as a voxel grid. As used herein, a voxel grid can refer to a set of cubes stacked (or arranged) together to represent objects in the space around the ego. Each voxel can contain information about a specific location within the surrounding space of the ego. Using the methods and systems discussed herein, the occupancy of each voxel can be evaluated. For example, an analysis server (using the AI models discussed herein) can determine whether each voxel is occupied by an object with mass. Voxel predictions can be aggregated into a dataset referred to herein as a queryable result. Using the queryable result, voxel information can be queried by a processor or downstream software module (e.g., autonomous driving software / processor) to identify the occupancy data of the surrounding environment of the ego.
[0106] In some embodiments, if any part of a voxel is occupied, the voxel can be designated as occupied. Thus, in some embodiments, each voxel can include a binary designation of 0 (unoccupied) or 1 (occupied). Alternatively, in some embodiments, the AI model can also predict detailed occupancy data within / inside a specific voxel. For example, a voxel with a binary value of 1 (occupied) can be further analyzed at a finer granularity level to determine the occupancy of each point within the voxel. For example, an object can be curved. While some voxels (associated with the object) are fully occupied, some other voxels can be partially occupied. These voxels can be divided into smaller voxels such that some of the smaller voxels are unoccupied. As described herein, this method can be used to identify the shape of an object.
[0107] Method 200 begins at step 210, where image data is received from one or more cameras of the self. Method 200 visually shows how an AI model (trained using the methods discussed herein) ingests image data and generates queryable outputs that can indicate the volume occupancy of various voxels within the surrounding environment of the self. Image data can refer to any data received from one or more images of the self.
[0108] The captured image data can then be characterized (step 220). An image characterizer or various characterization algorithms can be used to extract relevant and meaningful features from the received image data. Using an image characterizer, the image data can be transformed into a data representation that captures important information about the image content. This allows for more efficient analysis of the image data.
[0109] In some embodiments, the AI model can perform the characterization discussed herein. In some other embodiments, a convolutional neural network can be used to characterize the image data. In one non-limiting example, as shown, a RegNet (Regularized Neural Network) can be used to transform the data into a BiFPN (Bidirectional Feature Pyramid Network). However, other protocols can also be used. In some other embodiments, a transformer can be used to characterize the image data.
[0110] After the image data is encoded / characterized, a transformer can be used to change the image data from a 2D image to a 3D image (step 230). As described herein, in an example configuration, there can be eight different cameras communicating with the self. As a result, the image data can include eight different camera feeds (one feed corresponding to each camera or other sensor), and can include overlapping views. The transformer can aggregate these individual camera feeds and generate one or more 3D representations using the received camera feeds.
[0111] The transducer can ingest three separate inputs: an image key, an image value, and a 3D query. The image key and image value can refer to attributes associated with 2D image data received from the ego. For example, these values can be output via image characterization (step 220). The transducer can also use an image query from the 3D space. The depicted spatial attention module can use the 3D query to analyze the 2D image key and image value. As shown, the BiFPN generated in step 220 can be aggregated into a multi-camera query embedding and can be used to perform 3D space queries. In some embodiments, each voxel can have its own query. Using the 3D space query, the analysis server can identify regions within the 2D characterized image that correspond to specific parts of the 3D representation. The identified regions within the characterized image can then be analyzed to transform the multi-camera image data into a 3D representation for each voxel, which can result in a 3D representation of the ego's surrounding environment. Thus, the depicted spatial attention module can output a single 3D vector space representing the ego's surrounding environment. In effect, this moves all of the image data generated by all of the camera feeds into a top-down space or a 3D spatial representation of the ego's surrounding environment.
[0112] Steps 210 through 230 can be performed on each video frame received from each camera of the ego. For example, at each timestamp, steps 210 through 230 can be performed on eight different images received from eight different cameras of the ego. As a result, at each timestamp, method 200 can produce a 3D spatial representation of the eight images. In step 240, method 200 can fuse together the 3D spaces (for different timestamps). This fusion can be based on the timestamps of each set of images. For example, the 3D spatial representations can be fused based on their corresponding timestamps (e.g., in a sequential manner).
[0113] As shown, the 3D spatial representation at timestamp t can be fused with the 3D spatial representations of the ego's surrounding environment at t-1, t-2, and t-3. As a result, the output can have both spatial and temporal information. This concept is depicted in Figure 2 as spatio-temporal features.
[0114] Then, the spatio-temporal features can be transformed into different voxels using deconvolution (step 250). As described herein, various data points are characterized and fused together. In this step 250, method 200 can perform various mathematical operations to reverse the process such that the fused data can be transformed back into different voxels. As used herein, deconvolution can refer to a mathematical operation used to reverse the effects of convolution.
[0115] After applying deconvolution to the image data (which has been characterized, transformed, and fused), method 200 can then apply various trained AI modeling techniques discussed herein (e.g., FIGS. 3 to 4) to generate a volumetric output (step 260). The volumetric output can include binary data of different voxels, which indicates whether a particular voxel is occupied by an object with mass. Specifically, the volumetric output can include occupancy data (including binary data) indicating whether a voxel is occupied and / or occupancy flow data indicating the rate at which the voxel is moving (if any) (the velocity is calculated using time alignment).
[0116] The volumetric output can also include shape information (the shape of the mass of the occupied voxels). In some embodiments, the size of each voxel can be predetermined, but the size can be modified to produce a more fine-grained result. For example, the default size of different voxels can be 33 centimeters (per vertex). While this size is generally acceptable for voxels, the results can be improved by reducing the size of the voxels. For example, a 33 cm voxel can be appropriate if the voxel is detected outside the driving surface of the ego. However, the analysis server can reduce the size of the occupied voxels (e.g., to 10 cm) to within a threshold distance from the ego and / or the ego-driving surface. When the voxel occupancy data is identified, a regression model can be executed to identify the shape of the group of voxels. For example, a 33 cm voxel (which belongs to a curb) can be half occupied (e.g., only 16 cm of the voxel is occupied). The analysis server can use regression to determine how much of the voxel is occupied.
[0117] Additionally or alternatively, the analysis server can decode the sub-voxel values to identify the shape of the sub-voxels (inside the occupied voxels). For example, if a voxel is half occupied, the analysis server can define a set of sub-voxels and use the methods discussed herein to identify the volumetric output of the sub-voxels. When the sub-voxels are aggregated (back into the original voxel), the analysis server can determine the shape of the voxel. For example, each voxel can have eight vertices. In some embodiments, each vertex can be analyzed individually and have its embedding. Thus, any point within each vertex of the voxel can be queried individually. Thus, in this "continuous resolution" method, the analysis server can not define the size of the sub-voxels. In some embodiments, the analysis server can use a multivariable interpolation (e.g., trilinear interpolation) protocol to estimate the occupancy status of each sub-voxel and / or any point within each vertex.
[0118] The volumetric output may also include 3D semantic data indicative of an object occupying a voxel (or a group of voxels). The 3D semantics may indicate whether the voxel and / or a group of nearby voxels are occupied by a car, a curb, a building, or other objects. The 3D semantics may also indicate whether the voxel is occupied by a static mass or a moving mass. The 3D semantic data may be identified using various temporal attributes of the voxels. For example, if a group of voxels is identified as being occupied by a mass, the collective shape of the voxels may indicate that the voxels belong to a vehicle. If, at a previous timestamp, the identified group of voxels (now known to be a vehicle) was identified as being in motion, the group of voxels may have 3D semantics indicative of the group of voxels belonging to a moving vehicle. In another example, if a group of voxels is identified as having a shape corresponding to a curb and is not identified as having any motion, the group of voxels may have 3D semantics indicative of a static curb.
[0119] In some embodiments, certain shapes or 3D semantics may be prioritized. For example, certain objects such as other vehicles on the road or objects associated with the driving surface (e.g., a curb indicating the outer boundary of a road) may be analyzed thoroughly. In contrast, details of static objects such as buildings far from the self-driving surface may not be analyzed as thoroughly as the analysis of moving vehicles near the self. In some embodiments, certain objects having a particular size or shape may be ignored. For example, the analysis of road debris may not be as extensive as the analysis of moving vehicles near the self.
[0120] In some embodiments, method 200 may not need to perform object-level detection. For example, the self must navigate around voxels identified as static and occupied in front of the self, regardless of whether the voxel belongs to another vehicle, a pedestrian, or a traffic sign. Thus, the occupancy information may be object-agnostic. In some embodiments, an object detection model may be executed separately (e.g., in parallel), which may detect objects corresponding to groups of voxels.
[0121] In step 270, method 200 may generate a queryable data set that allows other software modules to query the occupancy status of different voxels. For example, a software module may transmit coordinate values (X, Y, and Z axes) of the self's surrounding environment and may receive any one of four types of occupancy data (e.g., volumetric output) generated using method 200. The queryable data set may be used to generate an occupancy map (e.g., Figures 3A to 3B ) or may be used to make autonomous navigation decisions for the self.
[0122] Additionally or alternatively, an analysis server may generate a map corresponding to the predicted occupancy status of different voxels. In a non-limiting example, the analysis server may use a multi-view 3D reconstruction protocol to visualize each voxel and its occupancy status. Figures 3A to 3BNon-limiting examples of maps or occupancy maps are presented (e.g., simulation 350). In some embodiments, simulation 350 can be displayed on the user interface of the vehicle itself. Simulation 350 can show Figure 3A the camera feed 300 depicted in Figure 3A . The camera feed 300 represents image data (either in real-time or near real-time) received from eight different cameras of the vehicle itself. Specifically, the camera feed 300 can include camera feeds 310a to 310c received from three different front cameras of the vehicle itself; camera feeds 320a to 320b received from two different right-facing cameras of the vehicle itself; camera feeds 330a to 330b received from two different left-facing cameras of the vehicle itself; and a camera feed 340 received from the rear camera of the vehicle itself.
[0123] Using the methods discussed herein, the analysis server can analyze the camera feed 300, divide the space around the vehicle itself into voxels, and generate the simulation 350 as a graphical representation of the surrounding environment of the vehicle itself (as Figure 3B shown). The simulation 350 can include a simulated vehicle itself (360) and its surrounding voxels. For example, the simulation 350 can include graphical indicators of different qualities of different voxels that occupy the area around the simulated vehicle 360. For example, the simulation 350 can include simulation qualities 370a to 370c.
[0124] Each simulation quality block 370a to 370c can represent an object depicted in the camera feed 300. For example, simulation quality 370a corresponds to quality 380a (a vehicle); simulation quality 370b corresponds to quality 380b (a vehicle); and simulation quality 370c can correspond to quality 380c (a building near the road). As shown, each simulation quality includes various voxels. In addition, the voxels depicted in the simulation 350 can have different graphical / visual characteristics corresponding to their volume output (e.g., occupancy data). For example, simulation quality 370c (e.g., a building) can have a first color indicating that it has been identified as static. Similarly, simulation quality 370b (e.g., a vehicle) can have a second color indicating that it is a parked or stationary vehicle. In contrast, simulation quality 370a (e.g., another vehicle) can have a third color and / or other visual characteristics indicating that it is predicted to be moving.
[0125] Additionally or alternatively, the analysis server can transmit the generated map to a downstream software application or another server. The prediction results can be further analyzed and used in various models and / or algorithms to perform various actions. For example, a software model or processor associated with the autonomous navigation system of the vehicle itself can receive the occupancy data predicted by the trained AI model, and navigation decisions can be made based on this data.
[0126] Figure 2BFIG. 201 is a flowchart of a method 201 performed in an AI-enabled visual data analysis system according to an embodiment. Method 201 may include steps 210 to 290. However, other embodiments may include additional or alternative steps, or one or more steps may be entirely omitted. Method 201 is described as being performed by an analysis server (e.g., a computer similar to analysis server 110a). However, one or more steps of method 201 may be performed by any number of computing devices (e.g., processors of the ego 140 and / or the ego computing device 141) operating in the Figures 1A to 1C distributed computing system described in. For example, one or more computing devices of the ego may perform some or all of the steps described in Figure 2B .
[0127] Using method 201, an AI model may be configured to generate more than one orthogonal projection of the surroundings of the ego. The AI model may only need image data to predict various surfaces and their corresponding surface properties near the ego. As shown, method 201 includes a volume output (step 260) that indicates the surface properties of different volumes around the ego.
[0128] As shown, steps 210 to 250 may be similar in Figure 2A and Figure 2B . However, method 201 may include additional steps that allow the AI model to predict the properties of the surfaces around the ego. Specifically, method 201 may include an additional step 280 and an additional step 290, in which a ground truth is generated in additional step 280, and a 3D representation (e.g., model rendering) of the surroundings of the ego is generated using the data predicted via the execution of methods 200 and 201 in additional step 290.
[0129] Method 201 allows the AI model to predict the 3D properties of various surfaces in the surroundings of the ego, rather than generating an orthogonal view of the surroundings of the ego. Using method 201, it may no longer be necessary to localize the ego to achieve autonomous navigation. Compared with traditional methods, method 201 may allow the AI model to receive image data in real time or near real time (dynamically) and analyze various surfaces near the ego. Therefore, the ego may be able to navigate itself without performing a localization protocol.
[0130] Images received from the vehicle's camera can include a 2D representation of the vehicle's surrounding environment. Such a representation is sometimes referred to as a 2D or planar lattice. The planar lattice can be transformed into different nodes with specific X-axis and Y-axis coordinate values. Using method 201, the AI model can predict the Z-axis coordinate value of each node within the planar lattice. Specifically, using method 201, the AI model can predict the feature vector of each point with different X-axis and Y-axis coordinate values within the image data. As used herein, the Z-axis coordinate value of each point or node can represent the elevation of that point relative to a flat surface with an elevation of 0 in the world.
[0131] In addition to predicting the elevation of each node, the AI model can also determine the class (surface property) of each node. For example, the AI model can determine whether the surface is drivable. Additionally, the AI model can determine the properties of the material of each surface (e.g., grass, dirt, asphalt, or concrete). Additionally, the AI model can determine whether the surface is a road or a sidewalk. Additionally, the AI model can determine the paint lines associated with different surfaces so that the AI model can infer whether the surface is a road surface or a curb.
[0132] Using the feature vector of each node, the AI model can generate a mesh representation corresponding to the vehicle's surrounding environment. As used herein, a mesh can refer to a series of interconnected nodes representing the vehicle's surrounding environment, where each node includes X, Y, and Z-axis coordinate values. Each node can also include data indicating its properties and class (e.g., whether the node within the surface is drivable, what the node is identified as, and what material the node is predicted to be).
[0133] In step 280, the AI model can generate a ground truth to be ingested by the deconvolution step (250). The vehicle's sensors can generate a point cloud of the vehicle's surrounding environment. The point cloud can include many points that represent 3D coordinate data associated with the vehicle's surrounding environment at different timestamps. In a non-limiting example, LiDAR data can be received from the vehicle, and the point cloud can represent the received LiDAR data points. The vehicle's camera can also transmit images of the vehicle's surrounding environment at different timestamps. The analysis server can use the different timestamps to identify the image data corresponding to different points within the point cloud. The analysis server can then project the data associated with the points within the image data, thereby identifying the image regions (with a set of pixels) corresponding to one or more points within the point cloud.
[0134] The analysis server can also use an auxiliary AI model (e.g., a neural network), such as a semantic segmentation network, to analyze pixels within the image data. For example, a set of pixels can be analyzed by the semantic segmentation network. The semantic segmentation network can then determine one or more attributes of the set of pixels. For example, using this paradigm, the analysis server can determine whether a set of pixels corresponds to a tree, sky, curb, or road. In some embodiments, the semantic segmentation network can determine whether a surface is drivable. In some embodiments, the semantic segmentation network can determine the material associated with a set of pixels. For example, the semantic segmentation model can determine whether the pixels within the image data correspond to dirt, water, concrete, or asphalt. In some other embodiments, the semantic segmentation network can identify whether a surface is painted; and if so, identify whether the surface is the color of the paint. Essentially, the semantics of each 3D point can be identified using the semantic segmentation network.
[0135] Using the semantic segmentation model, the analysis server can filter these points and cluster them into their respective categories (e.g., pixels representing a sidewalk, pixels representing a dirt road or an asphalt road). The analysis server can analyze different image data at different timestamps.
[0136] After executing the semantic segmentation model, the point cloud can be segmented based on the corresponding image data of the point cloud and / or its attributes (as predicted by the semantic segmentation model). As a result, the points associated with a specific surface and the image data associated with the same surface around the ego can be identified and isolated. Then, the analysis server can fit a mesh surface to the isolated data points. This can be because the AI model can perform more efficiently using a smooth surface, which can better reflect reality. In fact, mesh fitting can denoise the data and provide a more realistic representation of the surface around the ego. The fitted surface can be used as ground truth for training purposes.
[0137] The AI model can be trained using the image data received from the ego and the ground truth, such that when being trained, the AI model can analyze the image data received from the ego without requiring any sensor data. Effectively, using this specific training paradigm, the AI model can associate how pixels associated with a specific surface with specific attributes (e.g., an uphill dirt road with white paint) are represented. Thus, the AI model (during inference) can utilize only the image data without requiring other sensor data.
[0138] After being trained, the AI model can be configured to ingest image data and generate a grid with various nodes, where each node has a corresponding feature vector, including X and Y axis coordinate values (identified via the image data) and Z axis coordinate values predicted by the AI model. The AI model can also predict one or more attributes of each node. For example, a particular node can include a feature vector that includes a predicted elevation (e.g., 1 meter above the self). Additionally, the AI model can predict that the node is a road node (since the corresponding pixel is predicted to be a driving surface) and that there is paint on the node and the paint is yellow.
[0139] In some embodiments, it may be necessary to adjust the coordinate values (e.g., the Z axis coordinate indicating the elevation of the node) because the self itself has changed position and the Z coordinate may not have been modified. For example, when the self is navigating in the terrain, it can transmit the coordinates of the surrounding environment. However, the coordinates can be related to the sensors of the self or the self itself. Thus, if the self changes its vertical position (e.g., if the self is driving over a speed bump or a pothole), the coordinates received from the self can also change. However, the coordinates can change because they are relative to the coordinates of the self. For example, the same position can have different coordinate values when the self is driving on a flat surface compared to when the self is driving over a speed bump. Therefore, in some embodiments, the coordinates received from the self can be modified before being used to train the AI model.
[0140] To correct this issue, the coordinate values can be aligned with the surface of the self's surrounding environment (rather than the self). In this way, the noisy or incorrect data received due to the movement of the self can be smoothed. Essentially, the surface is processed independently, and the coordinate values are calculated (and ultimately predicted) based on the surface rather than the self.
[0141] In some embodiments, method 201 can be combined with method 200 (the occupancy detection paradigm) to identify objects located within an elevated surface. For example, an object can be detected on a surface that has been identified as having a higher or lower elevation than the self (e.g., a traffic cone is identified on a mountain in front of the self). In this example, the AI model can use method 201 to determine the attributes of the mountain in front of the self. Then, the attributes of the cone itself can be identified as if the cone were on a flat surface (e.g., the height of the mountain at that particular location can be subtracted). Then, the AI model can use method 200 to identify the voxels associated with the cone, thereby identifying the dimensions of the cone. Then the dimensions are added to the mountain identified using method 201. Thus, the AI model can bifurcate the identification of the surface and the object and then combine them to truly understand / predict the position and attributes of different objects located on different surfaces.
[0142] Dividing the detection into two different protocols (methods 200 and 201) also allows the ego to detect the occupancy status of different voxels when the ego exceeds the occupancy detection range of the ego in different voxels. For example, the ego can have a vertical occupancy detection range from -3 meters to +3 meters. This means that if different voxels are within the elevation range of -3 meters to +3 meters of the ego, the ego can identify the occupancy status of the different voxels. The occupancy detection range may not mean the lens of the camera cannot record objects outside the range; in contrast, this may mean that the AI model cannot be used to identify objects outside the detection range.
[0143] In these embodiments, the ego may not be able to predict any objects on a steep slope outside the occupancy detection range of the ego (e.g., a traffic cone on a downhill slope at an elevation of -4 meters relative to the ego). Using the methods discussed herein, the ego can first determine that the driving surface is -4 meters lower than the ego. Then, the AI model can separately determine the attributes of the voxels of the occupied space (the traffic cone) and subtract the height of the mountain from the height of the traffic cone. Effectively, in addition to providing more consistent results, method 201 can also be used to extend the occupancy detection range of the ego (used in method 200).
[0144] Using method 201, the AI model can receive image data from the ego's camera and transform the image data into a mesh representation of the ego's surrounding environment. Thus, the image received from the camera can be transformed into a 3D description of the surfaces surrounding the ego, such as the driving surface.
[0145] In some embodiments, the analysis server can use neural radiance field (NeRF) technology to reconstruct a rendering of the ego's surrounding environment (step 290). In some embodiments, the analysis server can use the captured image data to generate a map indicating various surfaces around the ego. The map can correspond to the predicted surfaces and their predicted attributes. In a non-limiting example, the analysis server can use a multi-view 3D reconstruction protocol to visualize each voxel and its surface state / attributes. Figures 4A to 4C Non-limiting examples of the map or surface map are presented (e.g., simulation 400).
[0146] In some embodiments, simulation 400 may be displayed on a user interface of the vehicle. Simulation 400 may show how to analyze camera feed 410 to generate a graphical representation of the vehicle's surrounding environment. Camera feed 410 represents image data received from five different cameras of the vehicle (either in real-time or near real-time). Each camera feed may be received from a different camera and may depict a different perspective / angle of the vehicle's surrounding environment. Specifically, camera feed 410 represents image data received from eight different cameras of the vehicle (either in real-time or near real-time). Camera feed 410 may include camera feeds 410a to 410c received from three different front cameras of the vehicle; camera feeds 410d to 410e received from two different right-side cameras of the vehicle; camera feeds 410f to 410g received from two different left-side cameras of the vehicle; and camera feed 410h received from the rear camera of the vehicle.
[0147] Using the methods discussed herein, an analysis server may analyze camera feeds 410 to 450 and generate simulation 400, which is a graphical representation of the vehicle's surrounding surface. Simulation 400 may include a simulated vehicle (420) and its surrounding surface. For example, simulation 400 may visually identify surfaces 430 and 440 using visual attributes (such as different colors (or other visual methods, such as shadow patterns)) to indicate that the AI model has identified surfaces 430 and 440 as drivable surfaces. Simulation 400 may also include surfaces 450 and 492, which are visually different from surfaces 430 and 440 (e.g., different colors or different shadow patterns) because surfaces 450 to 460 have been identified as curbs, and curbs are not drivable surfaces.
[0148] The different surfaces depicted in simulation 400 may visually replicate the predicted elevation (e.g., the Z coordinate value predicted using an AI model). For example, surface 430 (in front of the vehicle) visually indicates that the road in front of the vehicle is a downhill road. In contrast, surface 440 is visually depicted as an uphill road.
[0149] Now referring to Figure 4C , simulation 410 depicts the same surfaces depicted in simulation 400. Specifically, simulation 401 includes a simulated vehicle 420 traveling on surface 430 (the same surface 430 depicted in simulation 400) and surface 440 to the right of the simulated vehicle 420.
[0150] Additionally or alternatively, the analysis server may transmit the generated map to a downstream software application or another server. The prediction results may be further analyzed and used in various models and / or algorithms to perform various actions. For example, a software module or processor associated with the autonomous navigation of the vehicle may receive the occupancy data predicted by the trained AI model, and various navigation decisions may be made accordingly.
[0151] Figure 5 FIG. 500 is a flowchart of a method 500 performed in an AI-enabled visual data analysis system according to an embodiment. Method 500 may include steps 510 to 530. However, other embodiments may include additional or alternative steps, or one or more steps may be completely omitted. Method 500 is described as being performed by an analysis server (e.g., a computer similar to analysis server 110a). However, one or more steps of method 500 may be performed by any number of computing devices (e.g., the processors of vehicle 140 and / or vehicle computing device 141) operating in a distributed computing system as described in Figure 1A and Figure 1B For example, one or more computing devices may perform some or all of the steps described in Figure 5 locally. For example, a chip placed within the vehicle may perform method 500.
[0152] In step 510, the analysis server may input image data of the space around the vehicle object into an artificial intelligence model using the cameras of the vehicle object. The analysis server may collect and / or analyze data received from various cameras of the vehicle (e.g., outward-facing cameras). In another example, the analysis server may collect and aggregate the footage recorded by one or more cameras of the vehicle. The analysis server may then transmit the footage to an AI model trained using the methods discussed herein.
[0153] In step 520, the analysis server may predict the surface properties of one or more surfaces of the space around the vehicle object by executing the artificial intelligence model. The AI model may use the methods discussed herein to identify one or more surfaces around the vehicle. The AI model may also use the data received in step 610 to predict one or more surface properties (e.g., category, material, elevation) of the one or more surfaces.
[0154] In step 530, the analysis server may generate a data set based on the one or more surfaces and their corresponding surface properties. The analysis server may generate a data set including the one or more surfaces and their corresponding surface properties. The data set may be a queryable data set that can be used to transmit the predicted surface data occupancy status to different software modules.
[0155] Figure 6AFIG. 600 is a flow chart of a method performed in an AI - enabled data analysis system according to an embodiment. Method 600 may include steps 610 to 650. However, other embodiments may include additional or alternative steps, or may omit one or more steps entirely. Method 600 is described as being executed by a processor. In some embodiments, the processor may be an analysis server (e.g., a computer similar to analysis server 110a). Alternatively, the processor may be a processor of an autonomous entity (e.g., autonomous computing device 141).
[0156] In some embodiments, one or more steps of method 600 may be performed by any number of computing devices (e.g., processors of autonomous entities 140, analysis server 110a, and / or autonomous computing device 141) operating in a distributed computing system described in Figure 1A and Figure 1B . For example, one or more computing devices may perform some or all of the steps described in Figure 6A locally. For example, a chip placed within an autonomous entity may perform one or more steps of method 600, and the analysis server may perform one or more additional steps of method 500.
[0157] Using method 600, a processor (such as a processor of an autonomous entity) and / or a remote processor (such as an analysis server) may identify / generate a path of the autonomous entity without using external location data received from the autonomous entity. In a non - limiting example, the autonomous entity may be navigating indoors where GPS or other location - indicating data may not be as readily available as when the same autonomous entity is navigating outdoors. The autonomous entity may use method 600 to identify its path without disturbing various obstacles and without the need to receive an indication of the location (or other attributes) of the obstacles. Thus, using method 600, the autonomous entity may autonomously navigate in an indoor environment without the need for external data.
[0158] In step 610, the processor may retrieve image data of the space around the autonomous entity, which is captured by a camera of the autonomous entity. As described herein, the autonomous entity may be equipped with various sensors, including a set of cameras. The autonomous entity may communicate with and retrieve the captured data periodically (sometimes in real - time or near real - time). The captured data may indicate the environment / space in which the autonomous entity is navigating. In some embodiments, the captured data may include navigation data and / or camera feeds received from the autonomous entity, as Figure 7A shown and described.
[0159] At step 620, the processor may predict the occupancy attributes of a plurality of voxels corresponding to the space around the ego by executing an artificial intelligence model. The processor may execute various models discussed herein to analyze the environment in which the ego is located and / or navigating. In a non-limiting example, the ego may execute an occupancy model (discussed in Figure 2A to identify the occupancy status of various voxels within the environment. The processor may also execute a surface detection model (discussed in Figure 2B to identify various surfaces within the environment.
[0160] In some embodiments, the occupancy network and / or the surface detection model may be calibrated and tuned for indoor use. To tune the occupancy / surface model, the ego (e.g., a humanoid robot) and / or any other device (sometimes a human operator) may navigate in various indoor environments, such as an office space, while its various sensors (e.g., cameras and other sensors) collect image data and telemetry data. For example, an employee may carry various sensors (e.g., a telemetry sensor and a camera in a backpack) and walk around within the office space. As the employee walks around the office, the sensors may collect data about the environment, such as camera feeds 660a to 660c, as Figure 6B shown. The recorded data may then be applied to the occupancy and / or surface model so that after comparing the predictions of these models with the actual layout of the same office space, these models may be recalibrated / retrained. In some embodiments, various visual features may be fused with the results received via the model to obtain better results.
[0161] In some embodiments, the processor may use various models to create a real-world representation of the environment. Thus, the processor may create a grid representation of the environment. Additionally, using the telemetry data, the processor may calculate a velocity vector and a yaw rate associated with the ego. As described herein, the path planning module may use a combination of visual data and telemetry data (and additional extracted knowledge) to identify the trajectory of the ego, which may be used to generate a 3D representation of the environment.
[0162] Referring again to Figure 6A , at step 630, the processor may generate a 3D model corresponding to the space around the ego and the occupancy attributes of each voxel. Using the captured sensor data (step 610) and the analyzed and extracted data (step 620), the processor may generate a 3D model of the environment / space in which the ego is located.
[0163] Now referring to Figure 7A , a non-limiting description of the data received from the ego is presented. As described herein, the data may include navigation data and the camera feeds of the ego. However, in some embodiments, the retrieved data may include only image data (camera feeds).
[0164] As used herein, navigation data can include any data related to the navigation of an environment that is collected and / or retrieved by an agent (either autonomously or via a human operator). An agent can rely on a variety of sensors and technologies to collect comprehensive navigation data so that they can navigate autonomously in a variety of environments. Thus, an agent can collect a wide variety of information from the environment in which they navigate. Therefore, navigation data can include any data collected by Figures 1A to 1C any of the sensors discussed herein. Additionally, navigation data can also include any data that is extracted or analyzed using any sensor data, including high-definition maps, trajectory information, and the like. Non-limiting examples of navigation data can include visual inertial odometry (VIO), inertial measurement unit (IMU) data, and / or any data that can indicate the trajectory of an agent.
[0165] In some embodiments, the navigation data can be anonymized. Thus, the analysis server may not receive an indication of which data set / data point belongs to which agent within a group of agents. Anonymization can be performed locally on the agent, e.g., via an agent computing device. Alternatively, anonymization can be performed by another processor before the data is received by the analysis server. In some embodiments, the agent processor / computing device can transmit only data strings without transmitting any agent identification data that would allow the analysis server and / or any other processor to determine which agent generated which data set.
[0166] In addition to retrieving navigation data, the processor can also retrieve image data (e.g., camera feeds or video clips) as the agent navigates within different environments. The image data can include various features located within the environment. As used herein, features within the environment can refer to any physical item located within the environment in which one or more agents navigate. Thus, a feature can correspond to a natural or man-made object. Non-limiting examples of features can include walls, artworks, decorations, and the like.
[0167] Figures 7A to 7B visually depicts how a 3D model is generated. Although Figures 7A to 7B depicts an outdoor navigation scenario, the methods and systems discussed herein are applicable to both indoor and outdoor navigation. Thus, these figures or any other figures presented herein are not intended to be limiting. In some embodiments, the same methods and techniques can be tuned and calibrated for indoor environments.
[0168] Now refer to Figure 7A, the data 700 visually represents the navigation and image data retrieved from the self while the self navigates within the environment. The data 700 may include image data 702, 704, 706, 708, 712, 714, 716, and 718 (collectively referred to as the camera feed 701). When the self navigates within the environment, different cameras collect image data of the self's surrounding environment (e.g., the environment). The camera feed 701 may depict various features located within the environment. For example, the image data 702 depicts various lane lines (e.g., dashed lines dividing four lanes) and trees. The image data 704 depicts the same lane lines and trees from a different angle. The image data 706 depicts the same lane lines from another angle. Additionally, the image data 706 also depicts a building on the other side of the street. The image data 708, 712, 718, 716, and 714 depict the same lane lines. However, some of these image data also depict additional features, such as the traffic lights depicted in the image data 714, 708, and / or 712.
[0169] The navigation data 710 may represent the trajectory of the self from which the Figure 7A depicted image data was collected. The trajectory may be a two-dimensional or three-dimensional trajectory of the self, which is calculated using sensor data retrieved from the self. In some embodiments, various navigation data may be used to determine the trajectory of the self.
[0170] Now referring to Figure 7B , a non-limiting example of a 3D model and its corresponding camera feed is shown. As shown, the image data 720 to the image data 728 represent the camera feed captured by the cameras of the self navigating within the street. Using the camera feed in combination with other navigation data received from the self, the processor may generate a 3D model 730. The 3D model 730 may indicate the position of the self (732) driving within the environment 736. The environment 736 may be a 3D representation that includes features captured as a result of analyzing the camera feed and navigation data of the self. Thus, the environment 736 is similar to the environment in which the self navigates. For example, the sidewalk 738 corresponds to the sidewalk seen in the image data. The model 730 may include all the features identified within the environment, such as traffic lights, road signs, etc. Additionally, the model 730 may include a grid surface for the street on which the self navigates.
[0171] When the self navigates within the environment, the processor may periodically update the 3D model. For example, when the self is first located within a new environment, the self may collect data, and the processor may begin generating a 3D model of the environment. As the self navigates, the processor may continuously collect the self's navigation data and camera feed. Then, the processor may dynamically and continuously update the 3D model.
[0172] Referring again to Figure 6A, at step 650, after receiving the destination, the processor can locate the self by using key image features within the image data corresponding to the 3D model to identify the current position of the self, without receiving the position of the self from a position tracking sensor.
[0173] The processor can use key features extracted from the received image data to locate the self. In some embodiments, the processor can use only the image data because the self may not transmit any position tracking data (since the self is navigating indoors). The processor can first identify key features from the image data. The processor can use a keypoint detector network and a keypoint descriptor network to identify keypoints within the image data. Using these networks, the processor can identify unique points within the received image data. After identifying the keypoints, the processor can track the keypoint in consecutive images (when the self is navigating). Thus, the processor can use only the image data to track the self odometry. However, in some embodiments, other navigation data can also be used. Using this method, the processor can track the position of the self relative to the initial frame (keypoints within the initial frame) and / or the 3D model discussed herein. Thus, the processor can locate the self without the need for GPS or other position tracking data.
[0174] The processor can locate the self periodically. For example, when the self is moving within its path, the processor can locate the self multiple times. Thus, the six-degree pose of the self can be identified at any time. Using this method, the processor can also identify the progress of the self as it moves towards its destination.
[0175] At step 660, the processor can generate a path for the self to move from the current position to the destination.
[0176] The path can include the direction of motion of the self and the corresponding speed along the path such that the self does not collide with any object in the environment. As used herein, the speed can include the forward and / or lateral speed of the self and the yaw rate.
[0177] The processor can use standard trajectory optimization protocols to identify the path using the received destination and the current position (identified based on the localization). For example, the processor can use the Generalized Voronoi Diagram (GVD) protocol, the Rapidly-exploring Random Tree (RRT) protocol, and the Gradient Descent Algorithm (GDA) to identify the trajectory of the self. Additionally or alternatively, the processor can use the Iterative Linear Quadratic Regulator (ILQR) protocol to identify the path / trajectory for the self to reach the destination.
[0178] The processor can locate the self periodically such that the processor can determine the position of the self relative to the destination and whether the self is achieving its goal.
[0179] In some embodiments, an agent can be placed in a new environment. Thus, the agent can initiate an initialization phase, during which the agent navigates within the environment to identify the layout / map of the environment and generate an initial 3D model of the environment. However, even before the 3D model is generated, the agent can be able to navigate based on coordination or distance. For example, the agent can be able to navigate forward five meters without colliding with any obstacles and / or bypassing obstacles.
[0180] Now referring to Figure 8 , an example of an agent analyzing its surrounding environment is depicted. Although Figure 8 the example depicted and described in uses a humanoid agent, the methods and systems discussed herein are applicable to all agents that autonomously navigate, whether indoors or outdoors.
[0181] As shown, agent 840 includes various cameras and captures camera feeds from different angles. For example, camera feeds 800a to 800f represent six different camera feeds captured by agent 840. These camera feeds are collectively referred to as camera feed 800 herein. Using the methods and systems discussed herein, a processor of agent 840 (or another processor communicating with the agent, such as an analysis server) can identify various objects in its surroundings. For example, when agent 840 navigates in an environment represented by 3D model 830, agent 840 can execute various models, such as the AI models discussed herein (e.g., occupancy network or surface network), to analyze its surrounding environment. As described above, agent 840 can analyze the camera feed 800 (of the environment represented within 3D model 830) to identify office furniture (and other obstacles) within the environment represented by 3D model 830. Specifically, chairs 820a and 820b can be represented by representations 810a to 810b (using the occupancy network and surface detection models discussed herein). Using this method, agent 840 can identify the layout of the environment represented by 3D model 830.
[0182] Agent 840 can use various methods to render various images (after it has analyzed them). Rendering can enable a more precise 3D model. For example, as Figure 9 shown, synthetic view rendering techniques can be used to generate image 900. Additionally, as Figure 10 shown, volumetric depth rendering techniques can be used to generate image 1000. These images can be added to the 3D model. As described herein, these images can also be used to calibrate the various models used by agent 840.
[0183] In a non-limiting example, as Figure 11As shown, it can instruct the self-body 1110 to move to a specific destination (e.g., destination 1150). The self-body 1110 can analyze its camera feed and generate a 3D model 1100. The 3D model can include various occupancy and / or surface states of different voxels around the self-body 1110. For example, the 3D model 1100 can include walls 1120 and 1130, which are similar to the walls near the self-body 1110 in real life. The self-body 1110 can use the methods and systems discussed herein to locate itself. Specifically, the self-body 1110 can determine its current position as position 1140 using the 3D model 1100 and / or key image features corresponding to the decoration 1170.
[0184] Using its known current position 1140 and destination 1150, the self-body 1110 can identify a path 1160. As shown, the self-body 1110 may not be able to determine the shortest or straight-line path from its current position 1140 to the destination 1150. This is because, as shown in the 3D model 1100, the straight-line path would interfere with the wall 1130. Instead, the self-body 1110 can determine that the path should be curved, as shown. Specifically, the self-body 1110 can determine that the path 1160 should bend to the left at the end of the wall 1130. Thus, when the self-body 1110 navigates along the path 1160, the self-body 1110 can periodically locate itself to determine the optimal time / position for bending its path (e.g., position 1161 where the wall 1130 ends). For example, the self-body 1110 can use key image features corresponding to the wall decoration 1170 on the wall 1120 to locate itself.
[0185] The various illustrative logical blocks, modules, circuits, and algorithm steps described in connection with the embodiments disclosed herein can be implemented as electronic hardware, computer software, or combinations of both. To clearly illustrate this interchangeability of hardware and software, various illustrative components, blocks, modules, circuits, and steps have been described in terms of their functionality. Whether this functionality is implemented as hardware or software depends upon the particular application and design constraints imposed on the overall system. Skilled artisans may implement the described functionality in varying ways for each particular application, but such implementation decisions should not be interpreted as causing a departure from the scope of the present disclosure or the claims.
[0186] Embodiments implemented in computer software can be implemented in software, firmware, middleware, microcode, hardware description language, or any combination thereof. Code segments or machine-executable instructions can represent a process, function, subroutine, program, routine, subroutine, module, software package, class, or any combination of instructions, data structures, or program statements. By passing and / or receiving information, data, arguments, parameters, or memory contents, a code segment can be coupled to another code segment or hardware circuit. Information, arguments, parameters, data, etc. can be passed, forwarded, or transmitted in any suitable manner, including memory sharing, message passing, token passing, network transmission, etc.
[0187] The actual software code or dedicated control hardware used to implement these systems and methods does not limit the claimed features or this disclosure. Thus, the operation and behavior of the systems and methods can be described without reference to specific software code, and it can be understood that the software and control hardware can be designed to implement these systems and methods based on the description herein.
[0188] When implemented in software, these functions can be stored as one or more instructions or code on a non-transitory computer-readable or processor-readable storage medium. The steps of the methods or algorithms disclosed herein can be embodied in a processor-executable software module that can reside on a computer-readable or processor-readable storage medium. Non-transitory computer-readable or processor-readable media include both computer storage media and tangible storage media, and tangible storage media facilitate the transfer of a computer program from one place to another. The non-transitory processor-readable storage medium can be any available medium accessible by a computer. By way of example and not limitation, such non-transitory processor-readable media can include RAM, ROM, EEPROM, CD-ROM, or other optical disk storage, magnetic disk storage, or other magnetic storage devices, or any other tangible storage medium that can be used to store the desired program code in the form of instructions or data structures and can be accessed by a computer or processor. As used herein, disk and optical disk include compact disc (CD), laser disc, optical disc, digital versatile disc (DVD), Blu-ray disc, and floppy disk, where "disk" generally reproduces data magnetically, and "optical disc" reproduces data optically with a laser. Combinations of the above should also be included within the scope of computer-readable media. In addition, the operations of the methods or algorithms can reside as one or any combination or set of code and / or instructions on a non-transitory processor-readable medium and / or computer-readable medium, which can be incorporated into a computer program product.
[0189] The foregoing description of the disclosed embodiments is provided to enable a person skilled in the art to make or use the embodiments and their variations described herein. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the principles defined herein may be applied to other embodiments without departing from the spirit or scope of the subject matter disclosed herein. Accordingly, the disclosure is not intended to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the following claims and the principles and novel features disclosed herein.
[0190] Although various aspects and embodiments have been disclosed, other aspects and embodiments may also be contemplated. The disclosed aspects and embodiments are for illustrative purposes and are not intended to be limiting, and the true scope and spirit are indicated by the following claims.
Claims
1. A method, comprising: Retrieving, by a processor, image data of a space around an ego, the image data being captured by a camera of the ego; Predicting, by the processor, occupancy attributes of a plurality of voxels corresponding to the space around the ego by executing an artificial intelligence model; Generating, by the processor, a 3D model corresponding to the space around the ego and the occupancy attributes of each voxel; After receiving a destination, positioning, by the processor, the ego by using key image features in the image data corresponding to the 3D model to identify a current position of the ego, without receiving the position of the ego from a position tracking sensor; And Generating, by the processor, a path for the ego to travel from the current position to the destination.
2. The method according to claim 1 further comprises: Periodically positioning, by the processor, the ego during the path.
3. The method according to claim 1, wherein locating the autologous tissue comprises: Tracking the key image features in consecutive image data.
4. The method according to claim 1, wherein the key image points correspond to unique points in the image data.
5. The method according to claim 1, wherein generating the path comprises: Generating at least one of a trajectory, a yaw rate, a forward speed, or a lateral speed for the ego.
6. The method according to claim 1, wherein the path is generated using an Iterative Linear Quadratic Regulator (ILQR) protocol.
7. The method according to claim 1, wherein the 3D model also corresponds to surface attributes of at least one object in the space around the ego.
8. A computer system, comprising: A non-transitory computer-readable medium having a set of instructions that, when executed, cause a processor to: Retrieve image data of a space around an ego, the image data being captured by a camera of the ego; Predict occupancy attributes of a plurality of voxels corresponding to the space around the ego by executing an artificial intelligence model; Generate a 3D model corresponding to the space around the ego and the occupancy attributes of each voxel; After receiving a destination, position the ego by using key image features in the image data corresponding to the 3D model to identify a current position of the ego, without receiving the position of the ego from a position tracking sensor; and Generate a path for the ego to travel from the current position to the destination.
9. The computer system according to claim 8, wherein the set of instructions further causes the processor to periodically position the ego during the path.
10. The computer system according to claim 8, wherein locating the self comprises: Tracking the key image features in consecutive image data.
11. The computer system according to claim 8, wherein the key image points correspond to unique points in the image data.
12. The computer system according to claim 8, wherein generating the path comprises: Generating at least one of a trajectory, a yaw rate, a forward speed, or a lateral speed for the ego.
13. The computer system according to claim 8, wherein the path is generated using an Iterative Linear Quadratic Regulator (ILQR) protocol.
14. The computer system according to claim 8, wherein the 3D model also corresponds to surface attributes of at least one object in the space around the ego.
15. An ego, comprising: A processor, the processor being configured to: Retrieve image data of the space around the self, the image data being captured by a camera of the self; Predict occupancy attributes of a plurality of voxels corresponding to the space around the self by executing an artificial intelligence model; Generate a 3D model corresponding to the space around the self and the occupancy attributes of each voxel; After receiving a destination, locate the self by using key image features within the image data corresponding to the 3D model to identify the current position of the self without receiving the position of the self from a position tracking sensor; and Generate a path for the self to move from the current position to the destination.
16. The self according to claim 15, wherein the processor is further configured to periodically locate the self during the path.
17. The self-body according to claim 15, wherein positioning the self-body comprises: Track the key image features in consecutive image data.
18. The self according to claim 15, wherein the key image points correspond to unique points within the image data.
19. The self body according to claim 15, wherein generating the path includes: Generate at least one of a trajectory, a yaw rate, a forward speed, or a lateral speed for the self.
20. The self according to claim 15, wherein the path is generated using an Iterative Linear Quadratic Regulator (ILQR) protocol.