Artificial intelligence modeling techniques for vision-based surface determination
By using artificial intelligence-based models in autonomous vehicles or robots, analyzing the image data of the surrounding environment and predicting the occupancy attributes of the object, the problem of difficulty in effectively analyzing and predicting the surrounding environment of autonomous vehicles or robots in the prior art is solved, and efficient autonomous navigation is achieved.
Patent Information
- Application Number
- CN202380071207.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-09-30
- Filing Date
- 2023-09-28
- Publication Date
- 2025-05-30
AI Technical Summary
The prior art is difficult to effectively analyze the surrounding environment of autonomous vehicles or robots and predict the occupancy properties of objects in the environment, especially in complex and dynamic environments.
Using an artificial intelligence-based model, the surrounding environment image data of autonomous vehicles or robots is analyzed through the trained AI model, the occupancy data associated with the space around autonomous vehicles or robots is predicted, and the surface properties of the object are determined.
It realizes navigation without the need for autonomous vehicles or robots to position themselves, and can predict and analyze the occupied state in the surrounding environment in real time or near real time, improving the safety and reliability of autonomous navigation.
Smart Images

Figure CN120077414A_ABST
Abstract
Description
[0001] Cross - Reference to Related Applications
[0002] This application claims the benefit of U.S. Provisional Application No. 63 / 377,954, filed Sep. 30, 2022, which is hereby incorporated by reference in its entirety for all purposes. Technical Field
[0003] The present disclosure generally relates to artificial intelligence-based modeling techniques for analyzing image data and predicting occupancy attributes of an environment surrounding an ego. Background Art
[0004] Due to the rapid development of computer technology, autonomous navigation technologies for autonomous vehicles and robots (collectively referred to as ego) have become ubiquitous. These advancements allow for safer and more reliable autonomous navigation of the ego. The ego typically needs to navigate in complex and dynamic environments and terrains, which can include vehicles, traffic, pedestrians, cyclists, and various other static or dynamic obstacles. Understanding the ego's surrounding environment is necessary for informed and capable decision-making to avoid collisions. Summary of the Invention
[0005] For the above reasons, there is a desire for methods and systems that can analyze the ego's surrounding environment and predict the presence of objects with mass within the ego's surrounding environment. Specifically, a trained AI model used within a particular artificial intelligence (AI) architecture can predict occupancy data associated with the space surrounding the ego. As used herein, occupancy data or occupancy attributes can refer to whether a defined space is occupied by an object with mass (e.g., occupied or unoccupied).
[0006] There is also a desire for methods and systems that can analyze the ego's surrounding environment and identify / evaluate different surfaces of objects that occupy the ego's surrounding environment. Specifically, a trained AI model can determine the surface attributes of an object that occupies the space surrounding the ego.
[0007] Using the methods and systems discussed herein, a trained AI model can generate a data set that corresponds to a three-dimensional (3D) representation of the surrounding environment of the ego. The AI model can use image data received from the ego's camera and predict the 3D structure of the driving surface around the ego. As used herein, the driving surface is applicable to any navigation of the ego or any other vehicle on the surface, whether or not the ego is designated for driving and whether the navigation is performed autonomously or via an operator. The data set can be used by autonomous navigation software and / or a processor to navigate the ego. Using the methods and systems discussed herein, the AI model can determine whether the surface is navigable, whether the surface includes a hilltop (and if so, how steep the hilltop is), whether the hilltop is uphill or downhill, whether the road is hilly, the location of the lanes and / or shoulders, whether the road has any markings (paint lines), where there are any speed bumps or potholes, etc. That is, the AI model discussed herein can be used to analyze the driving surface. For example, the AI model can determine whether a high ramp is inclined or a simple curve; whether there are any elevation changes on the ramp, etc.
[0008] Using the AI model discussed herein allows the ego to navigate without the need for the ego to locate itself. This is (at least in part) due to the fact that the AI model can predict the surrounding environment of the ego in real-time using images captured by the ego's (multiple) cameras, even if no image data has ever been ingested to train the AI model. For example, using the methods and systems discussed herein, an AI model can be trained such that the trained AI model can determine the surface / occupancy properties of a road that the AI model has never "seen" before.
[0009] In an embodiment, a method can include: inputting, by a processor, image data of the space around an ego object into an artificial intelligence model using one or more cameras of the ego object; predicting, by the processor executing the artificial intelligence model, surface properties of one or more surfaces of the space around the ego object; and generating, by the processor, a data set based on the one or more surfaces and their corresponding surface properties.
[0010] The method can further include: generating, by the processor, an output that represents the space around the ego object and illustrates the one or more surfaces and their corresponding surface properties.
[0011] The one or more surfaces are visually different according to their corresponding surface properties.
[0012] The method can further include: displaying, by the processor, the output on a screen associated with the ego object.
[0013] The data set can be a queryable data set configured to send surface properties to an autonomous driving protocol of the ego object.
[0014] An artificial intelligence model can be trained using sensor properties of one or more surfaces.
[0015] An ego object can be an autonomous vehicle that executes a driving protocol based on a data set.
[0016] The surface property can correspond to elevation.
[0017] The surface property can indicate whether one or more surfaces are navigable surfaces.
[0018] The surface property can indicate the material associated with one or more surfaces.
[0019] The surface property can indicate whether one or more surfaces correspond to one or more curbs.
[0020] In another embodiment, an ego object can include: one or more cameras; a processor; a non-transitory computer-readable medium containing an artificial intelligence model configured to be executed by the first processor, wherein the processor is configured to: input image data of the space around the ego object into the artificial intelligence model using one or more cameras of the ego object; predict surface properties of one or more surfaces of the space around the ego object by executing the artificial intelligence model; and generate a data set based on the one or more surfaces and their corresponding surface properties.
[0021] The instruction processor can also cause the processor to generate an output that represents the space around the ego object and illustrates one or more surfaces and their corresponding surface properties.
[0022] One or more surfaces can be visually different according to their corresponding surface properties.
[0023] The instruction can also cause the processor to display the output on a screen associated with the ego object.
[0024] The data set can be a queryable data set configured to send surface properties to the autonomous driving protocol of the ego object.
[0025] An artificial intelligence model can be trained using sensor properties of one or more surfaces.
[0026] An ego object can be an autonomous vehicle that executes a driving protocol based on a data set.
[0027] The surface property can correspond to elevation.
[0028] The surface property can indicate whether one or more surfaces are navigable surfaces. Brief Description of the Drawings
[0029] Non-limiting embodiments of the present disclosure are described by way of example in relation to the accompanying drawings, which are schematic and are not intended to be drawn to scale. Unless indicated as representing background art, the drawings represent various aspects of the present disclosure.
[0030] Figure 1A Components of an AI-enabled visual data analysis system according to an embodiment are illustrated.
[0031] Figure 1B Various sensors associated with the body according to an embodiment are illustrated.
[0032] Figure 1C Components of a vehicle according to an embodiment are illustrated.
[0033] Figures 2A - 2B Illustrated is a flow chart of different processes performed in an AI-enabled visual data analysis system according to an embodiment.
[0034] Figures 3A - 3B Illustrated are different occupancy maps generated in an AI-enabled visual data analysis system according to an embodiment.
[0035] Figures 4A - 4C Illustrated are different views of a surface map generated in an AI-enabled visual data analytics system according to an embodiment.
[0036] Figure 5 A flow chart of a process for executing an AI model to generate a surface map according to an embodiment is illustrated. DETAILED DESCRIPTION
[0037] Reference will now be made to the illustrative embodiments depicted in the accompanying drawings, and specific language will be used herein to describe them. However, it is to be understood that it is not intended to limit the scope of the claims or the disclosure. Changes and further modifications to the inventive features described herein that may be conceived by those skilled in the relevant art and that possess the present disclosure and additional applications of the subject principles described herein will be considered to be within the scope of the subject matter disclosed herein. Other embodiments may be used and / or other changes may be made without departing from the spirit or scope of the present disclosure. The illustrative embodiments described in the specific embodiments are not meant to limit the subject matter presented.
[0038] By implementing the methods described herein, a system can use a trained AI model to determine the occupancy status of different voxels of an image (or video) of the surrounding environment of an ego. The ego can be an autonomous vehicle (e.g., a car, truck, bus, motorcycle, all-terrain vehicle, cart), a robot, or other automated device. The ego can be configured to operate within a production line, building, home, or medical center, or to transport people, deliver goods, perform military functions, etc. Within these environments, the ego can navigate between known or unknown paths to complete a particular task or travel to a particular destination. During operation, it is desirable to avoid collisions, so the ego seeks to understand the environment. For example, in the context of an autonomous vehicle or robot, the system can use a camera (or other vision sensor) to receive real-time or near-real-time images of the surrounding environment of the ego. The system can then execute the trained AI model to determine the occupancy status of the surrounding environment of the ego. The AI model can divide the surrounding environment of the ego into different voxels and then determine the occupancy status for each voxel. Thus, using the methods discussed herein, the system can generate a map of the surrounding environment of the ego. Using voxel data (e.g., the coordinates of each voxel) and the corresponding occupancy status, the AI model (or sometimes another model using data predicted by the AI model) can generate a map of the surrounding environment of the ego.
[0039] Figure 1A are non-limiting examples of components of a system in which the methods and systems discussed herein can be implemented. For example, an analysis server can train an AI model and use the trained AI model to generate an occupancy dataset and / or map for one or more egos. Figure 1A Illustrates the components of an AI-enabled visual data analysis system 100. The system 100 can include an analysis server 110a, a system database 110b, an administrator computing device 120, egos 140a to 140b (collectively referred to as (multiple) egos 140), ego computing devices 141a to 141c (collectively referred to as ego computing devices 141), and a server 160. The system 100 is not limited to the components described herein and can include additional or other components not shown for the sake of brevity, which will be considered within the scope of the embodiments described herein.
[0040] The above components can be connected via a network 130. Examples of the network 130 can include, but are not limited to, a private or public LAN, WLAN, MAN, WAN, and the Internet. The network 130 can include wired and / or wireless communication according to one or more standards and / or via one or more transmission media.
[0041] Communication on network 130 can be performed according to various communication protocols such as Transmission Control Protocol and Internet Protocol (TCP / IP), User Datagram Protocol (UDP), and IEEE communication protocols. In one example, network 130 can include wireless communication according to a Bluetooth specification set or another standard or proprietary wireless communication protocol. In another example, network 130 can also include communication over a cellular network, including for example GSM (Global System for Mobile Communications), CDMA (Code Division Multiple Access), or EDGE (Enhanced Data for Global Evolution) network.
[0042] System 100 illustrates an example of a system architecture and components that can be used to train and execute one or more AI models such as (multiple) AI models 110c. Specifically, as Figure 1A depicted and described herein, analysis server 110a can use the methods discussed herein to train (multiple) AI models 110c using data retrieved from self 140 (e.g., by using data streams 172 and 176). When (multiple) AI models 110c have been trained, each self in self 140 can access and execute (multiple) trained AI models 110c. For example, vehicle 140a having a self computing device 141a can send its camera feed to (multiple) trained AI models 110c and can determine the occupancy status of its surrounding environment (e.g., data stream 174). Moreover, data ingested and / or predicted by (multiple) AI models 110c with respect to self 140 (at inference time) can also be used to improve (multiple) AI models 110c. Thus, system 100 depicts a continuous loop that can periodically improve the accuracy of (multiple) AI models 110c. Moreover, system 100 depicts a loop where, in addition to the inference phase, data received by self 140 can also be used in the training phase.
[0043] Analysis server 110a can be configured to collect, process, and analyze navigation data (e.g., images captured during navigation) and various sensor data collected from self 140. Then, the collected data can be processed and prepared into a training dataset. Then, the training dataset can be used to train one or more AI models such as AI model 110c. Analysis server 110a can also be configured to collect visual data from self 140. Using AI model 110c (trained using the methods and systems discussed herein), analysis server 110a can generate a dataset and / or occupancy map for self 140. Analysis server 110a can display the occupancy map on self 140 and / or send the occupancy map / dataset to self computing device 141, administrator computing device 120, and / or server 160.
[0044] In Figure 1AIn [the figure], the AI model 110c is illustrated as a component of the system database 110b, but the AI model 110c can be stored in different or separate components, such as a cloud storage device or any other data repository accessible by the analytics server 110a.
[0045] The analytics server 110a can also be configured to display an electronic platform that illustrates various training attributes for training the AI model 110c. The electronic platform can be displayed on the administrator computing device 120 such that an analyst can monitor the training of the AI model 110c. An example of an electronic platform generated and hosted by the analytics server 110a can be a web-based application or website configured to display the training data set collected from the self 140 and / or the training status / metrics of the AI model 110c.
[0046] The analytics server 110a can be any computing device including a processor and a non-transitory machine-readable storage device capable of performing the various tasks and processes described herein. Non-limiting examples of such computing devices can include workstation computers, laptop computers, server computers, etc. Although the system 100 includes a single analytics server 110a, the system 100 can include any number of computing devices operating in a distributed computing environment, such as a cloud environment.
[0047] The self 140 can represent various electronic data sources that send data associated with their previous or current navigation sessions to the analytics server 110a. The self 140 can be any device configured for navigation, such as a vehicle 140a and / or a truck 140c. The self 140 is not limited to being a vehicle and can also include robotic devices. For example, the self 140 can include a robot 140b, which can represent a general-purpose, bipedal, autonomous humanoid robot capable of navigating various terrains. The robot 140b can be equipped with software for achieving balance, navigation, perception, or interacting with the physical world. The robot 140b can also include various cameras configured to send visual data to the analytics server 110a.
[0048] Although referred to herein as the "self", the self 140 may or may not be an autonomous device configured for autonomous navigation. For example, in some embodiments, the self 140 can be controlled by a human operator or a remote processor. The self 140 can include various sensors, such as Figure 1BThe sensors depicted herein. The sensors can be configured to collect data as the vehicle 140 navigates various terrains (e.g., roads). The analysis server 110a can collect the data provided by the vehicle 140. For example, the analysis server 110a can obtain navigation sessions and / or road / terrain data (e.g., images of the vehicle 140 navigating on a road) from various sensors, such that the collected data is ultimately used by the AI model 110c for training purposes.
[0049] As used herein, a navigation session corresponds to the journey of the vehicle 140's travel route, regardless of whether the journey is autonomous or human-controlled. In some embodiments, the navigation session can be used for data collection and model training purposes. However, in some other embodiments, the vehicle 140 can refer to a vehicle purchased by a consumer, and the purpose of the journey can be classified as daily use. A navigation session can begin when the vehicle 140 moves from a non-moving position by more than a threshold distance (e.g., 0.1 mile, 100 feet) or at a rate greater than a threshold rate (e.g., greater than 0 mph, greater than 1 mph, greater than 5 mph). A navigation session can end when the vehicle 140 returns to a non-moving position and / or is turned off (e.g., when the driver exits the vehicle).
[0050] The vehicle 140 can represent a collection of vehicles monitored by the analysis server 110a to train the AI model(s) 110c. For example, the driver of the vehicle 140a can authorize the analysis server 110a to monitor data associated with their corresponding vehicle. As a result, the analysis server 110a can utilize the various methods discussed herein to collect sensor / camera data and generate a training dataset to train the AI model(s) 110c accordingly. The analysis server 110a can then apply the trained AI model(s) 110c to analyze data associated with the vehicle 140 and predict an occupancy map for the vehicle 140. Moreover, additional / ongoing data associated with the vehicle 140 can also be processed and added to the training dataset so that the analysis server 110a can recalibrate the AI model(s) 110c accordingly. Thus, the system 100 depicts a loop in which navigation data received from the vehicle 140 can be used to train the AI model(s) 110c. The vehicle 140 can include a processor that executes the trained AI model(s) 110c for navigation purposes. During navigation, the vehicle 140 can collect additional data about its navigation session, and this additional data can be used to calibrate the AI model(s) 110c. That is, the vehicle 140 represents a vehicle that can be used to train, execute / use, and recalibrate the AI model(s) 110c. In a non-limiting example, the vehicle 140 represents vehicles purchased by customers that can use the AI model(s) 110c to navigate autonomously while improving the AI model(s) 110c.
[0051] The self - body 140 can be equipped with various technologies that allow the self - body to collect data from its surrounding environment and (possibly) navigate autonomously. For example, the self - body 140 can be equipped with an inference chip to run autonomous driving software.
[0052] Various sensors for each self - body 140 can monitor the data collected associated with different navigation sessions and send the data to the analysis server 110a. Figures 1B - 1C A block diagram of sensors integrated within the self - body 140 according to an embodiment is illustrated. The quantity and location of each sensor discussed with respect to Figures 1B - 1C can depend on the type of self - body discussed in Figure 1A For example, the robot 140b can include different sensors from the vehicle 140a or the truck 140c. For example, the robot 140b may not include an airbag activation sensor 170q. Also, the sensors of the vehicle 140a and the truck 140c can be positioned differently from that Figure 1C illustrated.
[0053] As discussed herein, the various sensors integrated within each self - body 140 can be configured to measure various data associated with each navigation session. The analysis server 110a can periodically collect the data monitored and collected by these sensors, where the data is processed according to the methods described herein and is used to train the AI model 110c and / or execute the AI model 110c to generate an occupancy map.
[0054] The self - body 140 can include a user interface 170a. The user interface 170a can refer to the user interface of the self - body computing device (such as Figure 1A the self - body computing device 141 in Figure 1B ). The user interface 170a can be implemented as a display screen, a head - up display, a touch screen, etc. integrated with or coupled to the interior of the vehicle. The user interface 170a can include input devices such as a touch screen, a knob, a button, a keyboard, a mouse, a gesture sensor, a steering wheel, etc. In various embodiments, the user interface 170a can be adapted to provide user input (such as a signal and / or sensor information) to other devices or sensors (such as
[0055] The user interface 170a can also be implemented with one or more logical devices that can be adapted to execute instructions, such as software instructions, to implement any of the various processes and / or methods described herein. For example, the user interface 170a can be adapted to form a communication link, send and / or receive communications (such as sensor signals, control signals, sensor information, user input, and / or other information), or perform various other processes and / or methods. In another example, a driver can use the user interface 170a to control the temperature of the vehicle body 140 or activate its features (such as the autonomous driving or steering system 170o). Thus, the user interface 170a can be combined with other sensors described herein to monitor and collect driving session data. The user interface 170a can also be configured to display various data generated / predicted by the analysis server 110a and / or the AI model 110c.
[0056] The orientation sensor 170b can be implemented as a compass, a float, an accelerometer, and / or one or more of any other digital or analog devices capable of measuring the orientation of the vehicle body 140 (such as the magnitude and direction of roll, pitch, and / or yaw relative to one or more reference orientations, such as gravity and / or magnetic north). The orientation sensor 170b can be adapted to provide heading measurements to the vehicle body 140. In other embodiments, the orientation sensor 170b can be adapted to provide roll, pitch, and / or yaw rates to the vehicle body 140 using a time series of orientation measurements. The orientation sensor 170b can be positioned and / or adapted to make orientation measurements relative to a particular coordinate system of the vehicle body 140.
[0057] The controller 170c can be implemented as any suitable logical device (such as a processing device, a microcontroller, a processor, an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), a memory storage device, a memory reader, or other device or combination of devices) that can be adapted to execute, store, and / or receive appropriate instructions, such as software instructions that implement control loops for controlling various operations of the vehicle body 140. Such software instructions can also implement methods for processing sensor signals, determining sensor information, providing user feedback (such as through the user interface 170a), querying devices for operating parameters, selecting operating parameters for devices, or performing any of the various operations described herein.
[0058] The communication module 170e can be implemented as any wired and / or wireless interface that is configured to transmit sensor data, configuration data, parameters, and / or other data and / or signals to Figure 1A any of the features shown (such as the analysis server 110a). As described herein, in some embodiments, the communication module 170e can be implemented in a distributed manner such that portions of the communication module 170e are inFigure 1B implemented within one or more of the illustrated elements and sensors. In some embodiments, the communication module 170e may delay transmission of sensor data. For example, when the vehicle 140 does not have network connectivity, the communication module 170e may store the sensor data in a temporary data storage device and send the sensor data when the vehicle 140 is identified as having appropriate network connectivity.
[0059] The speed sensor 170d may be implemented as an electronic pitot tube, a metering gear or wheel, a water speed sensor, a wind speed sensor, a wind rate sensor (e.g., direction and magnitude), and / or other devices capable of measuring or determining the linear speed of the vehicle 140 (e.g., in the surrounding medium and / or aligned with the longitudinal axis of the vehicle 140) and providing such a measurement as a sensor signal, which may be transmitted to various devices.
[0060] The gyroscope / accelerometer 170f may be implemented as one or more electronic sextants, semiconductor devices, integrated chips, accelerometer sensors, or other systems or devices capable of measuring the angular velocity / acceleration and / or linear acceleration of the vehicle 140 (e.g., direction and magnitude) and providing such a measurement as a sensor signal, which may be transmitted to various devices, such as the analysis server 110a. The gyroscope / accelerometer 170f may be positioned and / or adapted to make such measurements with respect to a particular coordinate system of the vehicle 140. In various embodiments, the gyroscope / accelerometer 170f may be implemented in a common housing and / or module with Figure 1B the other elements depicted to ensure a common reference frame or a known transformation between reference frames.
[0061] The Global Navigation Satellite System (GNSS) 170h may be implemented as a global positioning satellite receiver and / or other device capable of determining the absolute and / or relative position of the vehicle 140 based on wireless signals received, for example, from space and / or ground sources and capable of providing measurements such as sensor signals, which may be transmitted to various devices. In some embodiments, the GNSS 170h may be adapted to determine the rate, speed, and / or yaw rate of the vehicle 140 (e.g., using a time series of position measurements), such as the yaw component of the absolute rate and / or angular rate of the vehicle 140.
[0062] The temperature sensor 170i may be implemented as a thermistor, an electrical sensor, an electrical thermometer, and / or other devices capable of measuring the temperature associated with the vehicle 140 and providing such a measurement as a sensor signal. The temperature sensor 170i may be configured to measure the ambient temperature associated with the vehicle 140, such as the cockpit or dashboard temperature, and for example, this ambient temperature may be used to estimate the temperature of one or more elements of the vehicle 140.
[0063] The humidity sensor 170j may be implemented as a relative humidity sensor, an electrical sensor, an electrical relative humidity sensor, and / or another device capable of measuring the relative humidity associated with the body 140 and providing such measurement as a sensor signal.
[0064] The steering sensor 170g may be adapted to physically adjust the heading of the body 140 in accordance with one or more control signals provided by a logic device (such as the controller 170c) and / or user input. The steering sensor 170g may include one or more actuators and control surfaces (such as a rudder or other type of steering or trim mechanism) of the body 140, and may be adapted to physically adjust the control surface to various positive and / or negative steering angles / positions. The steering sensor 170g may also be adapted to sense the current steering angle / position of such steering mechanism, and provide such measurement.
[0065] The propulsion system 170k may be implemented as a propeller, a turbine or other thrust-based propulsion system, a mechanical wheeled and / or tracked propulsion system, a wind / sail-based propulsion system, and / or other types of propulsion systems that may be used to provide power to the body 140. The propulsion system 170k may also monitor the direction of the power and / or thrust of the body 140 relative to the reference coordinate system of the body 140. In some embodiments, the propulsion system 170k may be coupled to the sensor 170g and / or integrated with the steering sensor 170g.
[0066] The passenger restraint sensor 170l may monitor seatbelt detection and locking / unlocking assemblies and other passenger restraint subsystems. The passenger restraint sensor 170l may include various environmental and / or status sensors, actuators, and / or other devices that facilitate the operation of safety mechanisms associated with the operation of the body 140. For example, the passenger restraint sensor 170l may be configured to receive motion and / or status data from Figure 1B the other sensors depicted. The passenger restraint sensor 170l may determine whether a safety measurement (such as a seatbelt) is being used.
[0067] As Figure 1C depicted, the camera 170m may refer to one or more cameras integrated within the body 140, and may include multiple cameras integrated (or retrofitted) into the body 140. The camera 170m may be an internal or external facing camera of the body 140. For example, as Figure 1C depicted, the body 140 may include one or more internal facing cameras 170m-1. These cameras may monitor and collect footage of the passengers of the body 140. The body 140 may also include a front-facing side camera 170m-2, a camera 170m-3 (such as integrated within a doorframe), and a rear-facing side camera 170m-4.
[0068] Reference Figure 1B , the radar 170n and the ultrasonic sensor 170p can be configured to monitor the distance of the host 140 to other objects (such as other vehicles or immovable objects (e.g., trees or garage doors)). As Figure 1C depicted, the radar 170n and the ultrasonic sensor 170p can be integrated into the host 140. The host 140 may also include an autonomous driving or steering system 170o configured to autonomously navigate the host 140 using data collected via various sensors (such as the radar 170n, the speed sensor 170d, and / or the ultrasonic sensor 170p).
[0069] Accordingly, the autonomous driving or steering system 170o can analyze various data collected by one or more of the sensors described herein to identify driving data. For example, the autonomous driving or steering system 170o can calculate the risk of a frontal collision based on the speed of the host 140 and its distance to another vehicle on the road. The autonomous driving or steering system 170o can also determine whether the driver is touching the steering wheel. The autonomous driving or steering system 170o can send the analyzed data to various features discussed herein, such as the analysis server.
[0070] The airbag activation sensor 170q can predict or detect a collision and cause the activation or deployment of one or more airbags. The airbag activation sensor 170q can send data regarding the airbag deployment, including data associated with the event that caused the deployment.
[0071] Referring again to Figure 1A , the administrator computing device 120 can represent a computing device operated by a system administrator. The administrator computing device 120 can be configured to display data retrieved or generated by the analysis server 110a (such as various analysis metrics and risk scores), where the system administrator can monitor the various models utilized by the analysis server 110a, review feedback, and / or facilitate the training of the AI model(s) 110c maintained by the analysis server 110a.
[0072] The host(s) 140 can be any device configured to navigate various routes, such as the vehicle 140a or the robot 140b. As with respect to Figures 1B - 1CAs discussed, the ego 140 can include various telemetry sensors. The ego 140 can also include an ego computing device 141. Specifically, each ego can have its own ego computing device 141. For example, the truck 140c can have an ego computing device 141c. For simplicity, the ego computing devices are collectively referred to as the (plural) ego computing devices 141. The ego computing device 141 can control content presentation on the infotainment system of the ego 140, process commands associated with the infotainment system, aggregate sensor data, manage communication of data to an electronic data source, receive updates, and / or send messages. In one configuration, the ego computing device 141 communicates with an electronic control unit. In another configuration, the ego computing device 141 is an electronic control unit. The ego computing device 141 can include a processor and a non-transitory machine-readable storage medium capable of performing the various tasks and processes described herein. For example, the (plural) AI models 110c described herein can be stored and executed (or directly accessed) by the ego computing device 141. Non-limiting examples of the ego computing device 141 can include vehicle multimedia and / or display systems.
[0073] In one example of how the (plural) AI models 110c can be trained, the analytics server 110a can collect data from the ego 140 to train the (plural) AI models 110c. Before executing the (plural) AI models 110c to generate / predict an occupancy dataset, the analytics server 110a can use various methods to train the (plural) AI models 110c. Training allows the (plural) AI models 110c to ingest data from one or more cameras of one or more egos 140 (without receiving radar data) and predict occupancy data for the surrounding environment of the ego. The operations described in this example can be performed by any number of computing devices (such as the processors of the egos 140) operating in the Figure 1A and 1B distributed computing system described in.
[0074] The analytics server 110a can use the sensors of the ego 140 to generate a first dataset having a first set of data points, where each data point within the first set of data points corresponds to the position and image attributes of at least one voxel of the space around the ego 140, and the sensor attribute indicates whether at least one voxel is occupied by an object having mass.
[0075] To train the AI model(s) 110c, the analysis server 110a can first use one or more of the vehicles 140 to drive a specific route. While driving, the vehicle 140 can use one or more of its sensors (including one or more cameras) to generate navigation session data. For example, one or more vehicles 140 equipped with various sensors can navigate a designated route. As one or more of the vehicles 140 traverse the terrain, their sensors can capture continuous (or periodic) data of their surrounding environment. The sensors can indicate the occupancy status of the surrounding environment of one or more of the vehicles 140. For example, the sensor data can indicate various objects with mass in the surrounding environment of one or more of the vehicles 140 as they navigate their route.
[0076] The analysis server 110a can use the sensor data received from one or more of the vehicles 140 to a first training dataset. The first dataset can indicate the occupancy status of different voxels within the surrounding environment of one or more of the vehicles 140. As used herein in some embodiments, a voxel is a three-dimensional pixel that forms the building blocks of the surrounding environment of one or more of the vehicles 140. Within the first dataset, each voxel can encapsulate sensor data that indicates whether a mass is identified for that specific voxel. As used herein, a mass can indicate or represent any object identified using a sensor. For example, in some embodiments, the vehicle 140 can be equipped with LiDAR that identifies a mass by emitting laser pulses and measuring the time it takes for these pulses to reach an object (with mass) and return. The LiDAR sensor system can operate based on the principle of measuring the distance between the LiDAR sensor and an object in its field of view. Combined with other sensor data, this information can be analyzed to identify and characterize different masses or objects within the surrounding environment of one or more of the vehicles 140.
[0077] Various additional data can be used to indicate whether a voxel in the surrounding environment of one or more of the vehicles 140 is occupied by an object with mass. For example, in some embodiments, a digital map of the surrounding environment of one or more of the vehicles 140 (e.g., a digital map of the route the vehicle is traversing) can be used to determine the occupancy status of each voxel.
[0078] In operation, as depicted by the data stream 176, when one or more of the vehicles 140 are navigating, their sensors collect data and send the data to the analysis server 110a. For example, the vehicle 140 computing device 141 can use the data stream 176 to send the sensor data to the analysis server 110a.
[0079] The analysis server 110a can use the camera of the ego body 140 to generate a second data set with a second data point set, where each data point in the second data point set corresponds to the position and image attributes of at least one voxel in the space around the ego body 140.
[0080] The analysis server 110a can receive camera feeds of one or more ego bodies 140 that navigate the same route as in the first step. In some embodiments, the analysis server 110a can perform the first step and the second step simultaneously (or concurrently). Alternatively, two (or more) different ego bodies 140 can navigate the same route, where one ego body sends its sensor data and the second ego body 140 sends its camera feed.
[0081] One or more ego bodies 140 can include one or more high-resolution cameras that capture a continuous visual data stream from the surrounding environment of the one or more ego bodies 140 as the one or more ego bodies 140 navigate along the route. The analysis server 110a can then use the camera feed to generate a second data set, where the visual elements / depictions of different voxels in the surrounding environment of the one or more ego bodies 140 are included in the second data set.
[0082] In operation, as one or more ego bodies 140 navigate, their cameras collect data and send the data to the analysis server 110a, as depicted by the data stream 172. For example, the ego computing device 141 can use the data stream 172 to send image data to the analysis server 110a.
[0083] The analysis server 110a can use the first data set and the second data set to train the AI model, whereby the AI model 110c trains itself using the corresponding positions of each data point, associating each data point in the first data point set with the corresponding data point in the second data point set. After being trained, the AI model 110c is configured to receive a camera feed from a new ego body 140 and predict the occupancy status of at least one voxel of the camera feed.
[0084] Using the first data set and the second data set, the analysis server 110a can train the AI model(s) 110c such that the AI model(s) 110c can associate different visual attributes of voxels (within the camera feed in the second data set) with the occupancy status of the voxels (within the first data set). In this way, after being trained, the AI model(s) 110c can receive a camera feed (e.g., from a new ego body 140) without receiving sensor data and then determine the occupancy status of each voxel for the new ego body 140.
[0085] The analysis server 110a can generate a training dataset including a first dataset and a second dataset. The analysis server 110a can use the first dataset as the ground truth. For example, the first dataset can indicate the different positions of the voxels and their occupancy status. The second dataset can include visual (e.g., camera feed) illustrations of the same voxels. Using the first dataset, the analysis server 110a can label the data such that the (multiple) data records associated with each voxel corresponding to an object are indicated as having a positive occupancy status.
[0086] The labeling of the occupancy status of different voxels can be performed automatically and / or manually. For example, in some embodiments, the analysis server 110a can use a human reviewer to label the data. For example, as discussed herein, the camera feed from one or more cameras of a vehicle can be shown to a human reviewer on an electronic platform for labeling. Additionally or alternatively, the (multiple) AI models 110c can ingest the entire data, where the (multiple) AI models 110c identify the corresponding voxels, analyze the first digital map, and associate the (multiple) images of each voxel with its corresponding occupancy status.
[0087] Using the ground truth, the (multiple) AI models 110c can be trained to analyze the visual elements of each voxel and associate them with whether the voxel is occupied by a mass. Thus, the AI model 110c can retrieve the occupancy status of each voxel (using the first dataset) and use this information as the ground truth. The (multiple) AI models 110c can also retrieve the visual attributes of the same voxels using the second dataset.
[0088] In some embodiments, the analysis server 110a can use a supervised training method. For example, using the ground truth and the received visual data, the (multiple) AI models 110c can train themselves such that it can predict the occupancy status for a voxel using only the image of that voxel. As a result, when trained, the (multiple) AI models 110c can receive the camera feed, analyze the camera feed, and determine the occupancy status for each voxel within the camera feed (without the need to use radar).
[0089] The analysis server 110a can feed a series of training datasets into the (multiple) AI models 110c and obtain a set of predicted outputs (e.g., predicted occupancy status). The analysis server 110a can then compare the predicted data with the ground truth data to determine the differences and train the (multiple) AI models 110c by adjusting the internal weights and parameters of the AI models 110c proportional to the determined differences according to a loss function. The analysis server 110a can train the (multiple) AI models 110c in a similar manner until the predictions of the trained AI models 110c are accurate to a certain threshold (e.g., recall or precision).
[0090] Additionally or alternatively, the analytics server 110a may use an unsupervised method in which the training data set is not labeled. Since labeling the data within the training data set can be time-consuming and may require excessive computing power, the analytics server 110a may utilize unsupervised training techniques to train the AI model 110c.
[0091] After the AI model 110c is trained, the ego 140 may use it to predict occupancy data for the surroundings of one or more egos 140. For example, the AI model(s) 110c may divide the surroundings of the ego into different voxels and predict the occupancy status for each voxel. In some embodiments, the AI model(s) 110c (or the analytics server 110a that uses the data predicted by the AI model 110c) may generate an occupancy map or occupancy network representing the surroundings of one or more egos 140 at any given time.
[0092] In another example of how the AI model(s) 110c may be used, after training the AI model(s) 110c, the analytics server 110a (or the local chip of the ego 140) may collect data from the ego (e.g., one or more egos among the egos 140) to predict an occupancy data set for one or more egos 140. This example describes how the AI model(s) 110c may be used to predict occupancy data for one or more egos 140 in real time or near real time. This configuration may have a processor that executes the AI model, such as the analytics server 110a. However, one or more actions may be performed locally, for example, via a chip located within one or more egos 140. In operation, the AI model(s) 110c may be executed locally via the ego 140 such that the results can be used for autonomous navigation of the ego itself.
[0093] The processor may input image data of the space around the ego object 140 into the AI model 110c using a camera of the ego object 140. The processor may collect and / or analyze data received from various cameras (e.g., outward-facing cameras) of one or more egos 140. In another example, the processor may collect and aggregate footage recorded by one or more cameras of the ego 140. The processor may then send the footage to the AI model(s) 110c trained using the methods discussed herein.
[0094] The processor may predict the occupancy attributes of multiple voxels by executing the AI model 110c. The AI model(s) 110c may use the methods discussed herein and use the received image data to predict the occupancy status for different voxels surrounding one or more egos 140.
[0095] The processor can generate a data set based on multiple voxels and their corresponding occupancy attributes. The analysis server 110a can generate a data set including its occupancy status according to the corresponding coordinate values of different voxels. This data set can be a queryable data set that can be used to send the predicted occupancy status to different software modules.
[0096] In operation, one or more aut bodies 140 can collect image data from their cameras and send the image data to the processor (locally placed on one or more aut bodies 140) and / or the analysis server 110a, as depicted by the data stream 172. The processor can then execute the AI model(s) 110c to predict the occupancy data for one or more aut bodies 140. If the prediction is executed by the analysis server 110a, then the occupancy data can be sent to one or more aut bodies 140 using the data stream 174. If the processor is locally placed within one or more aut bodies 140, then the occupancy data is sent to the aut computing device 141 ( Figure 1A not shown in the figure).
[0097] Using the methods discussed herein, the training of the AI model(s) 110c can be performed such that the execution of the AI model(s) 110c can be locally executed (at inference time) on any aut body within the aut body 140. The collected data (e.g., navigation data collected during the navigation of the aut body 140, such as image data of the journey) can then be fed back into the AI model(s) 110c such that additional data can improve the AI model(s) 110c.
[0098] FIG. 2 illustrates a flowchart of a method 200 executed in an AI-enabled visual data analysis system according to an embodiment. The method 200 can include step 210 to step 270. However, other embodiments can include additional or alternative steps, or one or more steps can be omitted. The method 200 is executed by an analysis server (e.g., a computer similar to the analysis server 110a). However, one or more steps of the method 200 can be executed by any number of computing devices (e.g., the processors of the aut body 140 and / or the aut computing device 141) operating in the Figures 1A - 1C distributed computing system described herein. For example, one or more computing devices of the aut body can locally execute some or all of the steps described in FIG. 2.
[0099] FIG. 2 illustrates a model architecture of how image input can be ingested from an aut body (step 210) and analyzed to predict a queryable output (step 270). Using the methods and systems discussed herein, the analysis server can ingest only image data (e.g., a camera feed from the surrounding environment of the aut body) to generate a queryable output. Thus, the methods and systems discussed herein can operate without receiving any data from radar, LiDAR, etc.
[0100] The queryable output (generated in step 270) can be used for various purposes. In one example, the queryable output can be used for an autonomous driving module, where various navigation decisions can be made based on whether the spatial voxels surrounding the ego are predicted to be occupied. In another example, using the queryable output, an analysis server can generate a digital map depicting the occupancy state of the surrounding environment of the ego. For example, the analysis server can generate a three-dimensional (3D) geometric representation of the surrounding environment of the ego. For example, the digital map can be displayed on the ego's computing device.
[0101] As used herein, a voxel can refer to a volume pixel and can be the 3D equivalent of a pixel in 2D. Thus, a voxel can represent a defined point in a 3D grid within the volumetric space or environment surrounding the ego (e.g., around). In some embodiments, the space surrounding the ego can be divided into different voxels, referred to as a voxel grid. As used herein, a voxel grid can refer to a set of cubes stacked (or arranged) together to represent objects in the space surrounding the ego. Each voxel can contain information about a specific location within the surrounding space of the ego. Using the methods and systems discussed herein, the occupancy of each voxel can be evaluated. For example, an analysis server (using the AI model discussed herein) can determine whether each voxel is occupied by an object with mass. Voxel predictions can be aggregated into a dataset referred to herein as queryable results. Using the queryable results, voxel information can be queried by a processor or downstream software module (e.g., autonomous driving software / processor) to identify occupancy data of the surrounding environment of the ego.
[0102] In some embodiments, if any part of a voxel is occupied, the voxel can be designated as occupied. Thus, in some embodiments, each voxel can include a binary designation of 0 (unoccupied) or 1 (occupied). Alternatively, in some embodiments, the AI model can also predict detailed occupancy data inside / within a specific voxel. For example, a voxel with a binary value of 1 (occupied) can be further analyzed at a finer granularity level to determine the occupancy of each point within the voxel. For example, an object can be curved. While some voxels (associated with the object) are fully occupied, some other voxels can be partially occupied. These voxels can be divided into smaller voxels such that some of the smaller voxels are unoccupied. As described herein, this method can be used to identify the shape of an object.
[0103] Method 200 starts at step 210, where image data is received from one or more cameras of the self. Method 200 visually illustrates how an AI model (trained using the methods discussed herein) can ingest the image data and generate a queryable output that can indicate the volume occupancy of various voxels within the surrounding environment of the self. The image data can refer to any data received from one or more images of the self.
[0104] The captured image data can then be characterized (step 220). An image featureizer or various characterization algorithms can be used to extract relevant and meaningful features from the received image data. Using the image featureizer, the image data can be transformed into a data representation that captures important information about the image content. This allows for more efficient analysis of the image data.
[0105] In some embodiments, the AI model can perform the characterization discussed herein. In some other embodiments, a convolutional neural network can be used to characterize the image data. In one non-limiting example, as depicted, a RegNet (Regularized Neural Network) can be used to transform the data into a BiFPN (Bidirectional Feature Pyramid Network). However, other protocols can also be used. In some other embodiments, a transformer can be used to characterize the image data.
[0106] After the image data is encoded / characterized, a transformer can be used to change the image data from a 2D image to a 3D image (step 230). As discussed herein, in an example configuration, there can be eight different cameras communicating with the self. As a result, the image data can include eight different camera feeds (one feed corresponding to each camera or other sensor), and can include overlapping views. The transformer can aggregate these individual camera feeds and generate one or more 3D representations using the received camera feeds.
[0107] The transducer can ingest three separate inputs: an image key, an image value, and a 3D query. The image key and image value can refer to attributes associated with 2D image data received from the ego. For example, these values can be output via image characterization (step 220). The transducer can also use an image query from the 3D space. The depicted spatial attention module can use the 3D query to analyze the 2D image key and image value. As depicted, the BiFPN generated in step 220 can be aggregated into a multi-camera query embedding and can be used to perform a 3D spatial query. In some embodiments, each voxel can have its own query. Using the 3D spatial query, the analysis server can identify regions within the 2D characterized image corresponding to specific portions of the 3D representation. The identified regions within the characterized image can then be analyzed to transform the multi-camera image data into a 3D representation for each voxel, which can result in a 3D representation of the ego's surrounding environment. Thus, the depicted spatial attention module can output a single 3D vector space representing the ego's surrounding environment. Effectively, this moves all the image data generated by all camera feeds into a top-down spatial or 3D space representation of the ego's surrounding environment.
[0108] Steps 210 through 230 can be performed for each video frame received from each camera of the ego. For example, at each timestamp, steps 210 through 230 can be performed on eight different images received from eight different cameras of the ego. As a result, at each timestamp, method 200 can produce a 3D space representation of the eight images. In step 240, method 200 can fuse the 3D spaces (for different timestamps) together. This fusion can be performed based on the timestamps of each set of images. For example, the 3D space representations can be fused based on their corresponding timestamps (e.g., in a sequential manner).
[0109] As depicted, the 3D space representation at timestamp t can be fused with the 3D space representations of the ego's surrounding environment at t-1, t-2, and t-3. As a result, the output can have both spatial and temporal information. This concept is depicted in Figure 2 as spatio-temporal features.
[0110] Then, the spatio-temporal features can be transformed into different voxels using deconvolution (step 250). As discussed herein, various data points are characterized and fused together. In this step 250, method 200 can perform various mathematical operations to reverse the process such that the fused data can be transformed back into different voxels. As used herein, deconvolution can refer to a mathematical operation used to reverse the effects of convolution.
[0111] After applying deconvolution to the image data (which has been characterized, transformed, and fused), method 200 can then apply the various trained AI modeling techniques discussed herein (e.g., FIGS. 3 to 4) to generate a volumetric output (step 260). The volumetric output can include binary data for different voxels that indicates whether a particular voxel is occupied by an object with mass. Specifically, the volumetric output can include occupancy data (including binary data) indicating whether a voxel is occupied and / or occupancy flow data indicating the rate at which the voxel is moving (if any) (using temporal alignment to calculate the rate).
[0112] The volumetric output can also include shape information (the shape of the mass of the occupied voxels). In some embodiments, the size of each voxel can be predetermined, but the size can be corrected to produce a more fine-grained result. For example, the default size of different voxels can be 33 centimeters (per vertex). While this size is generally acceptable for voxels, the results can be improved by reducing the size of the voxels. For example, a 33 cm voxel may be appropriate if the voxel is detected outside the driving surface of the ego. However, the analysis server can reduce the size of the voxels that are occupied and within a threshold distance from the ego and / or the ego driving surface (e.g., reduce to 10 cm). When the voxel occupancy data is identified, a regression model can be executed to identify the shape of the group of voxels. For example, a 33 cm voxel (which belongs to a curb) may be half occupied (e.g., only 16 cm of the voxel is occupied). The analysis server can use regression to determine how much of the voxel is occupied.
[0113] Additionally or alternatively, the analysis server can decode the sub-voxel values to identify the shape of the sub-voxels (inside the occupied voxels). For example, if a voxel is half occupied, the analysis server can define a set of sub-voxels and use the methods discussed herein to identify the volumetric output for the sub-voxels. When the sub-voxels are aggregated (back into the original voxel), the analysis server can determine the shape for the voxel. For example, each voxel can have eight vertices. In some embodiments, each vertex can be analyzed individually and have its embedding. As a result, any point within each vertex of the voxel can be queried individually. Thus, in this "continuous resolution" method, the analysis server can not define the size for the sub-voxels. In some embodiments, the analysis server can use a multivariable interpolation (e.g., trilinear interpolation) protocol to estimate the occupancy status of each sub-voxel and / or any point within each vertex.
[0114] The volumetric output may also include 3D semantic data indicative of an object occupying a voxel (or group of voxels). The 3D semantics may indicate whether a voxel and / or a group of nearby voxels are occupied by a car, a street curb, a building, or other object. The 3D semantics may also indicate whether a voxel is occupied by a static mass or a moving mass. Various temporal attributes of the voxels may be used to identify the 3D semantic data. For example, if a group of voxels is identified as being occupied by a mass, the collective shape of the voxels may indicate that the voxels belong to a vehicle. If, at a previous timestamp, the identified group of voxels (now known to be a vehicle) was identified as being in motion, then the group of voxels may have 3D semantics indicative of the group of voxels belonging to a moving vehicle. In another example, if a group of voxels is identified as having a shape corresponding to a curb and is not identified as having any motion, then the group of voxels may have 3D semantics indicative of a static curb.
[0115] In some embodiments, certain shapes or 3D semantics may be prioritized. For example, certain objects, such as other vehicles on the road or objects associated with the driving surface (e.g., curbs indicating the outer bounds of the road), may be analyzed thoroughly. In contrast, details of static objects (such as buildings far from the self-driving surface nearby) may not be analyzed as thoroughly as moving vehicles near the self. In some embodiments, certain objects having a particular size or shape may be ignored. For example, road debris may not be analyzed as much as moving vehicles near the self.
[0116] In some embodiments, method 200 may not need to perform object-level detection. For example, the self must navigate around voxels identified as static and occupied in front of the self, regardless of whether the voxels belong to another vehicle, a pedestrian, or a traffic sign. Thus, the occupancy information may be object-agnostic. In some embodiments, an object detection model may be executed separately (e.g., in parallel), which may detect objects corresponding to various groups of voxels.
[0117] In step 270, method 200 may generate a queryable data set that allows other software modules to query the occupancy status of different voxels. For example, a software module may send coordinate values (X, Y, and Z axes) of the self's surrounding environment and may receive any of the four classes of occupancy data (e.g., volumetric output) generated using method 200. The queryable data set may be used to generate an occupancy map (e.g., Figures 3A - 3B ) or may be used to make autonomous navigation decisions for the self.
[0118] Additionally or alternatively, an analysis server may generate a map corresponding to the predicted occupancy status of different voxels. In a non-limiting example, the analysis server may use a multi-view 3D reconstruction protocol to visualize each voxel and its occupancy status. Figures 3A - 3BNon-limiting examples of maps or occupancy maps are presented (e.g., simulation 350). In some embodiments, simulation 350 may be displayed on the user interface of the vehicle itself. Simulation 350 may illustrate Figure 3A the camera feed 300 depicted in FIG. The camera feed 300 represents image data (either in real-time or near real-time) received from eight different cameras of the vehicle itself. Specifically, the camera feed 300 may include camera feeds 310a to 310c received from three different front cameras of the vehicle itself; camera feeds 320a to 320b received from two different right-facing cameras of the vehicle itself; camera feeds 330a to 330b received from two different left-facing cameras of the vehicle itself; and a camera feed 340 received from the rear camera of the vehicle itself.
[0119] Using the methods discussed herein, the analysis server may analyze the camera feed 300, divide the space around the vehicle itself into voxels, and generate the simulation 350 (as Figure 3B depicted) as a graphical representation of the surrounding environment of the vehicle itself. The simulation 350 may include a simulated vehicle itself (360) and its surrounding voxels. For example, the simulation 350 may include graphical indicators of different qualities for occupying different voxels around the simulated vehicle itself 360. For example, the simulation 350 may include simulation qualities 370a to 370c.
[0120] Each of the simulation qualities 370a to 370c may represent an object depicted in the camera feed 300. For example, simulation quality 370a corresponds to quality 380a (a vehicle); simulation quality 370b corresponds to quality 380b (a vehicle); and simulation quality 370c may correspond to quality 380c (a building near the road). As depicted, each simulation quality includes various voxels. Moreover, the voxels depicted within the simulation 350 may have different graphical / visual characteristics corresponding to their volume output (e.g., occupancy data). For example, simulation quality 370c (e.g., a building) may have a first color indicating that it has been identified as static. Similarly, simulation quality 370b (e.g., a vehicle) may have a second color indicating that it is a parked or stationary vehicle. In contrast, simulation quality 370a (e.g., another vehicle) may have a third color and / or other visual characteristics indicating that it is predicted to be moving.
[0121] Additionally or alternatively, the analysis server may send the generated map to a downstream software application or another server. The prediction results may be further analyzed and used in various models and / or algorithms to perform various actions. For example, a software model or processor associated with the autonomous navigation system of the vehicle itself may receive the occupancy data predicted by the trained AI model, and navigation decisions may be made based on this data.
[0122] Figure 2B FIG. illustrates a flowchart of method 201 executed in an AI-enabled visual data analysis system according to an embodiment. Method 201 may include steps 210 to 290. However, other embodiments may include additional or alternative steps, or one or more steps may be entirely omitted. Method 201 is described as being executed by an analysis server (e.g., a computer similar to analysis server 110a). However, one or more steps of method 201 may be executed by any number of computing devices (e.g., processors of the ego 140 and / or the ego computing device 141) operating in a distributed computing system described in Figures 1A to 1C For example, one or more computing devices of the ego may execute some or all of the steps described in Figure 2B locally.
[0123] Using method 201, the AI model may be configured to generate more than an orthographic projection of the ego's surrounding environment. The AI model may only require image data to predict various surfaces near the ego and their corresponding surface properties. As depicted, method 201 includes a volume output (step 260) indicating surface properties of different volumes surrounding the ego.
[0124] As depicted, steps 210 to 250 may be similar in Figure 2A and Figure 2B However, method 201 may include additional steps that allow the AI model to predict surface properties surrounding the ego. Specifically, method 201 may include additional step 280 and additional step 290, in which ground truth is generated in additional step 280, and a 3D representation (e.g., model rendering) of the ego's surrounding environment is generated using the data predicted by executing methods 200 and 201 in additional step 290.
[0125] Method 201 allows the AI model to predict the 3D properties of various surfaces in the environment surrounding the ego, rather than generating an orthographic view of the ego's surrounding environment. Using method 201, it may no longer be required to localize the ego to achieve autonomous navigation. Compared with conventional methods, method 201 may allow the AI model to receive image data in real time or near real time (on the fly) and analyze various surfaces near the ego. As a result, the ego may be able to navigate itself without executing a localization protocol.
[0126] Images received from the vehicle's camera can include a 2D representation of the vehicle's surrounding environment. This representation is sometimes referred to as a 2D or planar lattice. The planar lattice can be transformed into different nodes with specific X-axis and Y-axis coordinate values. Using method 201, the AI model can predict the Z-axis coordinate value for each node within the planar lattice. Specifically, using method 201, the AI model can predict the feature vector for each point within the image data with different X-axis and Y-axis coordinate values. As used herein, the Z-axis coordinate value for each point or node can represent the elevation of that point relative to the plane with an elevation of 0 in the world.
[0127] In addition to predicting the elevation for each node, the AI model can also determine the class (surface property) for each node. For example, the AI model can determine whether the surface is navigable. Additionally, the AI model can determine the properties of each surface material (e.g., grass, dirt, asphalt, or concrete). Additionally, the AI model can determine whether the surface is a road or a sidewalk. Moreover, the AI model can determine the paint lines associated with different surfaces, allowing the AI model to infer whether the surface is a road surface or a curb.
[0128] Using the feature vector for each node, the AI model can generate a grid representation corresponding to the vehicle's surrounding environment. As used herein, a grid can refer to a series of interconnected nodes representing the vehicle's surrounding environment, where each node includes X, Y, and Z-axis coordinate values. Each node can also include data indicating its properties and class (e.g., whether the node within the surface is navigable, what the node is identified as, and what material the node is predicted to be).
[0129] At step 280, the AI model can generate the ground truth to be ingested by the deconvolution step (250). The vehicle's sensors can generate a point cloud of the vehicle's surrounding environment. The point cloud can include many points that represent 3D coordinate data associated with the vehicle's surrounding environment at different timestamps. In a non-limiting example, LiDAR data can be received from the vehicle, and the point cloud can represent the received LiDAR data points. The vehicle's camera can also send images of the vehicle's surrounding environment at different timestamps. The analysis server can use the different timestamps to identify the image data corresponding to different points within the point cloud. The analysis server can then project the data associated with the points within the image data, thereby identifying the image regions (with pixel sets) corresponding to one or more points within the point cloud.
[0130] Additionally or alternatively, the analysis server can use the image data captured by one or more cameras of the vehicle instead of using LiDAR data. Using the captured image data, a point cloud can also be generated, for example, by triangulating the significant feature points detected within the image data.
[0131] The analysis server can also use an auxiliary AI model (such as a neural network), such as a semantic segmentation network, to analyze pixels within the image data. For example, a group of pixels can be analyzed by the semantic segmentation network. The semantic segmentation network can then determine one or more attributes of the pixel set. For example, using this paradigm, the analysis server can determine whether the pixel group corresponds to a tree, sky, curb, or road. In some embodiments, the semantic segmentation network can determine whether the surface is navigable. In some embodiments, the semantic segmentation network can determine the material associated with the pixel set. For example, the semantic segmentation model can determine whether the pixels within the image data correspond to dirt, water, concrete, or asphalt. In some other embodiments, the semantic segmentation network can identify whether the surface is painted; and if so, identify the color of the paint. Essentially, the semantics of each 3D point can be identified using the semantic segmentation network.
[0132] Using the semantic segmentation model, the analysis server can filter these points and cluster them into their respective categories (such as pixels representing a sidewalk, pixels representing a dirt road or an asphalt road). The analysis server can analyze different image data at different timestamps.
[0133] After executing the semantic segmentation model, the point cloud can be segmented based on the image data corresponding to the point cloud and / or its attributes (as predicted by the semantic segmentation model). As a result, the points associated with a specific surface and the image data associated with the same surface around the ego can be identified and isolated. Then, the analysis server can fit a mesh surface to the isolated data points. This may be because the AI model can perform more efficiently using a smooth surface, which may be more indicative of reality. In fact, mesh fitting can denoise the data and provide a more realistic representation of the surface around the ego. The fitted surface can be used as ground truth for training purposes.
[0134] The AI model can be trained using the image data received from the ego and the ground truth, such that the AI model can analyze the image data received from the ego without the need for any sensor data when being trained. Effectively, using this specific training paradigm, the AI model can associate how to represent pixels associated with a specific surface with specific attributes (such as an uphill dirt road with white paint). Therefore, the AI model (at inference time) can utilize only the image data and does not require other sensor data.
[0135] Once trained, the AI model can be configured to ingest image data and generate a lattice with various nodes, where each node has a corresponding feature vector that includes X and Y axis coordinate values (identified via the image data) and a Z axis coordinate value predicted by the AI model. The AI model can also predict one or more attributes for each node. For example, a particular node can include a feature vector that includes a predicted elevation (e.g., 1 meter above itself). Additionally, the AI model can predict that the node is a road node (since the corresponding pixel is predicted to be a driving surface), that there is paint on the node, and that the paint is yellow.
[0136] In some embodiments, it may be necessary to adjust the coordinate values (e.g., the Z axis coordinate indicating the elevation of the node) because the self itself has changed position and the Z coordinate may not be corrected. For example, when the self navigates in the terrain, it can send the coordinates of its surroundings. However, the coordinates can be related to the sensors of the self or the self itself. Therefore, if the self changes its vertical position (e.g., if the self is driving over a speed bump or a pothole), the coordinates received from the self may also change. However, the coordinates may change because they are related to the self coordinates. For example, if the self is driving on a flat surface, the same location can have different coordinate values compared to when the same self is driving over a speed bump. Therefore, in some embodiments, the coordinates received from the self can be corrected before they can be used to train the AI model.
[0137] To correct this issue, the coordinate values can be aligned with the surface of the self's surroundings itself (rather than the self). In this way, the noisy or incorrect data received due to the movement of the self can be smoothed. Essentially, the surface is processed independently and the coordinate values are calculated (and ultimately predicted) based on the surface rather than the self.
[0138] In some embodiments, method 201 can be combined with method 200 (the occupancy detection paradigm) to identify objects located within an elevated surface. For example, an object can be detected on a surface that has been identified as having a higher or lower elevation than the self (e.g., a traffic cone is identified on a hill in front of the self). In this example, the AI model can use method 201 to determine the attributes of the hill in front of the self. Then, the attributes of the cone itself can be identified as if the cone were located on a flat surface (e.g., the height of the hill at that particular location can be subtracted). Then, the AI model can use method 200 to identify the voxels associated with the cone, thereby identifying the dimensions of the cone. The dimensions are then added to the hill identified using method 201. Thus, the AI model can bifurcate the identification of the surface and the object and then combine them to truly understand / predict the position and attributes of different objects located on different surfaces.
[0139] Detecting the fork into two different protocols (Method 200 and Method 201) also allows the ego to detect the occupancy status of different voxels when the ego exceeds the ego occupancy detection range in different voxels. For example, the ego can have a vertical occupancy detection range of -3 meters to +3 meters. This indicates that if different voxels are within the -3-meter to +3-meter elevation range of the ego, the ego can identify their occupancy status. The occupancy detection range may not mean that the camera cannot record the footage of objects outside the range; in contrast, this may mean that the AI model cannot identify objects outside the detection range.
[0140] In these embodiments, the ego may not be able to predict any objects on a steep slope outside the ego occupancy detection range (e.g., a traffic cone on a downhill slope 4 meters below the ego's elevation). Using the methods discussed herein, the ego can first determine that the driving surface is 4 meters lower than the ego. Then, the AI model can separately determine the attributes of the voxels in the occupied space (traffic cone) and subtract the height of the hill from the height of the traffic cone. Effectively, in addition to providing more consistent results, Method 201 can also be used to extend the ego's occupancy detection range (used in Method 200).
[0141] Using Method 201, the AI model can receive image data from the ego's camera and transform the image data into a mesh representation of the ego's surrounding environment. Thus, the image received from the camera can be transformed into a 3D description of the surface around the ego, such as the driving surface.
[0142] In some embodiments, the analysis server can use neural radiance field (NeRF) technology to reconstruct the rendering of the ego's surrounding environment (Step 290). In some embodiments, the analysis server can use the captured image data to generate a map indicating various surfaces around the ego. This map can correspond to the predicted surface and its predicted attributes. In a non-limiting example, the analysis server can use a multi-view 3D reconstruction protocol to visualize each voxel and its surface state / attributes. Figures 4A to 4C A non-limiting example of a map or surface map is presented (e.g., Simulation 400).
[0143] In some embodiments, the simulation 400 can be displayed on the user interface of the vehicle. The simulation 400 can illustrate how the camera feeds 410 can be analyzed to generate a graphical representation of the vehicle's surrounding environment. The camera feeds 410 represent image data received from five different cameras of the vehicle (either in real-time or near real-time). Each camera feed can be received from a different camera and can depict a different view / angle of the vehicle's surrounding environment. Specifically, the camera feeds 410 represent image data received from eight different cameras of the vehicle (either in real-time or near real-time). The camera feeds 410 can include: camera feeds 410a to 410c received from three different front-facing cameras of the vehicle; camera feeds 410d to 410e received from two different right-facing cameras of the vehicle; camera feeds 410f to 410g received from two different left-facing cameras of the vehicle; and a camera feed 410h received from the rear camera of the vehicle.
[0144] Using the methods discussed herein, the analysis server can analyze the camera feeds 410 to 450 and generate the simulation 400, which is a graphical representation of the surfaces surrounding the vehicle. The simulation 400 can include a simulated vehicle (420) and the surfaces surrounding it. For example, the simulation 400 can visually identify the surfaces 430 and 440 using visual attributes (such as different colors (or other visual methods, such as shadow patterns)) to indicate that the AI model has identified the surfaces 430 and 440 as navigable surfaces. The simulation 400 can also include surfaces 450 and 492, which are visually different from the surfaces 430 and 440 (e.g., different colors or different shadow patterns), because the surfaces 450 to 460 have been identified as curbs, and curbs are not navigable surfaces.
[0145] The different surfaces depicted within the simulation 400 can visually replicate the predicted elevation (e.g., the Z coordinate values predicted using an AI model). For example, the surface 430 (in front of the vehicle) visually indicates that the road in front of the vehicle is a downhill road. In contrast, the surface 440 is visually depicted as an uphill road.
[0146] Now referring to Figure 4C , the simulation 410 depicts the same surfaces depicted within the simulation 400. Specifically, the simulation 401 includes the simulated vehicle 420 driving on the surface 430 (the same surface 430 depicted within the simulation 400) and the surface 440 to the right of the simulated vehicle 420.
[0147] Additionally or alternatively, the analysis server may send the generated map to a downstream software application or another server. The prediction results may be further analyzed and used in various models and / or algorithms to perform various actions. For example, a software model or processor associated with the autonomous navigation of the ego vehicle may receive the occupancy data predicted by the trained AI model, and various navigation decisions may be made accordingly.
[0148] Figure 5 FIG. 500 is a flowchart of a method 500 performed in an AI-enabled visual data analysis system according to an embodiment. Method 500 may include steps 510 to 530. However, other embodiments may include additional or alternative steps, or one or more steps may be entirely omitted. Method 500 is described as being performed by an analysis server (e.g., a computer similar to analysis server 110a). However, one or more steps of method 500 may be performed by any number of computing devices (e.g., processors of ego vehicle 140 and / or ego computing device 141) operating in a distributed computing system as described in Figure 1A and 1B For example, one or more computing devices may perform some or all of the steps described in Figure 5 locally. For example, a chip placed within the ego vehicle may perform method 500.
[0149] At step 510, the analysis server may input image data of the space around the ego object into an artificial intelligence model using a camera of the ego object. The analysis server may collect and / or analyze data received from various cameras of the ego vehicle (e.g., outward-facing cameras). In another example, the analysis server may collect and aggregate footage recorded by one or more cameras of the ego vehicle. The analysis server may then send the footage to an AI model trained using the methods discussed herein.
[0150] At step 520, the analysis server may predict surface properties of one or more surfaces of the space around the ego object by executing the artificial intelligence model. The AI model may use the methods discussed herein to identify one or more surfaces around the ego vehicle. The AI model may also use the data received at step 710 to predict one or more surface properties (e.g., category, material, elevation) for the one or more surfaces.
[0151] At step 530, the analysis server may generate a data set based on the one or more surfaces and their corresponding surface properties. The analysis server may generate a data set that includes the one or more surfaces and their corresponding surface properties. The data set may be a queryable data set that can be used to send the predicted surface data occupancy status to different software modules.
[0152] The various illustrative logical blocks, modules, circuits, and algorithmic steps described in connection with the embodiments disclosed herein can be implemented as electronic hardware, computer software, or combinations of both. To clearly illustrate this interchangeability of hardware and software, various illustrative components, blocks, modules, circuits, and steps have been described generally in terms of their functionality. Whether such functionality is implemented as hardware or software depends upon the particular application and design constraints imposed on the overall system. Skilled artisans may implement the described functionality in varying ways for each particular application, but such implementation decisions should not be interpreted as causing a departure from the scope of the present disclosure or the claims.
[0153] Embodiments implemented in computer software can be implemented in software, firmware, middleware, microcode, hardware description language, or any combination thereof. A code segment or machine-executable instruction can represent a procedure, a function, a subprogram, a program, a routine, a subroutine, a module, a software package, a class, or any combination of instructions, data structures, or program statements. By passing and / or receiving information, data, arguments, parameters, or memory contents, a code segment can be coupled to another code segment or a hardware circuit. Information, arguments, parameters, data, etc. can be passed, forwarded, or transmitted via any suitable means, including memory sharing, message passing, token passing, network transmission, etc.
[0154] The actual software code or specialized control hardware used to implement these systems and methods does not limit the claimed features or the present disclosure. Accordingly, the operation and behavior of the systems and methods are described without reference to the specific software code, it being understood that the software and control hardware can be designed to implement the systems and methods based on the description herein.
[0155] When implemented in software, functions can be stored as one or more instructions or code on a non-transitory computer-readable or processor-readable storage medium. The steps of the methods or algorithms disclosed herein can be implemented in processor-executable software modules that may reside on a computer-readable or processor-readable storage medium. Non-transitory computer-readable or processor-readable media include both computer storage media and tangible storage media, and tangible storage media facilitate the transfer of a computer program from one place to another. A non-transitory processor-readable storage medium can be any available medium that is accessible by a computer. By way of example and not limitation, such non-transitory processor-readable media can include RAM, ROM, EEPROM, CD-ROM or other optical disk storage, magnetic disk storage or other magnetic storage devices, or any other tangible storage medium that can be used to store the desired program code in the form of instructions or data structures and that can be accessed by a computer or a processor. As used herein, disk and optical disk include compact disk (CD), laser disk, optical disk, digital versatile disk (DVD), Blu-ray disk, and floppy disk, where a "disk" typically reproduces data magnetically, while an "optical disk" reproduces data optically with a laser. Combinations of the above should also be included within the scope of computer-readable media. Additionally, the operations of a method or algorithm can reside as code and / or instructions in one or any combination or collection on a non-transitory processor-readable medium and / or a computer-readable medium that can be incorporated into a computer program product.
[0156] The foregoing description of the disclosed embodiments is provided to enable a person skilled in the art to make or use the embodiments and variations described herein. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the principles defined herein can be applied to other embodiments without departing from the spirit or scope of the subject matter disclosed herein. Thus, the disclosure is not intended to be limited to the embodiments shown herein but is to be accorded the widest scope consistent with the following claims and the principles and novel features disclosed herein.
[0157] Although various aspects and embodiments have been disclosed, other aspects and embodiments are also contemplated. The various aspects and embodiments disclosed are for illustrative purposes and are not intended to be limiting, and the true scope and spirit are indicated by the following claims.
Claims
1. A method, comprising: inputting, by a processor using one or more cameras of a self - object, image data of a space around the self - object into an artificial intelligence model; predicting, by the processor executing the artificial intelligence model, surface properties of one or more surfaces of the space around the self - object; and generating, by the processor, a data set based on the one or more surfaces and corresponding surface properties of the one or more surfaces.
2. The method according to claim 1, further comprising: generating, by the processor, an output representing the space around the self - object and illustrating the one or more surfaces and corresponding surface properties of the one or more surfaces.
3. The method according to claim 2, wherein the one or more surfaces are visually different according to their respective surface properties.
4. The method according to claim 2, further comprising: displaying, by the processor, the output on a screen associated with the self - object.
5. The method according to claim 1, wherein the data set is a queryable data set configured to send the surface properties to an autonomous driving protocol of the self - object.
6. The method according to claim 1, wherein the artificial intelligence model is trained using sensor properties of the one or more surfaces.
7. The method according to claim 1, wherein the self - object is an autonomous vehicle that executes a driving protocol based on the data set.
8. The method according to claim 1, wherein the surface property corresponds to elevation.
9. The method according to claim 1, wherein the surface property indicates whether the one or more surfaces are navigable surfaces.
10. The method according to claim 1, wherein the surface property indicates a material associated with the one or more surfaces.
11. The method according to claim 1, wherein the surface property indicates whether the one or more surfaces correspond to one or more curbs.
12. A self - object, comprising: one or more cameras; a processor; a non - transient computer - readable medium containing an artificial intelligence model configured to be executed by the first processor, wherein the processor is configured to: input, using the one or more cameras of the self - object, image data of a space around the self - object into the artificial intelligence model; predict, by executing the artificial intelligence model, surface properties of one or more surfaces of the space around the self - object; and generate, based on the one or more surfaces and corresponding surface properties of the one or more surfaces, a data set.
13. The self - object according to claim 12, wherein the instructions are further capable of causing the processor to generate an output representing the space around the self - object and illustrating the one or more surfaces and corresponding surface properties of the one or more surfaces.
14. The self - object according to claim 13, wherein the one or more surfaces are visually different according to their respective surface properties.
15. The self-object according to claim 13, wherein the instructions are further capable of causing the processor to display the output on a screen associated with the self-object.
16. The self-object according to claim 12, wherein the data set is a queryable data set configured to send the surface property to an autonomous driving protocol of the self-object.
17. The self-object according to claim 12, wherein the artificial intelligence model is trained using the sensor properties of the one or more surfaces.
18. The self-object according to claim 12, wherein the self-object is an autonomous vehicle that executes a driving protocol based on the data set.
19. The self-object according to claim 12, wherein the surface property corresponds to altitude.
20. The self-object according to claim 12, wherein the surface property indicates whether the one or more surfaces are navigable surfaces.