Simulation of viewpoint capture from environment rendered with terrestrial live heuristic
By generating a three-dimensional model of the physical environment in a virtual environment and capturing the simulated viewpoint, the problem of difficulty in obtaining navigation training data in severe weather is solved, and the efficiency and accuracy of navigation training are improved.
Patent Information
- Application Number
- CN202380076829.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-09-30
- Filing Date
- 2023-09-29
- Publication Date
- 2025-06-17
AI Technical Summary
Existing mobile systems require a large amount of real-world input when learning to accurately navigate their surroundings, but in severe weather, it becomes difficult to obtain these inputs, resulting in time-consuming and inefficient navigation training.
By generating a three-dimensional model corresponding to the physical environment of the real world and capturing the simulated viewpoints in the virtual environment, multiple virtual instances and simulated viewpoints are provided to replace inputs in the real world.
It can also provide a large amount of navigation training data in severe weather and other situations, improving the efficiency and accuracy of navigation training in mobile systems.
Smart Images

Figure CN120167069A_ABST
Abstract
Description
[0001] Cross - reference to related patent applications
[0002] This application claims priority to U.S. Provisional Application No. 63 / 377,954, filed on September 30, 2022, the entire content of which is incorporated herein by reference for all purposes. Technical Field
[0003] This implementation generally relates to computer rendering, including but not limited to simulations captured from a viewpoint in an environment rendered using ground truth heuristics. Background Art
[0004] Consumers are increasingly demanding mobile systems that can more accurately navigate their surroundings. Mobile systems can require a large amount of reliable input from the real world to learn to accurately navigate their surroundings. Training a mobile system to accurately navigate its surroundings is very time-consuming and requires a large amount of input corresponding to various situations. However, the amount of such input required by many mobile systems exceeds the amount of input that can be effectively obtained from the real world. For example, input for a specific real-world location in a tropical region cannot be obtained under snowy or blizzard conditions. Summary of the Invention
[0005] The technical solution is at least directed to capturing simulated sensor data from a virtual environment corresponding to a physical environment. For example, a computing system can generate a three-dimensional model corresponding to one or more physical aspects of one or more physical environments in the real world. For example, the computing system can obtain one or more input models, each input model corresponding to the boundary or shape of a zone (e.g., a roadway, a median) measured from the "ground truth" of a physical environment having similar characteristics. The computing system can generate a virtual environment corresponding to these aspects. The computing system can capture one or more viewpoints within the virtual environment corresponding to the movement of an autonomous object in the physical environment corresponding to the virtual environment. For example, the viewpoints can correspond to respective cameras, each camera having a specific orientation relative to the autonomous object and being movable within the virtual environment based on the position or orientation of the autonomous object. The computing system can also modify one or more aspects of the virtual environment to create a change in the virtual environment, each change in the virtual environment corresponding to the ground truth roadway of the physical environment but not yet or unable to be captured by measuring or imaging the physical environment. Thus, the computing system can create many virtual instances of the physical environment and create many simulated viewpoints corresponding to the many virtual instances of the physical environment, which can provide a technical improvement to generate at least a number of simulated viewpoints that exceeds the number available in the real world and the number that can be manually drawn. For example, the computing system can at least provide a technical improvement that allows an autonomous object that can navigate the real world more accurately. Thus, the embodiments herein provide a technical solution for simulating viewpoints captured from an environment rendered using ground truth heuristics.
[0006] In one embodiment, a system can include a memory and one or more processors. The system can retrieve a camera feed of an autonomous object navigating within a physical environment. The system can generate a three-dimensional (3D) model based on the camera feed according to one or more first environmental metrics, the 3D model including a first surface corresponding to one or more physical paths through the physical environment, the one or more first environmental metrics indicating the boundaries of the one or more physical paths. The system can generate one or more geometric two-dimensional (2D) objects on the first surface according to one or more second environmental metrics, the second environmental metrics indicating the one or more physical paths. The system can identify one or more viewpoints oriented to capture corresponding portions of the 3D model according to one or more viewpoint metrics of a camera indicating a physical object, the physical object being configured to move along the one or more physical paths. The system can render one or more simulated environment 2D images from the one or more corresponding portions of the 3D model of the physical environment, each of the one or more simulated environment 2D images corresponding to a respective one of the viewpoints. The system can train an artificial intelligence model according to the camera feed and the one or more simulated environment two-dimensional images.
[0007] The system can modify a portion of the first surface corresponding to a portion of a physical path among one or more physical paths according to one or more third environmental metrics, and the one or more third environmental metrics indicate the condition of the physical path.
[0008] The system can modify the topology of the portion of the first surface according to one or more third environmental metrics.
[0009] The system can modify the opacity of at least a portion of a geometric 2D object among geometric 2D objects located in the portion of the first surface according to one or more third environmental metrics.
[0010] The system can generate one or more 3D objects that satisfy a positioning heuristic at one or more corresponding positions in a second surface other than the first surface in the 3D model according to a positioning heuristic indicating the type of a physical object in the physical environment.
[0011] The system can include the type of a physical object corresponding to at least one of a geographical type, a climate type, or an architectural type.
[0012] The system can generate one or more 3D objects that satisfy an environmental heuristic at one or more corresponding positions in the 3D model according to an environmental heuristic indicating the atmospheric condition and the weather in the physical environment.
[0013] The system can divide a regional model into multiple regional segments according to a block heuristic indicating the amount of computing resources. The regional model indicating the physical region can include the physical environment, and each regional segment corresponds to a corresponding portion of the regional model. The 3D model corresponds to the regional segment among the regional segments.
[0014] The system can execute a first subset of the regional segments by a first computing resource. The system can execute a second subset of the regional segments by a second computing resource and simultaneously with the execution of the first computing resource.
[0015] In another embodiment, a method may include retrieving a camera feed of an autonomous object navigating within a physical environment. The method may include generating a three-dimensional (3D) model based on one or more first environmental metrics, the 3D model may include a first surface corresponding to one or more physical paths through the physical environment, and the one or more first environmental metrics indicate the boundaries of the one or more physical paths. The method may include generating one or more geometric two-dimensional (2D) objects on the first surface based on the camera feed according to one or more second environmental metrics, the second environmental metrics indicating the one or more physical paths. The method may include identifying one or more viewpoints oriented to capture corresponding portions of the 3D model according to one or more viewpoint metrics of a camera indicative of a physical object, the physical object being configured to move along one or more physical paths. The method may include rendering one or more 2D images from the one or more corresponding portions of the 3D model of the physical environment, each of the one or more 2D images corresponding to a respective one of the viewpoints. The method may include training an artificial intelligence model based on the camera feed and one or more simulated environmental 2D images.
[0016] The method may include modifying a portion of the first surface corresponding to a portion of a physical path among the physical paths according to one or more third environmental metrics, the one or more third environmental metrics indicating the condition of the physical path.
[0017] The method may include modifying the topology of a portion of the first surface according to one or more third environmental metrics.
[0018] The method may include modifying the opacity of at least a portion of the geometric 2D objects among the geometric 2D objects located on the portion of the first surface according to one or more third environmental metrics.
[0019] The method may include generating one or more 3D objects that satisfy a positioning heuristic at one or more corresponding locations in a second surface other than the first surface of the 3D model according to a positioning heuristic indicative of the type of a physical object in the physical environment.
[0020] The method may include the type of a physical object corresponding to at least one of a geographical type, a climate type, or an architectural type.
[0021] The method may include generating one or more 3D objects that satisfy an environmental heuristic at one or more corresponding locations in the 3D model according to an environmental heuristic indicative of an atmospheric condition, the weather in the physical environment.
[0022] The method may include segmenting a regional model into a plurality of regional segments according to a block heuristic that indicates a resource amount, the regional model indicating a physical region may include a physical environment, and each of the regional segments corresponds to a respective portion of the regional model, and the 3D model corresponds to a regional segment among the regional segments.
[0023] The method may include executing a first subset of the regional segments by a first computing resource. The method may include executing a second subset of the regional segments by a second computing resource and concurrently with the execution of the first computing resource.
[0024] In yet another embodiment, a non-transitory computer-readable medium may include one or more instructions stored thereon and executable by a processor. The processor may retrieve a camera feed of an ego object navigating within a physical environment. The processor may generate a three-dimensional (3D) model according to one or more first environmental metrics, the three-dimensional model may include a first surface corresponding to one or more physical paths through the physical environment, and the one or more first environmental metrics indicate boundaries of the one or more physical paths. The processor may generate one or more geometric two-dimensional (2D) objects on the first surface according to one or more second environmental metrics, the second environmental metrics indicating the one or more physical paths. The processor may identify one or more viewpoints oriented to capture corresponding portions of the 3D model according to one or more viewpoint metrics of a camera of a physical object configured to move along the one or more physical paths. The processor may render one or more simulated environment 2D images from one or more corresponding portions of the 3D model of the physical environment, each of the one or more simulated environment 2D images corresponding to a respective one of the viewpoints. The processor may train an artificial intelligence model according to the camera feed and the one or more simulated environment 2D images.
[0025] The computer-readable medium may include one or more instructions executable by a processor. The processor may modify a portion of the first surface corresponding to a portion of a physical path among the physical paths according to one or more third environmental metrics, the one or more third environmental metrics indicating a condition of the physical path. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] These and other aspects and features of the present implementation are depicted by way of example in the drawings discussed herein. The present implementation may be directed to, but is not limited to, the examples depicted in the drawings discussed herein. Accordingly, the present disclosure is not limited to any of the drawings or portions thereof depicted or referenced herein, or any aspect described herein with respect to any of the drawings depicted or referenced herein.
[0027] Figure 1A Illustrated are components of an AI-enabled visual data analysis system according to an embodiment.
[0028] Figure 1BIllustrates various sensors associated with the self according to an embodiment.
[0029] Figure 1C Illustrates components of a vehicle according to an embodiment.
[0030] Figure 2 Illustrates a flowchart of a simulation captured from a viewpoint in an environment rendered using ground truth heuristics according to an embodiment.
[0031] Figure 3 Illustrates a system architecture according to an embodiment.
[0032] Figure 4 Illustrates ground truth visualization according to an embodiment.
[0033] Figure 5 Illustrates road visualization according to an embodiment.
[0034] Figure 6 Illustrates road surface visualization according to an embodiment.
[0035] Figure 7A Illustrates road line visualization according to an embodiment.
[0036] Figure 7B Illustrates road line visualization according to an embodiment.
[0037] Figure 8A Illustrates external surface visualization according to an embodiment.
[0038] Figure 8B Illustrates filled external surface visualization according to an embodiment.
[0039] Figure 9A Illustrates traffic object visualization according to an embodiment.
[0040] Figure 9B Illustrates filled traffic object visualization according to an embodiment.
[0041] Figure 10 Illustrates filled traffic environment according to an embodiment.
[0042] Figure 11A Illustrates visualization of a modified environment scene according to an embodiment.
[0043] Figure 11B Illustrates visualization of a modified environment scene according to an embodiment.
[0044] Figure 12 Illustrates a rendered video object according to an embodiment.
[0045] Figure 13Illustrates a segmented geographical model according to an embodiment.
[0046] Figure 14 Illustrates a segmented architecture according to an embodiment. Detailed implementation
[0047] Aspects of the present technical solution are described herein with reference to the accompanying drawings, which are illustrative examples of the present technical solution. The following drawings and examples are not intended to limit the scope of the present technical solution to this implementation or a single implementation, and other implementations according to this implementation are also possible, for example, by interchanging some or all of the elements described or illustrated. When some elements of this implementation can be implemented in part or in whole using known components, only those parts of such known components that are necessary for understanding this implementation are described, and detailed descriptions of other parts of such known components are omitted so as not to obscure this implementation. Unless explicitly set forth herein, terms in the specification and claims should not be given an uncommon or special meaning. In addition, the present technical solution and this implementation cover existing and future known equivalents of known components mentioned herein by description, illustration, or example.
[0048] For example, the system can render multiple different three-dimensional environments that have one or more aspects corresponding to the physical characteristics of the physical environment. For example, the system can obtain a ground truth data model corresponding to one or more aspects of the physical environment detected from the physical environment. For example, aspects of the physical environment can include, but are not limited to, the boundary between a roadway and the surrounding land, the surface characteristics of the roadway, the surface markings of the roadway, traffic signs or lights at a specific location relative to the roadway, and the traffic pattern through the roadway.
[0049] The system can modify aspects of the virtual environment to generate many more instances than exist in a particular physical location or the conditions of a physical location. For example, the system can modify the virtual environment to modify a specific roadway marking to change the indication of traffic flow, change the wear level of the roadway surface, the water level of the roadway surface, the roadway markings or traffic signs, or change one or more objects around the roadway corresponding to weather, biomes, or urban density levels. The system can include dynamic objects that can be captured by one or more viewpoints. For example, the dynamic objects can include dynamic traffic objects, and the dynamic traffic objects include traffic lights or gates. In another example, the dynamic objects can include dynamic environmental objects, and the dynamic environmental objects include trees, branches, traffic cones, or other obstacles on the roadway that can affect the traffic pattern, or vehicles, pedestrians, or any combination thereof. Thus, the present technical solution can at least provide a technical improvement in creating numerous permutations of real-world environments that would otherwise be unavailable and cannot be manually detected or mapped.
[0050] The system can allocate the generation of multiple portions of a physical area to at least achieve a technical improvement in creating numerous arrangements of a real-world environment that would otherwise be unavailable and undetectable or drawable manually. For example, the system can divide a large geographical area into multiple physical locations and allocate one or more instructions to one or more processors or processor cores to render the corresponding physical locations. For example, the system can allocate various instructions to render various corresponding physical locations based on one or more aspects of the physical locations or one or more modifications to one or more aspects.
[0051] Figure 1A are non-limiting examples of components of a system in which the methods and systems discussed herein can be implemented. For example, an analytics server can train an artificial intelligence (AI) model and use the trained AI model to generate occupancy datasets and / or maps for one or more avatars. Figure 1A Illustrates components of an AI-enabled visual data analysis system 100. The system 100 can include an analytics server 110a, a system database 110b, an administrator computing device 120, avatars 140a to 140b (collectively referred to as (the) avatars 140), avatar computing devices 141a to 141c (collectively referred to as avatar computing devices 141), and a server 160. The system 100 is not limited to the components described herein and can include additional or other components not shown for the sake of brevity, which will be considered within the scope of the embodiments described herein.
[0052] The above components can be connected via a network 130. Examples of the network 130 can include, but are not limited to, private or public LANs, WLANs, MANs, WANs, and the Internet. The network 130 can include wired and / or wireless communication according to one or more standards and / or via one or more transmission media.
[0053] Communication over the network 130 can be performed according to various communication protocols such as Transmission Control Protocol and Internet Protocol (TCP / IP), User Datagram Protocol (UDP), and IEEE communication protocols. In one example, the network 130 can include wireless communication according to a Bluetooth specification set or another standard or proprietary wireless communication protocol. In another example, the network 130 can also include communication over a cellular network, including, for example, GSM (Global System for Mobile Communications), CDMA (Code Division Multiple Access), or EDGE (Enhanced Data for Global Evolution) networks.
[0054] The system 100 illustrates an example of a system architecture and components that can be used to train and execute one or more AI models such as (the) AI models 110c. Specifically, as Figure 1AAs depicted and described herein, the analytics server 110a can use the methods discussed herein to train the AI model(s) 110c using data retrieved from the vehicle 140 (e.g., by using data streams 172 and 174). When the AI model(s) 110c have been trained, each vehicle in the vehicle 140 can access and execute the trained AI model(s) 110c. For example, a vehicle 140a having a vehicle computing device 141a can transmit its camera feed to the trained AI model(s) 110c and can determine the occupancy status of its surrounding environment (e.g., data stream 174). In addition, the data ingested and / or predicted by the AI model(s) 110c with respect to the vehicle 140 (at inference time) can also be used to improve the AI model(s) 110c. Thus, the system 100 depicts a continuous loop that can periodically improve the accuracy of the AI model(s) 110c. In addition, the system 100 depicts a loop in which, in addition to the inference phase, the data received by the vehicle 140 can also be used in the training phase.
[0055] The analytics server 110a can be configured to collect, process, and analyze navigation data (e.g., images captured during navigation) and various sensor data collected from the vehicle 140. The collected data can then be processed and prepared into a training dataset. The training dataset can then be used to train one or more AI models, such as the AI model 110c. The analytics server 110a can also be configured to collect visual data from the vehicle 140. Using the AI model 110c (trained using the methods and systems discussed herein), the analytics server 110a can generate a dataset and / or an occupancy map for the vehicle 140. The analytics server 110a can display the occupancy map on the vehicle 140 and / or transmit the occupancy map / dataset to the vehicle computing device 141, the administrator computing device 120, and / or the server 160.
[0056] In Figure 1A , the AI model 110c is illustrated as a component of the system database 110b, but the AI model 110c can be stored in a different or separate component, such as a cloud storage device or any other data repository accessible to the analytics server 110a.
[0057] The analytics server 110a can also be configured to display an electronic platform that illustrates various training attributes for training the AI model 110c. The electronic platform can be displayed on the administrator computing device 120 such that an analyst can monitor the training of the AI model 110c. An example of an electronic platform generated and hosted by the analytics server 110a can be a web-based application or website configured to display the training dataset collected from the vehicle 140 and / or the training status / metrics of the AI model 110c.
[0058] The analysis server 110a can be any computing device including a processor and a non-transitory machine-readable storage device capable of performing the various tasks and processes described herein. Non-limiting examples of such computing devices can include workstation computers, laptop computers, server computers, etc. Although the system 100 includes a single analysis server 110a, the system 100 can include any number of computing devices operating in a distributed computing environment, such as a cloud environment.
[0059] The vehicle body 140 can represent various electronic data sources that transmit data associated with its previous or current navigation session to the analysis server 110a. The vehicle body 140 can be any device configured for navigation, such as the vehicle 140a and / or the truck 140c. The vehicle body 140 is not limited to being a vehicle and can also include robotic devices. For example, the vehicle body 140 can include the robot 140b, which can represent a general-purpose, bipedal, autonomous humanoid robot capable of navigating various terrains. The robot 140b can be equipped with software for achieving balance, navigation, perception, or interacting with the physical world. The robot 140b can also include various cameras configured to transmit visual data to the analysis server 110a.
[0060] Although referred to herein as the "vehicle body", the vehicle body 140 may or may not be an autonomous device configured for autonomous navigation. For example, in some embodiments, the vehicle body 140 can be controlled by a human operator or a remote processor. The vehicle body 140 can include various sensors, such as Figure 1B the sensors depicted in. The sensors can be configured to collect data as the vehicle body 140 navigates various terrains (e.g., roads). The analysis server 110a can collect the data provided by the vehicle body 140. For example, the analysis server 110a can obtain navigation session and / or road / terrain data (e.g., images of the vehicle body 140 navigating on a road) from various sensors, such that the collected data is ultimately used by the AI model 110c for training purposes.
[0061] As used herein, a navigation session corresponds to the journey of the ego 140's travel route, regardless of whether the journey is autonomous or human-controlled. In some embodiments, the navigation session can be used for data collection and model training purposes. However, in some other embodiments, the ego 140 can refer to a vehicle purchased by a consumer, and the purpose of the journey can be classified as daily use. A navigation session can begin when the ego 140 moves from a non-mobile position by more than a threshold distance (e.g., 0.1 miles, 100 feet) or at a rate greater than a threshold rate (e.g., greater than 0 mph, greater than 1 mph, greater than 5 mph). A navigation session can end when the ego 140 returns to a non-mobile position and / or is turned off (e.g., when the driver exits the vehicle).
[0062] The ego 140 can represent a collection of egos monitored by the analytics server 110a to train the AI model(s) 110c. For example, a driver of the vehicle 140a can authorize the analytics server 110a to monitor data associated with their respective vehicle. As a result, the analytics server 110a can use the various methods discussed herein to collect sensor / camera data and generate a training dataset to train the AI model(s) 110c accordingly. The analytics server 110a can then apply the trained AI model(s) 110c to analyze data associated with the ego 140 and predict an occupancy map for the ego 140. Additionally, additional / ongoing data associated with the ego 140 can also be processed and added to the training dataset so that the analytics server 110a can recalibrate the AI model(s) 110c accordingly. Thus, the system 100 depicts a cycle in which navigation data received from the ego 140 can be used to train the AI model(s) 110c. The ego 140 can include a processor that executes the trained AI model(s) 110c for navigation purposes. During navigation, the ego 140 can collect additional data about its navigation session, and this additional data can be used to calibrate the AI model(s) 110c. That is, the ego 140 represents an ego that can be used to train, execute / use, and recalibrate the AI model(s) 110c. In a non-limiting example, the ego 140 represents vehicles purchased by customers that can use the AI model(s) 110c for autonomous navigation while improving the AI model(s) 110c.
[0063] The ego 140 can be equipped with various technologies that allow the ego to collect data from its surrounding environment and (possibly) navigate autonomously. For example, the ego 140 can be equipped with an inference chip to run autonomous driving software.
[0064] Various sensors for each ego 140 can monitor the data collected associated with different navigation sessions and transmit the data to the analytics server 110a.Figures 1B to 1C FIG. 1 is a block diagram of a sensor integrated within the body 140 according to an embodiment. Figures 1B to 1C The number and location of each sensor discussed may depend on Figure 1A For example, robot 140b may include different sensors than vehicle 140a or truck 140c. For example, robot 140b may not include airbag activation sensor 170q. In addition, the sensors of vehicle 140a and truck 140c may be different from those of vehicle 140a and truck 140c. Figure 1C The shown are positioned differently.
[0065] As discussed herein, various sensors integrated within each self 140 can be configured to measure various data associated with each navigation session. The analysis server 110a can periodically collect data monitored and collected by these sensors, where the data is processed according to the methods described herein and used to train the AI model 110c and / or execute the AI model 110c to generate an occupancy map.
[0066] The host 140 may include a user interface 170a. The user interface 170a may refer to a host computing device (e.g., Figure 1A The user interface 170a may be implemented as a display screen, a head-up display, a touch screen, etc. that is integrated with or coupled to the interior of the vehicle. The user interface 170a may include input devices such as a touch screen, a knob, a button, a keyboard, a mouse, a gesture sensor, a steering wheel, etc. In various embodiments, the user interface 170a may be adapted to provide input to other devices or sensors (e.g., Figure 1B A sensor (such as controller 170c) is shown) providing user input (eg, as a signal and / or sensor information).
[0067] The user interface 170a can also be implemented with one or more logic devices that can be suitable for executing instructions, such as software instructions, to implement any of the various processes and / or methods described herein. For example, the user interface 170a can be suitable for forming a communication link, transmitting and / or receiving communications (e.g., sensor signals, control signals, sensor information, user input, and / or other information), or performing various other processes and / or methods. In another example, the driver can use the user interface 170a to control the temperature of the body 140 or activate its features (e.g., autonomous driving or steering system 170o). Therefore, the user interface 170a can be combined with other sensors described herein to monitor and collect driving session data. The user interface 170a can also be configured to display various data generated / predicted by the analysis server 110a and / or the AI model 110c.
[0068] The orientation sensor 170b can be implemented as a compass, a float, an accelerometer, and / or one or more of other digital or analog devices capable of measuring the orientation of the body 140 (e.g., the magnitude and direction of roll, pitch, and / or yaw relative to one or more reference orientations such as gravity and / or magnetic north). The orientation sensor 170b can be adapted to provide heading measurements to the body 140. In other embodiments, the orientation sensor 170b can be adapted to provide roll, pitch, and / or yaw rate to the body 140 using a time series of orientation measurements. The orientation sensor 170b can be positioned and / or adapted to make orientation measurements relative to a particular coordinate system of the body 140.
[0069] The controller 170c can be implemented as any suitable logic device (e.g., a processing device, a microcontroller, a processor, an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), a memory storage device, a memory reader, or other device or combination of devices) that can be adapted to execute, store, and / or receive appropriate instructions, such as software instructions implementing control loops for controlling various operations of the body 140. Such software instructions can also implement methods for processing sensor signals, determining sensor information, providing user feedback (e.g., via the user interface 170a), querying the device for operating parameters, selecting operating parameters for the device, or performing any of the various operations described herein.
[0070] The communication module 170e can be implemented as any wired and / or wireless interface configured to transfer sensor data, configuration data, parameters, and / or other data and / or signals to Figure 1A any of the features shown (e.g., the analytics server 110a). As described herein, in some embodiments, the communication module 170e can be implemented in a distributed manner such that portions of the communication module 170e are implemented within Figure 1B one or more of the elements and sensors shown. In some embodiments, the communication module 170e can delay the transfer of sensor data. For example, when the body 140 does not have a network connection, the communication module 170e can store the sensor data in a temporary data storage device and transmit the sensor data when the body 140 is identified as having an appropriate network connection.
[0071] The speed sensor 170d can be implemented as an electronic pitot tube, a metering gear or wheel, a water speed sensor, a wind speed sensor, a wind rate sensor (e.g., direction and magnitude), and / or other devices capable of measuring or determining the linear speed of the body 140 (e.g., in the surrounding medium and / or aligned with the longitudinal axis of the body 140) and providing such measurement as a sensor signal that can be transmitted to various devices.
[0072] The gyroscope / accelerometer 170f can be implemented as one or more electronic sextants, semiconductor devices, integrated chips, accelerometer sensors, or other systems or devices capable of measuring the angular velocity / acceleration and / or linear acceleration (e.g., direction and magnitude) of the vehicle body 140 and providing such measurements as sensor signals (which can be transmitted to various devices such as the analysis server 110a). The gyroscope / accelerometer 170f can be positioned and / or adapted to make such measurements with respect to a specific coordinate system of the vehicle body 140. In various embodiments, the gyroscope / accelerometer 170f can be implemented in a common housing and / or module with Figure 1B the other elements depicted to ensure a common reference frame or known transformation between reference frames.
[0073] The Global Navigation Satellite System (GNSS) 170h can be implemented as a global positioning satellite receiver and / or another device capable of determining the absolute and / or relative position of the vehicle body 140 based on wireless signals received, for example, from space and / or ground sources and capable of providing measurements such as sensor signals (which can be transmitted to various devices). In some embodiments, the GNSS 170h can be adapted to determine the rate, speed, and / or yaw rate of the vehicle body 140 (e.g., using a time series of position measurements), such as the yaw component of the absolute rate and / or angular rate of the vehicle body 140.
[0074] The temperature sensor 170i can be implemented as a thermistor, electrical sensor, electrical thermometer, and / or other device capable of measuring the temperature associated with the vehicle body 140 and providing such measurement as a sensor signal. The temperature sensor 170i can be configured to measure the ambient temperature associated with the vehicle body 140, such as the cockpit or dashboard temperature. For example, this ambient temperature can be used to estimate the temperature of one or more elements of the vehicle body 140.
[0075] The humidity sensor 170j can be implemented as a relative humidity sensor, electrical sensor, electrical relative humidity sensor, and / or another device capable of measuring the relative humidity associated with the vehicle body 140 and providing such measurement as a sensor signal.
[0076] The steering sensor 170g can be adapted to physically adjust the heading of the vehicle body 140 according to one or more control signals and / or user inputs provided by a logic device such as the controller 170c. The steering sensor 170g can include one or more actuators and control surfaces of the vehicle body 140 (e.g., a rudder or other type of steering or adjustment mechanism) and can be adapted to physically adjust the control surface to various positive and / or negative steering angles / positions. The steering sensor 170g can also be adapted to sense the current steering angle / position of such a steering mechanism and provide such measurement.
[0077] The propulsion system 170k can be implemented as a propeller, a turbine, or other thrust-based propulsion systems, mechanical wheeled and / or tracked propulsion systems, wind / sail-based propulsion systems, and / or other types of propulsion systems that can be used to power the vehicle body 140. The propulsion system 170k can also monitor the direction of the motive power and / or thrust of the vehicle body 140 relative to a reference coordinate system of the vehicle body 140. In some embodiments, the propulsion system 170k can be coupled to the sensor 170g and / or integrated with the steering sensor 170g.
[0078] The passenger restraint sensor 170l can monitor seat belt detection and locking / unlocking components as well as other passenger restraint subsystems. The passenger restraint sensor 170l can include various environmental and / or status sensors, actuators, and / or other devices that facilitate the operation of safety mechanisms associated with the operation of the vehicle body 140. For example, the passenger restraint sensor 170l can be configured to receive motion and / or status data from Figure 1B the other sensors depicted. The passenger restraint sensor 170l can determine whether safety measures (e.g., seat belts) are being used.
[0079] As Figure 1C depicted, the camera 170m can refer to one or more cameras integrated within the vehicle body 140, and can include multiple cameras integrated (or retrofitted) into the vehicle body 140. The camera 170m can be an internal or external-facing camera of the vehicle body 140. For example, as Figure 1C depicted, the vehicle body 140 can include one or more internal-facing cameras that can monitor and collect footage of the passengers of the vehicle body 140. The vehicle body 140 can include eight external-facing cameras. For example, the vehicle body 140 can include a front camera 170m-1, front side cameras 170m-2, 170m-3, rear side cameras 170m-4 on each front fender, cameras 170m-5 on each side (e.g., integrated within the B-pillar), and a rear camera 170m-6.
[0080] Reference Figure 1B , the radar 170n and ultrasonic sensor 170p can be configured to monitor the distance of the vehicle body 140 to other objects (such as other vehicles or immovable objects (e.g., trees or garage doors)). The vehicle body 140 can also include an autonomous driving or steering system 170o configured to navigate the vehicle body 140 autonomously using data collected via various sensors (e.g., the radar 170n, speed sensor 170d, and / or ultrasonic sensor 170p).
[0081] Accordingly, the autonomous driving or steering system 170o can analyze various data collected by one or more sensors described herein to identify driving data. For example, the autonomous driving or steering system 170o can calculate the risk of a forward collision based on the speed of the vehicle body 140 and its distance to another vehicle on the road. The autonomous driving or steering system 170o can also determine whether the driver is touching the steering wheel. The autonomous driving or steering system 170o can transmit the analyzed data to various features discussed herein, such as the analysis server.
[0082] The airbag activation sensor 170q can predict or detect a collision and cause the activation or deployment of one or more airbags. The airbag activation sensor 170q can transmit data regarding the airbag deployment, including data associated with the event that caused the deployment.
[0083] Referring again to Figure 1A , the administrator computing device 120 can represent a computing device operated by a system administrator. The administrator computing device 120 can be configured to display data retrieved or generated by the analysis server 110a (e.g., various analysis metrics and risk scores), where the system administrator can monitor the various models utilized by the analysis server 110a, review feedback, and / or facilitate the training of the AI model(s) 110c maintained by the analysis server 110a.
[0084] The vehicle body(ies) 140 can be any device configured to navigate various routes, such as the vehicle 140a or the robot 140b. As discussed with respect to Figures 1B to 1C , the vehicle body 140 can include various telemetry sensors. The vehicle body 140 can also include a vehicle body computing device 141. Specifically, each vehicle body can have its own vehicle body computing device 141. For example, the truck 140c can have a vehicle body computing device 141c. For the sake of brevity, the vehicle body computing devices are collectively referred to as the vehicle body computing device(s) 141. The vehicle body computing device 141 can control the content presentation on the infotainment system of the vehicle body 140, process commands associated with the infotainment system, aggregate sensor data, manage the communication of data to electronic data sources, receive updates, and / or transmit messages. In one configuration, the vehicle body computing device 141 communicates with an electronic control unit. In another configuration, the vehicle body computing device 141 is an electronic control unit. The vehicle body computing device 141 can include a processor and a non-transitory machine-readable storage medium capable of performing the various tasks and processes described herein. For example, the AI model(s) 110c described herein can be stored and executed (or directly accessed) by the vehicle body computing device 141. Non-limiting examples of the vehicle body computing device 141 can include vehicle multimedia and / or display systems.
[0085] In one example of how to train the AI model(s) 110c, the analysis server 110a can collect data from the simulated ego 140 to train the AI model(s) 110c. Before executing the AI model(s) 110c to generate / predict a training data set, the analysis server 110a can generate training data. The training allows the AI model(s) 110c to ingest data from one or more simulated cameras of one or more simulated egos 140 (without receiving radar data). The operations described in this example can be performed by Figure 1A and Figure 1B any number of computing devices operating in the distributed computing system described in.
[0086] To train the AI model(s) 110c, the analysis server 110a can first employ one or more egos in the ego 140 to drive a specific simulated route in a virtual environment. While driving, the ego 140 can use one or more of its simulated sensors (including one or more simulated cameras) to generate navigation session data. For example, one or more simulated egos in the simulated ego 140 equipped with various simulated sensors can navigate a specified route in a virtual environment. As one or more egos in the ego 140 traverse the terrain, their simulated sensors can capture continuous (or periodic) data of their surrounding environment. The simulated sensors can capture visual information of the surrounding environment of one or more egos 140.
[0087] The analysis server 110a can generate a training data set using the data collected from the ego 140 (e.g., the simulated camera feed received from the ego 140). The training data set can indicate video from one or more simulated cameras that corresponds to the viewpoints of the simulated ego 140 within the surrounding environment of one or more egos in the ego 140.
[0088] In operation, as one or more egos 140 navigate, their sensors collect data and transmit the data to the analysis server 110a, as depicted by the data stream 172.
[0089] In some embodiments, one or more egos 140 can include one or more high-resolution cameras that capture a continuous visual data stream from the surrounding environment of one or more egos 140 as one or more egos 140 navigate a route. The analysis server 110a can then use the camera feed to generate a second data set that includes visual elements / depictions of different voxels of the surrounding environment of one or more egos 140.
[0090] In operation, when one or more avatars 140 are navigating, their cameras collect data and transmit the data to the analysis server 110a, as depicted by data stream 172. For example, the avatar computing device 141 may use data stream 172 to transmit image data to the analysis server 110a.
[0091] Figure 2 Depicts an example method of a simulation captured from a viewpoint in an environment rendered using ground truth heuristics according to the present disclosure. Figures 1A to 1C At least one of the systems or vehicles of any of them may execute method 200.
[0092] At 210, method 200 may retrieve a camera feed of an avatar object navigating within a physical environment. For example, the camera feed may correspond to a series of images or videos captured from a viewpoint within a virtual environment by the virtual environment. For example, the avatar object may correspond to a simulated avatar object traveling within the virtual environment.
[0093] At 220, method 200 may generate a 3D model according to one or more first environmental metrics, the 3D model including a first surface corresponding to one or more physical paths through the physical environment, and the one or more first environmental metrics indicating boundaries of the one or more physical paths. For example, the boundaries may correspond to one or more points, vectors, planes, or volumes defining the edges and surface topologies of one or more roadway surfaces. For example, the boundaries may be obtained or generated from the detection of the surface topology of the physical environment via one or more sensors. For example, the boundaries may define a road and a curb grid.
[0094] At 230, method 200 may generate one or more geometric 2D objects on the first surface according to one or more second environmental metrics, and the second environmental metrics indicate the one or more physical paths. For example, the 2D objects may correspond to one or more points, vectors, planes, textures, images, or patterns defining two-dimensional objects on one or more roadway surfaces. For example, the 2D objects may indicate lane markings, direction markings, or any combination thereof. For example, the 2D objects may be obtained or generated from the detection of the surface image of the physical environment via one or more sensors. For example, the 2D objects may define lane paint decals and direction road markings.
[0095] At 240, method 200 may identify one or more viewpoints oriented to capture corresponding portions of the 3D model based on one or more viewpoint metrics of a camera of a physical object configured to move along one or more physical paths. For example, the camera of the physical object may be a simulated viewpoint in a virtual environment and capture a rendered video object. The viewpoints may include a front view, a right front view, a right rear view, a left front view, and a left rear view. Each viewpoint may correspond to a 2D image or 2D video captured by a simulated ego vehicle traveling through the virtual environment.
[0096] At 250, method 200 may render one or more 2D images from one or more corresponding portions of the 3D model, each of the one or more 2D images corresponding to a respective one of the one or more viewpoints. For example, method 200 may render a 2D video by capturing images from the simulated viewpoints over time.
[0097] At 260, method 200 may train an artificial intelligence model based on the camera feed and one or more simulated environment 2D images. For example, AI model 110c may be used to train the scenario. For example, the system may obtain one or more models generated from boundary measurements of various roads, medians, crosswalks, sidewalks, bike lanes, or any other features of an outdoor built environment. The system may add textures and objects to the roadway and surrounding space to generate a visualization of an intersection as it appears in the real world. The system may change, remove, or add street signs, pavement, street markings, weather conditions, road obstacles, traffic markings, or any combination thereof to create multiple variations of the intersection while maintaining the measurement characteristics of the intersection as it appears in the real world. When changing the environment, the system may change the time of day, weather, or other conditions. In this way, the system may at least provide a technical improvement for providing multiple variations of a real-world location to provide a quantity and type of training data far exceeding what can be obtained by a manual process in the real world, at least including because many of the environments in the environment that can be generated by the technical solution do not exist and cannot exist in the real world.
[0098] Figure 3 Depicts an example system architecture in accordance with the present disclosure. As illustrated by the example in Figure 3 The system architecture 300 may at least include a ground truth model 310 and a geospatial model 320, each indicating in various ways but not limited to a road and curb grid 330, lane paint decals 332, islands 334, scene arrangements 340, directional road markings 342, traffic lights and stop signs 350, buildings 360, and areas and zones outside the drivable area 362. For example, AI model 110c may have or be configured according to system architecture 300.
[0099] The ground truth model 310 can include a data structure indicating physical aspects of a physical environment in the real world. For example, the ground truth model 310 can correspond to a space or geographical identifier in a coordinate space. For example, the ground truth model 310 can correspond to one or more sets of one or more spatial elements. For example, the spatial elements can be defined or determined according to a coordinate system indicating or convertible to indicate the physical environment. For example, the spatial elements can correspond to one or more of latitude, longitude, or altitude, or can be defined relative to other spatial elements. The ground truth model 310 can correspond to a set of models each indicating different aspects of the physical environment. The ground truth model 310 can include a road boundary model 312, a road line model 314, an intermediate edge model 316, and a lane map model 318. For example, the ground truth model 310 can be derived from measurements of the physical environment in the real world.
[0100] The road boundary model 312 can indicate the structure of one or more outer edges of one or more roadways in the physical environment. For example, the road boundary model 312 can correspond to one or more points, vectors, planes, or volumes defining the edges and surface topology of one or more roadway surfaces. For example, the road boundary model 312 can be obtained or generated via detection of the surface topology of the physical environment by one or more sensors. For example, the road boundary model can define a road and a curb grid 330. The road line model 314 can indicate the structure of one or more markings on one or more roadways in the physical environment. For example, the road line model 314 can correspond to one or more points, vectors, planes, textures, images, or patterns defining two-dimensional objects on one or more roadway surfaces. For example, the road line model 314 can indicate lane markings, direction markings, or any combination thereof. For example, the road line model 314 can be obtained or generated via detection of the surface image of the physical environment by one or more sensors. For example, the road line model 314 can define a lane paint decal 332 and a direction road marking 342.
[0101] The intermediate edge model 316 can indicate the structure of one or more inner edges of one or more roadways in the physical environment. For example, the road boundary model 312 or the intermediate edge model 316 can correspond to one or more points, vectors, planes, or volumes defining the edges and surface topology of one or more roadway surfaces at least partially located within one or more road boundary models of the road boundary model 312. For example, the intermediate edge model 316 can be obtained or generated via detection of the surface topology of the physical environment by one or more sensors. For example, the intermediate edge model 316 can define an island 334 corresponding to a non-drivable area surrounded by the roadway.
[0102] The lane map model 318 can indicate the movement patterns of one or more movable objects in a lane in a physical environment. For example, the lane map model 318 can correspond to one or more points, vectors, planes, or volumes that define a passage along one or more lane surfaces according to one or more markings of the road line model 314. For example, the road boundary model 312 can be obtained or generated from the detection of the surface topology of the physical environment via one or more sensors.
[0103] The geospatial model 320 can include a data structure that indicates physical aspects of the physical environment in the real world that are different from the ground truth model 310. For example, the geospatial model 320 can at least partially correspond to the ground truth model 310 in one or more aspects of its structure and operation. The geospatial model 320 can be obtained from external sensors or sensors different from those used to detect the ground truth model 310. For example, sensors on a vehicle traversing one or more lanes can detect data to generate the ground truth model 310, and sensors on a satellite or data corresponding to a geographic information system (GIS) can generate the geospatial model 320. The geospatial model 320 can include a map model 322, an environment model 324, and a scene arrangement model 326.
[0104] The map model 322 can indicate the structure of one or more objects along one or more lanes in the physical environment. For example, the road boundary model 312 can correspond to one or more points, vectors, planes, or volumes that define the shape and location of one or more objects different from the lane surface. For example, the map model 322 can be obtained or generated from the detection of an image of the physical environment on or around the lane via one or more sensors. For example, the map model 322 can define one or more of the traffic signal and stop sign data 350.
[0105] The environment model 324 can indicate the structure of one or more objects along one or more lanes in the physical environment. For example, the road boundary model 312 can correspond to one or more points, vectors, planes, or volumes that define the shape and location of one or more objects different from the road surface. For example, the environment model 324 can be obtained or generated from the detection of an image of the physical environment on or around the roadway via one or more sensors.
[0106] The scene arrangement model 326 can include one or more instructions for modifying one or more of the ground truth model 310 or the geospatial model 320 corresponding to a specific 3D model of the physical environment. For example, the scene arrangement model can include one or more instructions to populate or transform one or more of the environmental models 324 into one or more scene arrangements in the scene arrangement 340 according to one or more environmental heuristics. For example, the environmental heuristics can correspond to weather, biomes, or urban density levels. For example, the scene arrangement model 326 can modify or replace one or more of the buildings 360 and the sites and areas 362 outside the drivable area according to the environmental heuristics.
[0107] Figure 4 Depicts an example ground truth visualization according to the present disclosure. As illustrated by the example in Figure 4 The ground truth model 400 can at least include a road boundary model 410, a road line model 420, a median edge model 430, and a lane map model 440. The ground truth model 400 can correspond to a 3D model including at least one of the ground truth model 310 or the geospatial model 320, but is not limited thereto.
[0108] The road boundary model 410 can correspond to the rendering or instance of the road boundary model among the road boundary models 312 that indicates the road boundary of the road 330. For example, the road boundary model 410 can include the contours indicating one or more intersecting roadways. The road line model 420 can correspond to the rendering or instance of the road line model among the road line models 314 that indicates one or more lane paint decals 332 in the lane paint decals. For example, the road line model 420 can include the contours indicating one or more markings on the corresponding surfaces of one or more intersecting roadways. The median edge model 430 can correspond to the rendering or instantiation of the median edge model of the island that indicates one or more islands among the median edge models 316. For example, the median edge model 430 can include the contours indicating one or more islands at least partially within one or more intersecting roadways of the road boundary model 410. The lane map model 440 can correspond to the rendering or instantiation of the movement path that indicates one or more movement paths of corresponding vehicles, pedestrians, or any combination thereof among the lane map models 318. For example, the lane map model 440 can include the passageways for one or more vehicles along one or more intersecting roadways. For example, the lane map model 440 can include the passageways for one or more pedestrians crossing one or more intersecting roadways.
[0109] Figure 5 Depicts an example road visualization according to the present disclosure. As illustrated by the example in Figure 5As illustrated by the example in, the road model 500 may at least include a road topology model 510. The road topology model 510 may correspond to the rendering or instantiation of the surface topology of the road surface based on the road boundary model among the road boundary models 312 that indicate the curb grid 330. For example, the road topology model 510 may include a grid surface indicating the surface elevation of one or more intersecting roadways at one or more corresponding points along the roadways. The road boundary model 520 may at least partially correspond to the road boundary model 410 in one or more aspects of structure and operation.
[0110] Figure 6 depicts an example road surface visualization according to the present disclosure. As illustrated by the example in Figure 6 As illustrated by the example in, the road surface model 600 may at least include a textured road topology model 610. The textured road topology model 610 may correspond to the road topology model 510 with a rendered road surface that indicates the road surface at a physical location. For example, the textured road topology model 610 may include a road surface having a texture corresponding to or indicating asphalt. For example, the textured road topology model 610 may include one or more surface features, or may have one or more surface features existing thereon. The textured road topology model 610 may include topological features 620 and environmental features 630. The topological features 620 may correspond to a portion of the textured road topology model 610 that has a topology indicating a road feature having a shape different from the plane of the road topology model 510 corresponding to the road feature around the road feature. For example, the road feature may correspond to a pothole or a speed bump, but is not limited thereto. The environmental features 630 may correspond to a portion of the textured road topology model 610 that has environmental characteristics different from the road characteristics of the road topology model 510 corresponding to the road feature around the road feature. For example, the environmental characteristics may correspond to a puddle having a first reflectivity greater than the second reflectivity of the asphalt of the textured road topology model 610, but is not limited thereto. The road boundary model 640 may at least partially correspond to the road boundary model 410 in one or more aspects of structure and operation.
[0111] Figure 7A depicts an example road line visualization according to the present disclosure. As illustrated by the example in Figure 7AAs illustrated by the example in, the road line model 700A may include at least a textured road topology model 710A. The textured road topology model 710A may at least partially correspond to the textured road topology model 610 in one or more aspects of structure and operation. The road line model 700A may include a road line model 712 and the textured road topology model 610, and may correspond to the state of the 3D model before rendering one or more lane paint decals corresponding to the lane paint decals 332 of the road line model 712. The road line model 712 may at least partially correspond to the road line model 420 in one or more aspects of structure and operation.
[0112] Figure 7B depicts an example road line visualization according to the present disclosure. As illustrated by the example in Figure 7B As illustrated by the example in, the road line model 700B may include at least a line-mapped road surface 710B. The line-mapped road surface 710B may at least partially correspond to the textured road topology model 710A in one or more aspects of structure and operation, and may include one or more lane paint decals corresponding to the lane paint decals 332 of the road line model 420. Thus, the road line model 700B may correspond to the state of the 3D model after rendering one or more lane paint decals corresponding to the lane paint decals 332 of the road line model 420.
[0113] Figure 8A depicts an example external surface visualization according to the present disclosure. As illustrated by the example in Figure 8A As illustrated by the example in, the external surface model 800A may include at least a line-mapped road surface 802, an intermediate edge model 804, an intermediate region 810A, and a blocking region 820A. The intermediate region 810A may correspond to a part of the physical environment corresponding to one or more islands 334. The intermediate region 810A may be at least partially surrounded by at least a part of the road line model 420. The blocking region 820A may correspond to a part of the physical environment corresponding to one or more 362 outside the drivable area. The blocking region 820A may be at least partially located outside at least a part of the road line model 420. For example, the blocking region 820A may correspond to a part of the physical environment indicating an urban block outside the right-of-way (including but not limited to a roadway). The line-mapped road surface 802 may at least partially correspond to the line-mapped road surface 710B in one or more aspects of structure and operation. The intermediate edge model 804 may at least partially correspond to the intermediate edge model 430 in one or more aspects of structure and operation.
[0114] Figure 8B depicts an example filled external surface visualization according to the present disclosure. As illustrated by the example in Figure 8BAs illustrated by the example in, the filled outer surface model 800B may include at least a filled intermediate region 810B and a filled blocking region 820B.
[0115] The filled intermediate region 810B may at least partially correspond to the intermediate region 810A in one or more aspects of structure and operation, and may include one or more objects that are at least partially disposed within one or more islands 334 corresponding to the intermediate region 810A and the filled intermediate region 810B. For example, the filled intermediate region 810B may correspond to one or more islands 334 having one or more tree objects thereon and including a grass surface having a color and texture different from that of the road surface 710B. For example, the tree object may include a dynamic component that can interact with the 3D environment according to one or more scene arrangements in the scene arrangement 340. For example, the tree object may include a branch object or a leaf object that covers at least a portion of the filled intermediate region 810B or the road surface 710B. The filled blocking region 820B may at least partially correspond to the blocking region 820A in one or more aspects of structure and operation, and may include one or more objects that are at least partially disposed within the filled blocking region 820B. For example, the filled blocking region 820B may correspond to one or more regions outside the drivable area 330, and the one or more regions have one or more building objects or tree objects thereon and include a sidewalk surface having a color and texture different from that of the road surface 710B. For example, one or more objects in the filled intermediate region 810B and the filled blocking region 820B may be modified or replaced according to one or more scene arrangements in the scene arrangement 340.
[0116] Figure 9A depicts an example traffic object model according to the present disclosure. As shown by the example in Figure 9A the traffic object model 900A may include at least a line-mapped road surface 902 and a map data model 910. The map data model 910 may at least partially correspond to the map model 322 in one or more aspects of structure and operation. For example, the map data model 910 may correspond to the map data model among the map models 332 for the road surface 710B. The map data model 910 may include one or more indicators corresponding to one or more traffic objects corresponding to a specific physical location. For example, the map data model 910 may include location data for one or more traffic lights or stop signs 350. The line-mapped road surface 902 may at least partially correspond to the line-mapped road surface 802 in one or more aspects of structure and operation.
[0117] Figure 9B depicts an example filled traffic object model according to the present disclosure. As shown by the example inFigure 9B As illustrated in the example of, the populated traffic object model 900B may include at least a dynamic traffic object 920 and a static traffic object 930. The dynamic traffic object 920 may correspond to an object having multiple states that can be rendered in a 3D model corresponding to a physical environment. For example, the dynamic traffic object 920 may correspond to a traffic signal that can be rendered to indicate multiple states according to a three-light stop light for vehicle traffic. For example, the dynamic traffic object 920 may correspond to a traffic signal that can be rendered to indicate multiple states according to a multi-state illuminated sign for pedestrian or bicycle traffic. The static traffic object 930 may correspond to an object having multiple states that can be rendered in a 3D model corresponding to a physical environment. For example, the static traffic object 930 may correspond to a stop sign that can be rendered in a single state according to a stop sign texture. For example, the static traffic object 930 may correspond to a street sign that can be rendered in a single state according to a street sign texture and a location indicated by the map data model 910.
[0118] Figure 10 depicts a populated traffic environment according to the present disclosure. As shown by Figure 10 the example of, the populated traffic environment 1000 may include at least a line-mapped road surface 1002, a lane graph model 1004, a roadway traffic object 1010, and an intersection traffic object 1020. The roadway traffic object 1010 may correspond to a 3D object that can move along one or more paths of the lane graph model 440 corresponding to the road surface 1002. For example, the roadway traffic object 1010 may correspond to a vehicle including a car, a bicycle, a motorcycle, a truck, or any combination thereof, but is not limited thereto. The intersection traffic object 1020 may correspond to a 3D object that can move along one or more paths of the lane graph model 440 corresponding to the road surface 1002. For example, the roadway traffic object 1010 may correspond to a vehicle including a car, a bicycle, a motorcycle, a truck, or any combination thereof, but is not limited thereto. The line-mapped road surface 1002 may at least partially correspond to the line-mapped road surface 710B in one or more aspects of structure and operation. The lane graph model 1004 may at least partially correspond to the lane graph model 1004 in one or more aspects of structure and operation.
[0119] Figure 11A depicts an example model with a modified environmental scene according to the present disclosure. As shown by Figure 11AAs illustrated by the example in, a model with a modified environmental scene 1100A can include at least a line-mapped road surface 1102, environmental features 1104, a modified weather environment 1110A, and a place environment 1120A. The line-mapped road surface 1102 can at least partially correspond to the line-mapped road surface 710B in one or more aspects of structure and operation. The environmental feature 1004 can at least partially correspond to the environmental feature 630 in one or more aspects of structure and operation.
[0120] The modified weather environment 1110A can at least partially correspond to the populated traffic environment 1000 in one or more aspects of structure and operation, and can include one or more visual characteristics that correspond to weather conditions different from the physical environment in the real world captured by one or more sensors. For example, the modified weather environment 1110A can include modifications to the road surface 710B to include additional puddle objects, modifications to tree objects to include additional leaf objects or branch objects on the road surface 710B, or can include one or more objects or filters corresponding to a reduced visible distance within the populated traffic environment 1000. Thus, the modified weather environment 1110A can at least provide a technical improvement to eliminate or minimize additional sensor captures of the physical environment in multiple weather states according to the ground truth model 310 while maintaining the accuracy of the model relative to the corresponding physical environment in the real world.
[0121] The place environment 1120A can at least partially correspond to the populated traffic environment 1000 in one or more aspects of structure and operation, and can include one or more visual characteristics corresponding to the place associated with the physical environment in the real world. For example, the place environment 1120A can be linked to an urban environment and can be linked to multiple objects that are correspondingly linked to the urban environment. The multiple objects can include, for example, models of buildings 360 having a predetermined size, height, or footprint. Thus, the place environment 1120A can correspond to a place or biome corresponding to the physical environment in the real world captured by one or more sensors.
[0122] Figure 11B Depicts an example model with a modified environmental scene according to the present disclosure. As illustrated by the example in Figure 11B A model with a modified environmental scene 1100B can include at least a weather environment 1110B and a modified place environment 1120B.
[0123] The weather environment 1110B may at least partially correspond to the populated traffic environment 1000 in one or more aspects of structure and operation, and may include one or more visual characteristics that correspond to the weather conditions of the physical environment in the real world captured by one or more sensors. For example, the weather environment 1110B may include a road surface 710B without puddle objects, without leaf objects or branch objects on the road surface 710B, or without objects or filters corresponding to a reduced visible distance within the populated traffic environment 1000.
[0124] The modified venue environment 1120B may at least partially correspond to the populated traffic environment 1000 in one or more aspects of structure and operation, and may include one or more visual characteristics corresponding to a venue different from the physical environment in the real world. For example, the modified venue environment 1120B may be linked to a rural environment or a tropical environment, and may be linked to a plurality of objects correspondingly linked to the corresponding environment. The plurality of objects may include, for example, a model of a building 360 having a predetermined size, height, or footprint. Thus, the modified venue environment 1120B may correspond to a venue or biome different from the physical environment in the real world captured by one or more sensors. Thus, the modified venue environment 1120B may at least provide a technical improvement in creating additional locations with different physical characteristics that do not exist in the real world according to the ground truth model 310 while maintaining the accuracy of the model relative to the corresponding physical environment in the real world.
[0125] Figure 12 Depicts a video object rendered according to an example of the present disclosure. As shown by Figure 12As illustrated by the examples in, the rendered video object 1200 may include at least a front view video object 1210, a right front view video object 1220, a right rear view video object 1230, a left front view video object 1240, and a left rear view video object 1250. Each of the video objects 1210, 1220, 1230, 1240, and 1250 may correspond to a 2D image or 2D video captured by the simulated ego vehicle 140 when traveling in one or more of the environments 900B, 1000, 1100A, or 1100B. The front view video object 1210 may correspond to a 2D image or 2D video captured by the simulated ego vehicle 140 from a first simulated viewpoint that is oriented at the front of the simulated ego vehicle 140 and facing away from the front of the simulated ego vehicle 140. The right front view video object 1220 may correspond to a 2D image or 2D video captured by the simulated ego vehicle 140 from a second simulated viewpoint that is oriented at the right front of the simulated ego vehicle 140 and facing away from the right front of the simulated ego vehicle 140. The right rear view video object 1230 may correspond to a 2D image or 2D video captured by the simulated ego vehicle 140 from a third simulated viewpoint that is oriented at the right rear of the simulated ego vehicle 140 and facing away from the right rear of the simulated ego vehicle 140. The left front view video object 1240 may correspond to a 2D image or 2D video captured by the simulated ego vehicle 140 from a fourth simulated viewpoint that is oriented at the left front of the simulated ego vehicle 140 and facing away from the left front of the simulated ego vehicle 140. The left rear view video object 1250 may correspond to a 2D image or 2D video captured by the simulated ego vehicle 140 from a fifth simulated viewpoint that is oriented at the left rear of the simulated ego vehicle 140 and facing away from the left rear of the simulated ego vehicle 140.
[0126] Figure 13 depicts an example segmented geographic model according to the present disclosure. As shown by Figure 13As illustrated by the example in, the segmented geography model 1300 can include at least a geography region model 1310. The geography region model 1310 can correspond to a geographic region indicating a physical area. For example, the geography region model 1310 can correspond to a physical area for San Francisco, California. The geography region model 1310 can be linked to, define, or include one or more ground truth models 310 and one or more geospatial models 320 corresponding to multiple physical locations of the geographic region. The geography region model 1310 can include region segments 1320. Each region segment 1320 can correspond to a respective part of the geography region model 1310. For example, the size of each region segment 1320 in the region segments 1320 can be determined based on the quantity or type of computing resources required to render a 3D model according to the ground truth model 310 and one or more geospatial models 320 for each region segment among the region segments 1320. For example, one or more region segments in the region segments 1320 can be set to include the size of a physical area that can be allocated within the computing limits of each processor or processor core of a multi-processor system, and the physical area includes the ground truth model 310 and one or more geospatial models 320. The multi-processor system can correspond to any computer environment discussed herein.
[0127] Figure 14 depicts an example segmented architecture according to the present disclosure. As illustrated by the example in Figure 14 As illustrated by the example in, the segmented architecture 1400 can include at least a block creator 1410, a block extractor 1420, a block loader 1430, a rendering engine 1440, and one or more rendered video objects 1450, and can obtain one or more models among the ground truth model 1402 and the geospatial model 1404. The ground truth model 1402 can at least partially correspond to the ground truth model 310 in one or more aspects of structure and operation. The geospatial model 1404 can at least partially correspond to the geospatial model 320 in one or more aspects of structure and operation.
[0128] The block creator 1410 can determine the size of one or more blocks according to at least one of the ground truth model 310 and the geospatial model 320. For example, the block creator 1410 can identify the quantity of computing resources required to render these parts according to the ground truth model 310 and the geospatial model 320 associated with each part among one or more parts of the geography region model 1310. Then, the block creator 1410 can generate the region segments 1320 according to the determined size and according to the computing limits of each processor or processor core of the multi-processor system.
[0129] The block extractor 1420 can identify models corresponding to each of the region segments in the region segment 1320. For example, the block extractor 1420 can identify portions of the ground truth model 310 and the geospatial model 320 that have coordinates within the respective boundaries of each of the region segments in the region segment 1320. The block extractor 1420 can include model geometry 1422 and model instances 1424. The model geometry 1422 can correspond to portions of the ground truth model 310 and the geospatial model 320 that indicate physical locations in the real world. For example, the model geometry 1422 can correspond to one or more of the models 410, 420, 430, and 440. The model instances 1424 can correspond to portions of the ground truth model 310 and the geospatial model 320 that indicate objects in physical locations in the real world. For example, the model instances 1424 can correspond to one or more of the models indicating 350 and 360.
[0130] The block loader 1430 can assign portions of the ground truth model 310 and the geospatial model 320 that indicate physical locations in the real world to corresponding processors or processor cores. For example, the block loader 1430 can identify a processor core having a computational limit corresponding to or greater than the computational requirements for the corresponding portion among the portions. The rendering engine 1440 can convert one or more of the region segments in the region fragment 1320 into one or more corresponding 3D models. The rendering engine can generate, for example, the region segment 1320 according to the environments 400 to 100B discussed herein.
[0131] In an example of operation, the embodiments herein can be used in a training scenario. For example, the system can obtain one or more models generated from measurements of the boundaries of various roads, medians, crosswalks, sidewalks, bike lanes, or any other features of the outdoor built environment. The system can obtain information generated from measurements taken while traveling through a roadway (e.g., using a vehicle equipped with a camera, radar, lidar, or any combination) or from external detection of aspects of the roadway or the surrounding environment (e.g., via satellite imagery or a geospatial database for a region). The system can apply the measurements to generate a model of a given intersection in a given city block. The system can add textures and objects to the roadway and the surrounding space to generate a visualization of the intersection as it appears in the real world. The system can include traffic pattern paths that control the movement of vehicles, pedestrians, and other simulated roadway objects through the intersection. The system can change, remove, or add street signs, road surfaces, street markings, weather conditions, road obstacles, traffic markings, or any combination thereof to create multiple variations of the intersection while maintaining the measurement characteristics of the intersection as it appears in the real world. In this way, the system can at least provide a technical improvement for providing multiple variations of a real-world location to provide a quantity and type of training data far exceeding what can be obtained by a manual process in the real world, at least including that many of the environments that can be generated by this technical solution do not exist and cannot exist in the real world. For example, it is not possible in the real world to have an intersection with both a two-way traffic pattern and a one-way traffic pattern on the same roadway. However, this technical solution can create training data that matches both of these traffic configurations.
[0132] Some illustrative implementations have now been described. The foregoing is illustrative and not restrictive, and has been presented by way of example. Specifically, although many of the examples presented herein involve specific combinations of method acts or system elements, these acts and these elements can be combined in other ways to achieve the same goal. The acts, elements, and features discussed in connection with one implementation are not intended to exclude similar roles in other implementations.
[0133] The words and terms used herein are for the purpose of description and should not be regarded as limiting. The use herein of "including / comprising", "having", "containing", "involving", "characterized by", "characterized in that", and variations thereof is intended to cover the items listed hereinafter, their equivalents, and additional items, as well as alternative implementations consisting only of the items listed hereinafter. In one implementation, the systems and methods described herein consist of one, each combination of more than one, or all of the recited elements, acts, or components.
[0134] References to "or" may be construed as inclusive, such that any item described using "or" may indicate any of a single, plural, and all of the stated items. References to at least one item in a conjunctive list of items may be construed as an inclusive "or" to indicate any of a single, plural, and all of the stated items. For example, a reference to "at least one of 'A' and 'B'" may include only 'A', only 'B', and both 'A' and 'B'. Such references used in conjunction with "comprising" or other open terms may include additional items. References to "is" (singular) or "are" (plural) may be construed as not limiting the implementations or actions associated with the term. Unless otherwise specified herein, the term "is" (singular) or "are" (plural) or any tense or derivative thereof is interchangeable and synonymous with "may be" as used herein.
[0135] The direction indicators depicted herein are example directions to facilitate understanding of the examples discussed herein and are not limited to the direction indicators depicted herein. Unless otherwise specified herein, any direction indicator depicted herein may be modified to the opposite direction or may be modified to include both the depicted direction and the direction opposite to the depicted direction. Although the operations are depicted in a particular order in the figures, these operations are not required to be performed in the particular order shown or in a sequential order, and not all of the illustrated operations are required to be performed. The actions described herein may be performed in a different order. When a technical feature in the figures, detailed description, or any claim is followed by a reference sign, the reference sign has been included to enhance the understandability of the figures, detailed description, and claims. Accordingly, neither the reference sign nor its absence has any limiting effect on the scope of any claim element.
[0136] Accordingly, the scope of the systems and methods described herein is indicated by the appended claims rather than the foregoing description. The scope of the claims includes equivalents of the meaning and scope of the appended claims.
Claims
1. A system, comprising: A non-transitory memory and one or more processors, the one or more processors being configured to: Retrieve a camera feed of an ego object navigating within a physical environment; Generate a three-dimensional (3D) model based on the camera feed according to one or more first environmental metrics, the 3D model including a first surface corresponding to one or more physical paths traversing the physical environment, the one or more first environmental metrics indicating boundaries of the one or more physical paths; Generate one or more geometric two-dimensional (2D) objects on the first surface according to one or more second environmental metrics, the second environmental metrics indicating the one or more physical paths; Identify one or more viewpoints oriented to capture corresponding portions of the 3D model according to one or more viewpoint metrics of a camera of a physical object configured to move along the one or more physical paths; Render one or more simulated environment 2D images from the one or more corresponding portions of the 3D model of the physical environment, each of the one or more simulated environment 2D images corresponding to a respective one of the viewpoints; And Train an artificial intelligence model by the processor based on the camera feed and the one or more simulated environment 2D images.
2. The system according to claim 1, wherein the processor is configured to: Modify a part of the first surface corresponding to a part of the physical path among the physical paths according to one or more third environmental metrics, the one or more third environmental metrics indicating the condition of the physical path.
3. The system according to claim 1, wherein the processor is configured to: Modify the topology of the part of the first surface according to one or more third environmental metrics.
4. The system according to claim 1, wherein the processor is configured to: Modify the opacity of at least a part of the geometric 2D objects among the geometric 2D objects located at the part of the first surface according to one or more third environmental metrics.
5. The system according to claim 1, wherein the processor is configured to: Generate one or more 3D objects that satisfy the positioning heuristic at one or more corresponding positions in a second surface other than the first surface in the 3D model according to a positioning heuristic indicating the type of physical object in the physical environment.
6. The system according to claim 5, wherein the type of the physical object corresponds to at least one of the following: geographical type, climate type, or architectural type.
7. The system according to claim 1, wherein the processor is configured to: Generate one or more 3D objects that satisfy the environmental heuristic at one or more corresponding positions in the 3D model according to an environmental heuristic indicating the atmospheric condition and the weather in the physical environment.
8. The system according to claim 1, wherein the processor is configured to: Segment the regional model into a plurality of regional segments according to a block heuristic indicating the amount of computing resources, the regional model indicating the physical region includes the physical environment, and each of the regional segments corresponds to a corresponding part of the regional model, and the 3D model corresponds to the regional segments among the regional segments.
9. The system according to claim 8, wherein the processor is configured to: execute a first subset of the region segments by a first computing resource; and execute a second subset of the region segments by a second computing resource and concurrently with the execution of the first computing resource.
10. A method, comprising: The processor retrieves a camera feed of an ego object navigating within a physical environment; The processor generates a three-dimensional (3D) model based on the camera feed according to one or more first environmental metrics, the 3D model including a first surface corresponding to one or more physical paths traversing the physical environment, the one or more first environmental metrics indicating boundaries of the one or more physical paths; The processor generates one or more geometric two-dimensional (2D) objects on the first surface according to one or more second environmental metrics, the second environmental metrics indicating the one or more physical paths; The processor identifies one or more viewpoints oriented to capture corresponding portions of the 3D model according to one or more viewpoint metrics of a camera of a physical object configured to move along the one or more physical paths; The processor renders one or more simulated environment 2D images from the one or more corresponding portions of the 3D model of the physical environment, each of the one or more simulated environment 2D images corresponding to a respective one of the viewpoints; And The processor trains an artificial intelligence model based on the camera feed and the one or more simulated environment 2D images.
11. The method according to claim 10, further comprising: Modify a portion of the first surface corresponding to a portion of a physical path among the physical paths according to one or more third environmental metrics, the one or more third environmental metrics indicating a condition of the physical path.
12. The method according to claim 10, further comprising: Modify a topology of the portion of the first surface according to one or more third environmental metrics.
13. The method according to claim 10, further comprising: Modify an opacity of at least a portion of the geometric 2D objects among the geometric 2D objects located on the portion of the first surface according to one or more third environmental metrics.
14. The method according to claim 10, further comprising: Generate one or more 3D objects that satisfy the positioning heuristic at one or more corresponding positions in a second surface of the 3D model other than the first surface, according to a positioning heuristic that indicates the type of physical object in the indicated physical environment.
15. The method according to claim 14, wherein the type of the physical object corresponds to at least one of: a geographical type, a climate type, or an architectural type.
16. The method according to claim 10, further comprising: Generate one or more 3D objects that satisfy the environmental heuristic at one or more corresponding positions in the 3D model, according to an environmental heuristic that indicates the atmospheric conditions and the weather in the indicated physical environment.
17. The method according to claim 10, further comprising: Segment the regional model into multiple regional segments according to a block heuristic that indicates the amount of computing resources, the regional model indicating the physical region includes the physical environment, and each of the regional segments corresponds to a corresponding part of the regional model, the 3D model corresponding to a regional segment among the regional segments.
18. The method according to claim 17, further comprising: Execute a first subset of the regional segments by a first computing resource; And Execute a second subset of the regional segments by a second computing resource and simultaneously with the execution of the first computing resource.
19. A non-transitory computer-readable medium, the non-transitory computer-readable medium comprising one or more instructions stored thereon and executable by a processor to perform the following actions: retrieve, by the processor, a camera feed of an ego object navigating within a physical environment; Generate a three-dimensional (3D) model by the processor based on the camera feed according to one or more first environmental metrics, the 3D model including a first surface corresponding to one or more physical paths through the physical environment, the one or more first environmental metrics indicating the boundaries of the one or more physical paths; Generate one or more geometric two-dimensional (2D) objects on the first surface by the processor according to one or more second environmental metrics, the second environmental metrics indicating the one or more physical paths; Identify, by the processor, one or more viewpoints that are oriented to capture corresponding parts of the 3D model according to one or more viewpoint metrics of a camera of a physical object configured to move along the one or more physical paths; Render one or more simulated environment 2D images from the one or more corresponding parts of the 3D model of the physical environment, each of the one or more simulated environment 2D images corresponding to a corresponding viewpoint among the viewpoints; And Train an artificial intelligence model by the processor according to the camera feed and the one or more simulated environment 2D images.
20. The computer-readable medium according to claim 19, wherein the computer-readable medium further comprises one or more instructions executable by the processor to: Modify a portion of the first surface corresponding to a portion of the physical path among the physical paths according to one or more third environmental metrics, the one or more third environmental metrics indicating the condition of the physical path.