Tagging training data using large capacity navigation data
Through the navigation data and image data collected by the autonomous navigation system, the automatic labeling method and system are used to solve the problem of inefficient manual labeling in the prior art, efficient and accurate data labeling is achieved, and the training of multiple AI models is supported.
Patent Information
- Application Number
- CN202380071065.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-09-30
- Filing Date
- 2023-09-29
- Publication Date
- 2025-05-30
AI Technical Summary
In the prior art, autonomous navigation systems require manual labeling of large amounts of navigation data, which leads to inefficiency, time-consuming and susceptible to human errors, affecting the performance of the AI model.
A method and system for automatic labeling is proposed. Using the navigation data and image data collected by the autonomous navigation system, automatic labeling of navigation data is realized through high-precision trajectory and structure recovery, multi-stroke reconstruction protocol and 3D model generation.
This method can significantly reduce the time and resources required to train AI models, reduce the need for artificial intervention, improve the efficiency and accuracy of data labeling, and support various types of data labeling, such as object detection, kinematic analysis, etc.
Smart Images

Figure CN120077403A_ABST
Abstract
Description
[0001] Cross - Reference to Related Applications
[0002] This application claims priority to U.S. Provisional Application No. 63 / 377,954, filed Sep. 30, 2022, which is hereby incorporated by reference in its entirety for all purposes. Technical Field
[0003] The present disclosure generally relates to using navigation data to train artificial intelligence models. Background Art
[0004] Due to the rapid development of computer technology, autonomous navigation technologies for autonomous vehicles and robots (collectively referred to as "autos") have become ubiquitous. These advancements allow for safer and more reliable autonomous navigation of autos. Autos often use complex artificial intelligence (AI) models to identify their surrounding environment (e.g., objects occupying the surrounding environment of the auto and drivable surfaces), and to make navigation decisions.
[0005] Generating these AI models presents various technical challenges. For example, labeling training data is typically an inefficient and resource-intensive process because it requires human labelers to classify and annotate thousands of data points captured within navigation data. For example, a human reviewer can review camera footage of various vehicles navigating in a specific area, and label various objects such as lane markings, sidewalks, etc. This process is both time-consuming and expensive. Moreover, the process is prone to errors because it highly depends on the subjective knowledge and understanding of human labelers.
[0006] For the reasons above, manually labeling data is inefficient, time-consuming, and susceptible to human error, resulting in potential inaccuracies that can adversely affect the performance of AI models. Summary of the Invention
[0007] For the reasons above, there is a desire for methods and systems that can efficiently label navigation data. For example, there is a need for automated systems / methods to ingest navigation data captured by one or more autos / vehicles (e.g., camera feeds of vehicles driving in an area), and automatically label the navigation data while reducing (or sometimes eliminating) the need for human intervention.
[0008] The methods and systems discussed herein provide a tagging framework that allows for automated tagging, which is also model agnostic. Non-limiting examples of models that can be trained using the automatically tagged data are AI models related to autonomous navigation of an ego vehicle. The methods and systems discussed herein can provide a framework by which large amounts of data can be automatically tagged. Various ego vehicles are now capable of collecting data from a large number of trips (also referred to herein as “navigation sessions”), potentially involving millions of data points. Compared to other tagging methods, the framework discussed herein can significantly reduce the overall training time and the processing power required to tag such large amounts of data. The methods and systems discussed herein also provide a scalable automated tagging method.
[0009] The automated tagging process discussed herein can include three steps. The first step can involve using image data and navigation data captured by a collection of ego vehicles (e.g., multi-camera visual inertial odometry or VIO) for high-precision trajectory and structure recovery. The second step can involve performing a multi-trip reconstruction protocol, in which multiple trips (from different ego vehicles) and their corresponding data are aligned and aggregated. To achieve this, the methods and systems discussed herein can utilize a coarse alignment protocol, a pairwise matching protocol, a joint optimization protocol, and a surface refinement protocol. The multi-trip reconstruction can be finalized by a human analyst. As a result, a model (sometimes 3D) representing the environment can be created. The protocols involved in generating the model can be parallelized to improve efficiency. The third step can involve automatically tagging new trips using the generated model.
[0010] The automated tagging method discussed herein can allow an automated framework to tag various types of data that can be used for object detection models, kinematic analysis models, shape analysis models, occupancy / surface detection models, etc. Thus, the methods and systems discussed herein can be applied to various models because these methods are model agnostic. Using the methods and systems discussed can eliminate the need for human intervention when training AI models.
[0011] In an embodiment, the method includes: retrieving, by a processor, a collection of navigation data and image data from a collection of ego vehicles navigating within an environment including at least one feature; generating, by the processor, a three-dimensional (3D) model of the environment using the navigation data and image data of at least a subset of the collection of ego vehicles, the 3D model including a virtual representation of at least one feature of the environment; identifying, by the processor, machine learning labels associated with at least one feature within the image data; receiving, by the processor, second navigation data and second image data from a second ego vehicle not included in the collection of ego vehicles, the second ego vehicle navigating within the environment, the second image data including at least one feature; and automatically generating, by the processor, machine learning labels for at least one feature depicted within the second image data.
[0012] The method may further include: filtering, by a processor, a set of navigation data and image data into a subset of the set of navigation data and image data based on the trajectory of each ego within the ego set.
[0013] The method may further include: using, by a processor, a 3D model to locate a second ego using second navigation data or second image data.
[0014] At least one feature may be at least one of a drivable surface, a traffic sign, a traffic light, a lane line, or a 3D structure.
[0015] The method may further include: sending, by a processor, machine learning labels and second image data to an artificial intelligence model.
[0016] The artificial intelligence model may be an occupancy detection model.
[0017] The method may further include: executing, by a processor, the artificial intelligence model to predict machine learning labels associated with at least one feature within the image data.
[0018] The method may further include: receiving, by a processor, verification from a human reviewer regarding machine learning labels predicted by the artificial intelligence model and associated with at least one feature within the image data.
[0019] In another embodiment, a computer-readable medium includes a set of instructions that, when executed, cause a processor to: retrieve a set of navigation data and image data from a set of egos navigating within an environment that includes at least one feature; generate a three-dimensional (3D) model of the environment using the navigation data and image data of at least a subset of the set of egos, the 3D model including a virtual representation of at least one feature of the environment; identify machine learning labels associated with at least one feature within the image data; receive second navigation data and second image data from a second ego not included within the set of egos, the second ego navigating within the environment and the second image data including at least one feature; and automatically generate machine learning labels for at least one feature depicted within the second image data.
[0020] The set of instructions may further cause the processor to: filter, based on the trajectory of each ego within the set of egos, the set of navigation data and image data into a subset of the set of navigation data and image data.
[0021] The set of instructions may further cause the processor to: use the second navigation data or the second image data to locate the second ego based on the 3D model.
[0022] At least one feature is at least one of a drivable surface, a traffic sign, a traffic light, a lane line, or a 3D structure.
[0023] The instruction set also causes the processor to: send the machine learning label and the second image data to the artificial intelligence model.
[0024] The artificial intelligence model is an occupancy detection model.
[0025] The instruction set can also cause the processor to: execute the artificial intelligence model to predict a machine learning label associated with at least one feature within the image data.
[0026] The instruction set can also cause the processor to: receive verification from a human reviewer regarding the machine learning label associated with at least one feature within the image data predicted by the artificial intelligence model.
[0027] The system includes: an ego set; a processor in communication with the ego set, the processor being configured to: retrieve a set of navigation data and image data from the ego set navigating within an environment including at least one feature; use the navigation data and image data of at least a subset of the ego set to generate a three-dimensional (3D) model of the environment, the 3D model including a virtual representation of at least one feature of the environment; identify a machine learning label associated with at least one feature within the image data; receive second navigation data and second image data from a second ego not included within the ego set, the second ego navigating within the environment, the second image data including at least one feature; and automatically generate a machine learning label for at least one feature depicted within the second image data.
[0028] The processor can also be configured to: filter the set of navigation data and image data into a subset of the set of navigation data and image data according to the trajectory of each ego within the ego set.
[0029] The processor can also be configured to: use the second navigation data or the second image data to locate the second ego according to the 3D model.
[0030] The at least one feature is at least one of a drivable surface, a traffic sign, a traffic light, a lane line, or a 3D structure. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] Non-limiting embodiments of the present disclosure are described by way of example in connection with the accompanying drawings, which are schematic and not intended to be drawn to scale. Unless indicated as representing the background art, the drawings represent various aspects of the present disclosure.
[0032] Figure 1A Illustrates components of an AI-enabled data analysis system according to an embodiment.
[0033] Figure 1B Illustrates various sensors associated with an ego according to an embodiment.
[0034] Figure 1CIllustrates components of a vehicle according to an embodiment.
[0035] Figure 2 Illustrates a flowchart of a process executed in an AI-enabled data analysis system according to an embodiment.
[0036] Figure 3 Illustrates data received from a self according to an embodiment.
[0037] Figure 4 Illustrates a camera feed received from a self and a corresponding 3D model according to an embodiment.
[0038] Figure 5 Illustrates a 3D model generated by an AI-enabled data analysis system according to an embodiment.
[0039] Figures 6 - 7 Illustrates different camera feeds received from one or more selves according to an embodiment. Detailed Description
[0040] Now, reference will be made to the illustrative embodiments depicted in the drawings, and specific language will be used to describe them. However, it is to be understood that no limitation of the scope of the claims or the disclosure is thereby intended. Changes and further modifications to the features of the inventions illustrated herein, as well as additional applications of the principles of the subject matter illustrated herein, which would be obvious to those of ordinary skill in the relevant art and having the present disclosure, will be considered within the scope of the subject matter disclosed herein. Other embodiments may be used and / or other changes may be made without departing from the spirit or scope of the present disclosure. The illustrative embodiments described in the detailed description are not meant to limit the subject matter presented.
[0041] Figure 1A is a non-limiting example of components of a system in which the methods and systems discussed herein may be implemented. For example, an analysis server may train an AI model and use the trained AI model to generate occupancy datasets and / or maps for one or more selves. Figure 1A Illustrates components of an AI-enabled data analysis system 100. System 100 may include an analysis server 110a, a system database 110b, an administrator computing device 120, selves 140a to 140b (collectively referred to as (multiple) selves 140), self computing devices 141a to 141c (collectively referred to as self computing devices 141), and a server 160. System 100 is not limited to the components described herein and may include additional or other components not shown for the sake of brevity, which will be considered within the scope of the embodiments described herein.
[0042] The above components can be connected via network 130. Examples of network 130 can include, but are not limited to, private or public LANs, WLANs, MANs, WANs, and the Internet. Network 130 can include wired and / or wireless communication according to one or more standards and / or via one or more transmission media.
[0043] Communication over network 130 can be performed according to various communication protocols such as Transmission Control Protocol and Internet Protocol (TCP / IP), User Datagram Protocol (UDP), and IEEE communication protocols. In one example, network 130 can include wireless communication according to a Bluetooth specification set or another standard or proprietary wireless communication protocol. In another example, network 130 can also include communication over a cellular network, including, for example, GSM (Global System for Mobile Communications), CDMA (Code Division Multiple Access), or EDGE (Enhanced Data for Global Evolution) networks.
[0044] System 100 illustrates an example of a system architecture and components that can be used to train and execute one or more AI models such as (multiple) AI models 110c. Specifically, as Figure 1A depicted and described herein, analysis server 110a can use the methods discussed herein to train (multiple) AI models 110c using data retrieved from the selfs 140 (e.g., by using data streams 172 and 174). When (multiple) AI models 110c have been trained, each self in selfs 140 can access and execute (multiple) trained AI models 110c. For example, vehicle 140a with a self computing device 141a can send its camera feed to (multiple) trained AI models 110c and can determine the occupancy status of its surrounding environment (e.g., data stream 174). Moreover, data ingested and / or predicted by (multiple) AI models 110c with respect to selfs 140 (at inference time) can also be used to improve (multiple) AI models 110c. Thus, system 100 depicts a continuous loop that can periodically improve the accuracy of (multiple) AI models 110c. Moreover, system 100 depicts a loop in which, in addition to the inference phase, data received from selfs 140 can also be used in the training phase.
[0045] The analysis server 110a can be configured to collect, process, and analyze navigation data (e.g., images captured during navigation) and various sensor data collected from the vehicle 140. The collected data can then be processed and prepared into a training dataset. The training dataset can then be used to train one or more AI models, such as the AI model 110c. The analysis server 110a can also be configured to collect visual data from the vehicle 140. Using the AI model 110c (trained using the methods and systems discussed herein), the analysis server 110a can generate a dataset and / or an occupancy map for the vehicle 140. The analysis server 110a can display the occupancy map on the vehicle 140 and / or send the occupancy map / dataset to the vehicle computing device 141, the administrator computing device 120, and / or the server 160.
[0046] In Figure 1A , the AI model 110c is illustrated as a component of the system database 110b, but the AI model 110c can be stored in a different or separate component, such as a cloud storage device or any other data repository accessible to the analysis server 110a.
[0047] The analysis server 110a can also be configured to display an electronic platform that illustrates various training attributes for training the AI model 110c. The electronic platform can be displayed on the administrator computing device 120 such that an analyst can monitor the training of the AI model 110c. An example of an electronic platform generated and hosted by the analysis server 110a can be a web-based application or website configured to display the training dataset collected from the vehicle 140 and / or the training status / metrics of the AI model 110c.
[0048] The analysis server 110a can be any computing device that includes a processor and a non-transitory machine-readable storage device capable of performing the various tasks and processes described herein. Non-limiting examples of such computing devices can include workstation computers, laptop computers, server computers, etc. Although the system 100 includes a single analysis server 110a, the system 100 can include any number of computing devices operating in a distributed computing environment, such as a cloud environment.
[0049] The vehicle 140 can represent various electronic data sources that send data associated with its previous or current navigation session to the analytics server 110a. The vehicle 140 can be any device configured for navigation, such as a car 140a and / or a truck 140c. The vehicle 140 is not limited to being a vehicle and can also include robotic devices. For example, the vehicle 140 can include a robot 140b, which can represent a general-purpose, bipedal, autonomous humanoid robot capable of navigating various terrains. The robot 140b can be equipped with software for achieving balance, navigation, perception, or interaction with the physical world. The robot 140b can also include various cameras configured to send visual data to the analytics server 110a.
[0050] Although referred to herein as a "vehicle", the vehicle 140 may or may not be an autonomous device configured for autonomous navigation. For example, in some embodiments, the vehicle 140 can be controlled by a human operator or a remote processor. The vehicle 140 can include various sensors, such as Figure 1B the sensors depicted in. The sensors can be configured to collect data as the vehicle 140 navigates various terrains (e.g., roads). The analytics server 110a can collect the data provided by the vehicle 140. For example, the analytics server 110a can obtain navigation session and / or road / terrain data (e.g., an image of the vehicle 140 navigating on a road) from various sensors, such that the collected data is ultimately used by the AI model 110c for training purposes.
[0051] As used herein, a navigation session corresponds to the journey of the vehicle 140's travel route, regardless of whether the journey is autonomous or human-controlled. In some embodiments, the navigation session can be used for data collection and model training purposes. However, in some other embodiments, the vehicle 140 can refer to a vehicle purchased by a consumer, and the purpose of the journey can be classified as daily use. A navigation session can begin when the vehicle 140 moves more than a threshold distance (e.g., 0.1 miles, 100 feet) or at a threshold rate (e.g., more than 0 mph, more than 1 mph, more than 5 mph) from a non-moving position. A navigation session can end when the vehicle 140 returns to a non-moving position and / or is turned off (e.g., when the driver exits the vehicle).
[0052] The self 140 can represent a collection of selves monitored by the analytics server 110a for training the AI model(s) 110c. For example, a driver of the vehicle 140a can authorize the analytics server 110a to monitor data associated with their respective vehicle. As a result, the analytics server 110a can utilize the various methods discussed herein to collect sensor / camera data and generate a training dataset to train the AI model(s) 110c accordingly. The analytics server 110a can then apply the trained AI model(s) 110c to analyze data associated with the self 140 and predict an occupancy map for the self 140. Moreover, additional / ongoing data associated with the self 140 can also be processed and added to the training dataset so that the analytics server 110a can recalibrate the AI model(s) 110c accordingly. Thus, the system 100 depicts a loop in which navigation data received from the self 140 can be used to train the AI model(s) 110c. The self 140 can include a processor that executes the trained AI model(s) 110c for navigation purposes. During navigation, the self 140 can collect additional data about its navigation session, and this additional data can be used to calibrate the AI model(s) 110c. That is, the self 140 represents a self that can be used to train, execute / use, and recalibrate the AI model(s) 110c. In a non-limiting example, the self 140 represents vehicles purchased by customers that can use the AI model(s) 110c for autonomous navigation while improving the AI model(s) 110c.
[0053] The self 140 can be equipped with various technologies that allow the self to collect data from its surrounding environment and (possibly) navigate autonomously. For example, the self 140 can be equipped with an inference chip to run autonomous driving software.
[0054] Various sensors for each self 140 can monitor the data collected associated with different navigation sessions and send that data to the analytics server 110a. Figures 1B - 1C A block diagram of sensors integrated within the self 140 according to an embodiment is illustrated. The quantity and location of each sensor discussed with respect to Figures 1B - 1C can depend on the type of self discussed in Figure 1A . For example, the robot 140b can include different sensors than the vehicle 140a or the truck 140c. For example, the robot 140b may not include an airbag activation sensor 170q. Moreover, the sensors of the vehicle 140a and the truck 140c can be positioned differently than Figure 1C illustrated.
[0055] As discussed herein, the various sensors integrated within each ego 140 can be configured to measure various data associated with each navigation session. The analytics server 110a can periodically collect the data monitored and collected by these sensors, where the data is processed according to the methods described herein and is used to train the AI model 110c and / or execute the AI model 110c to generate an occupancy map.
[0056] The ego 140 can include a user interface 170a. The user interface 170a can refer to the user interface of an ego computing device (such as Figure 1A the ego computing device 141 in). The user interface 170a can be implemented as a display screen integrated with or coupled to the interior of the vehicle, a head-up display, a touch screen, etc. The user interface 170a can include input devices such as a touch screen, a knob, a button, a keyboard, a mouse, a gesture sensor, a steering wheel, etc. In various embodiments, the user interface 170a can be adapted to provide user input (such as as a signal and / or sensor information) to other devices or sensors of the ego 140 (such as Figure 1B the sensors illustrated) (such as the controller 170c).
[0057] The user interface 170a can also be implemented with one or more logic devices that can be adapted to execute instructions, such as software instructions, to implement any of the various processes and / or methods described herein. For example, the user interface 170a can be adapted to form a communication link, send and / or receive communications (such as sensor signals, control signals, sensor information, user input, and / or other information), or execute various other processes and / or methods. In another example, a driver can use the user interface 170a to control the temperature of the ego 140 or activate its features (such as the autonomous driving or steering system 170o). Thus, the user interface 170a can combine with the other sensors described herein to monitor and collect driving session data. The user interface 170a can also be configured to display various data generated / predicted by the analytics server 110a and / or the AI model 110c.
[0058] The orientation sensor 170b can be implemented as a compass, a float, an accelerometer, and / or one or more of other digital or analog devices capable of measuring the orientation of the ego 140 (such as the magnitude and direction of roll, pitch, and / or yaw relative to one or more reference orientations such as gravity and / or magnetic north). The orientation sensor 170b can be adapted to provide heading measurements to the ego 140. In other embodiments, the orientation sensor 170b can be adapted to provide roll, pitch, and / or yaw rates to the ego 140 using a time series of orientation measurements. The orientation sensor 170b can be positioned and / or adapted to make orientation measurements relative to a specific coordinate system of the ego 140.
[0059] The controller 170c can be implemented as any suitable logic device (such as a processing device, a microcontroller, a processor, an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), a memory storage device, a memory reader, or other device or combination of devices), which can be adapted to execute, store, and / or receive appropriate instructions, such as software instructions implementing a control loop for controlling various operations of the self 140. Such software instructions can also implement methods for processing sensor signals, determining sensor information, providing user feedback (such as via the user interface 170a), querying the device for operating parameters, selecting operating parameters for the device, or performing any of the various operations described herein.
[0060] The communication module 170e can be implemented as any wired and / or wireless interface configured to transmit sensor data, configuration data, parameters, and / or other data and / or signals to Figure 1A any of the features shown (such as the analytics server 110a). As described herein, in some embodiments, the communication module 170e can be implemented in a distributed manner such that portions of the communication module 170e are implemented within Figure 1B one or more of the elements shown and the sensor. In some embodiments, the communication module 170e can delay the transmission of sensor data. For example, when the self 140 does not have network connectivity, the communication module 170e can store the sensor data in a temporary data storage device and send the sensor data when the self 140 is identified as having appropriate network connectivity.
[0061] The speed sensor 170d can be implemented as an electronic pitot tube, a metering gear or wheel, a water speed sensor, a wind speed sensor, a wind rate sensor (such as direction and magnitude), and / or other devices capable of measuring or determining the linear speed of the self 140 (such as in the surrounding medium and / or aligned with the longitudinal axis of the self 140) and providing such measurement as a sensor signal (which can be transmitted to various devices).
[0062] The gyroscope / accelerometer 170f can be implemented as one or more electronic sextants, semiconductor devices, integrated chips, accelerometer sensors, or other systems or devices capable of measuring the angular velocity / acceleration and / or linear acceleration of the self 140 (such as direction and magnitude) and providing such measurement as a sensor signal (which can be transmitted to various devices, such as the analytics server 110a). The gyroscope / accelerometer 170f can be positioned and / or adapted to make such measurements relative to a particular coordinate system of the self 140. In various embodiments, the gyroscope / accelerometer 170f can be associated with Figure 1BThe other elements depicted are implemented in a common housing and / or module to ensure a common frame of reference or known transformations between frames of reference.
[0063] The Global Navigation Satellite System (GNSS) 170h may be implemented as a Global Positioning Satellite receiver and / or another device capable of determining the absolute and / or relative position of the vehicle body 140 based on wireless signals received, for example, from space and / or terrestrial sources and capable of providing measurements such as sensor signals (which may be transmitted to various devices). In some embodiments, the GNSS 170h may be adapted to determine the rate, velocity, and / or yaw rate of the vehicle body 140 (e.g., using a time series of position measurements), such as the yaw component of the absolute rate and / or angular rate of the vehicle body 140.
[0064] The temperature sensor 170i may be implemented as a thermistor, an electrical sensor, an electrical thermometer, and / or other devices capable of measuring the temperature associated with the vehicle body 140 and providing such measurement as a sensor signal. The temperature sensor 170i may be configured to measure the ambient temperature associated with the vehicle body 140, such as the cockpit or dashboard temperature, and for example, this ambient temperature may be used to estimate the temperature of one or more elements of the vehicle body 140.
[0065] The humidity sensor 170j may be implemented as a relative humidity sensor, an electrical sensor, an electrical relative humidity sensor, and / or another device capable of measuring the relative humidity associated with the vehicle body 140 and providing such measurement as a sensor signal.
[0066] The steering sensor 170g may be adapted to physically adjust the heading of the vehicle body 140 according to one or more control signals provided by a logic device (such as the controller 170c) and / or user input. The steering sensor 170g may include one or more actuators and control surfaces (such as a rudder or other types of steering or trim mechanisms) of the vehicle body 140, and may be adapted to physically adjust the control surface to various positive and / or negative steering angles / positions. The steering sensor 170g may also be adapted to sense the current steering angle / position of such a steering mechanism and provide such measurement.
[0067] The propulsion system 170k may be implemented as a propeller, a turbine, or other thrust-based propulsion systems, mechanical wheeled and / or tracked propulsion systems, wind / sail-based propulsion systems, and / or other types of propulsion systems that may be used to provide power to the vehicle body 140. The propulsion system 170k may also monitor the direction of the power and / or thrust of the vehicle body 140 relative to the reference coordinate system of the vehicle body 140. In some embodiments, the propulsion system 170k may be coupled to the sensor 170g and / or integrated with the steering sensor 170g.
[0068] The passenger restraint sensor 170l can monitor seat belt detection and locking / unlocking assemblies, as well as other passenger restraint subsystems. The passenger restraint sensor 170l can include various environmental and / or status sensors, actuators, and / or other devices that facilitate the operation of safety mechanisms associated with the operation of the vehicle 140. For example, the passenger restraint sensor 170l can be configured to receive motion and / or status data from Figure 1B the other sensors depicted. The passenger restraint sensor 170l can determine whether a safety measure (e.g., seat belt) is being used.
[0069] As Figure 1C depicted, the camera 170m can refer to one or more cameras integrated within the vehicle 140, and can include multiple cameras integrated (or retrofitted) into the vehicle 140. The camera 170m can be an internal or external-facing camera of the vehicle 140. For example, as Figure 1C depicted, the vehicle 140 can include one or more internal-facing cameras that can monitor and collect footage of the passengers of the vehicle 140. The vehicle 140 can include eight external-facing cameras. For example, the vehicle 140 can include a front camera 170m-1, front side cameras 170m-2, 170m-3, rear side cameras 170m-4 on each front fender, cameras 170m-5 on each side (e.g., integrated within the B-pillar), and a rear camera 170m-6.
[0070] Referring Figure 1B , the radar 170n and ultrasonic sensors 170p can be configured to monitor the distance of the vehicle 140 to other objects, such as other vehicles or immovable objects (e.g., trees or garage doors). The vehicle 140 can also include an autonomous driving or steering system 170o configured to autonomously navigate the vehicle 140 using data collected via various sensors, such as the radar 170n, speed sensors 170d, and / or ultrasonic sensors 170p.
[0071] Thus, the autonomous driving or steering system 170o can analyze various data collected by one or more of the sensors described herein to identify driving data. For example, the autonomous driving or steering system 170o can calculate the risk of a forward collision based on the speed of the vehicle 140 and its distance to another vehicle on the road. The autonomous driving or steering system 170o can also determine whether the driver is touching the steering wheel. The autonomous driving or steering system 170o can send the analyzed data to various features discussed herein, such as an analysis server.
[0072] The airbag activation sensor 170q can predict or detect a collision and cause the activation or deployment of one or more airbags. The airbag activation sensor 170q can send data regarding airbag deployment, including data associated with the event that caused the deployment.
[0073] Referring again to Figure 1A , the administrator computing device 120 can represent a computing device operated by a system administrator. The administrator computing device 120 can be configured to display data retrieved or generated by the analytics server 110a (such as various analytics metrics and risk scores), where the system administrator can monitor the various models utilized by the analytics server 110a, review feedback, and / or facilitate the training of the AI model(s) 110c maintained by the analytics server 110a.
[0074] (The) ego 140 can be any device configured to navigate various routes, such as a vehicle 140a or a robot 140b. As discussed with respect to Figures 1B - 1C , the ego 140 can include various telemetry sensors. The ego 140 can also include an ego computing device 141. Specifically, each ego can have its own ego computing device 141. For example, a truck 140c can have an ego computing device 141c. For simplicity, the ego computing devices are collectively referred to as the ego computing device(s) 141. The ego computing device 141 can control content presentation on the infotainment system of the ego 140, process commands associated with the infotainment system, aggregate sensor data, manage communication of data to electronic data sources, receive updates, and / or send messages. In one configuration, the ego computing device 141 communicates with an electronic control unit. In another configuration, the ego computing device 141 is an electronic control unit. The ego computing device 141 can include a processor and a non-transitory machine-readable storage medium capable of performing the various tasks and processes described herein. For example, the AI model(s) 110c described herein can be stored and executed (or directly accessed) by the ego computing device 141. Non-limiting examples of the ego computing device 141 can include vehicle multimedia and / or display systems.
[0075] In one example of how the AI model(s) 110c can be trained, the analytics server 110a can collect data from the ego 140 to train the AI model(s) 110c. Before executing the AI model(s) 110c to generate / predict an occupancy dataset, the analytics server 110a can use various methods to train the AI model(s) 110c. Training allows the AI model(s) 110c to ingest data from one or more cameras of one or more egos 140 (without receiving radar data) and predict occupancy data for the surrounding environment of the ego. The operations described in this example can be performed by the Figure 1A and 1Bexecuted by any number of computing devices (e.g., the processors of the vehicle 140) operating in the distributed computing system described in
[0076] To train the AI model(s) 110c, the analytics server 110a may communicate with one or more vehicles in the vehicle 140 that drive a specific route. For example, one or more vehicles may be selected for training purposes. One or more vehicles may drive a specific route autonomously or via a human operator. As a result of the navigation of one or more vehicles, various data points may be collected and used for training purposes. For example, while driving, the vehicle 140 may use one or more of its sensors (including one or more cameras) to generate navigation session data. For example, one or more vehicles 140 equipped with various sensors may navigate a designated route. As one or more vehicles 140 traverse the terrain, their sensors may capture continuous (or periodic) data of their surrounding environment. The sensors may indicate the occupancy status of the surrounding environment of one or more vehicles 140. For example, the sensor data may indicate various objects with mass in the surrounding environment of one or more vehicles 140 as they navigate their route.
[0077] The analytics server 110a may use the data collected from the vehicle 140 (e.g., the camera feeds received from the vehicle 140) to generate a training data set. The training data set may indicate the occupancy status of different voxels within the surrounding environment of one or more vehicles 140. As used herein in some embodiments, a voxel is a three-dimensional pixel that forms the building blocks of the surrounding environment of one or more vehicles 140. Within the training data set, each voxel may encapsulate sensor data that indicates whether a mass is identified for that specific voxel. As used herein, a mass may indicate or represent any object identified using a sensor. For example, in some embodiments, the vehicle 140 may be equipped with sensors that can identify masses near the vehicle 140.
[0078] In some embodiments, the training data set may include data received from the cameras of the vehicle 140. The data received from the camera(s) may have a set of data points, where each data point corresponds to the position and image attributes of at least one voxel of the space surrounding the vehicle 140. The training data set may also include 3D geometric data to indicate whether the voxels of the surrounding environment of one or more vehicles 140 are occupied by an object with mass.
[0079] In operation, as one or more vehicles 140 navigate, their sensors collect data and send the data to the analytics server 110a, as depicted by the data stream 172.
[0080] In some embodiments, one or more avatars 140 may include one or more high-resolution cameras that capture a continuous visual data stream from the surroundings of the one or more avatars 140 as the one or more avatars 140 navigate through a route. The analysis server 110a may then use the camera feed to generate a second data set that includes visual elements / depictions of different voxels of the surroundings of the one or more avatars 140.
[0081] In operation, as the one or more avatars 140 navigate, their cameras collect data and send the data to the analysis server 110a, as depicted by data stream 172. For example, the avatar computing device 141 may use data stream 172 to send image data to the analysis server 110a.
[0082] The analysis server 110a may use the first data set and the second data set to train an AI model, whereby the AI model 110c trains itself using the respective positions of each data point, associating each data point in the first set of data points with the corresponding data point in the second set of data points, wherein after being trained, the AI model 110c is configured to receive a camera feed from a new avatar 140 and predict the occupancy state of at least one voxel of the camera feed.
[0083] Using the first data set and the second data set, the analysis server 110a may train the AI model(s) 110c such that the AI model(s) 110c can associate different visual attributes of a voxel (within the camera feed in the second data set) with the occupancy state of that voxel (within the first data set). In this way, after being trained, the AI model(s) 110c can receive a camera feed (e.g., from a new avatar 140) without receiving sensor data and then determine the occupancy state of each voxel for the new avatar 140.
[0084] The analysis server 110a may generate a training data set that includes the first data set and the second data set. The analysis server 110a may use the first data set as ground truth. For example, the first data set may indicate the different positions of voxels and their occupancy states. The second data set may include visual (e.g., camera feed) depictions of the same voxels. Using the first data set, the analysis server 110a may label the data such that the data record(s) associated with each voxel corresponding to an object are indicated as having a positive occupancy state.
[0085] The occupancy status of different voxels can be marked automatically and / or manually. For example, in some embodiments, the analysis server 110a can use human reviewers to mark the data. For example, as discussed herein, the camera feeds from one or more cameras of a vehicle can be presented to a human reviewer on an electronic platform for marking. Additionally or alternatively, the AI model(s) 110c can ingest the entire data, where the AI model(s) 110c identify the corresponding voxels, analyze the first digital map, and associate the image(s) of each voxel with its corresponding occupancy status.
[0086] Using the methods and systems discussed herein, the analysis server 110a can automatically mark the data, such that the training process for the AI model(s) 110c is performed more efficiently.
[0087] Using the ground truth, the AI model(s) 110c can be trained to analyze the visual elements of each voxel and associate them with whether the voxel is occupied by a mass. Thus, the AI model 110c can retrieve the occupancy status of each voxel (using the first dataset) and use this information as the ground truth. The AI model(s) 110c can also retrieve the visual attributes of the same voxels using the second dataset.
[0088] In some embodiments, the analysis server 110a can use a supervised training method. For example, using the ground truth and the received visual data, the AI model(s) 110c can train itself such that it can predict the occupancy status for a voxel using only the image of that voxel. As a result, when trained, the AI model(s) 110c can receive the camera feed, analyze the camera feed, and determine the occupancy status for each voxel within the camera feed (without the need to use radar).
[0089] The analysis server 110a can feed a series of training datasets into the AI model(s) 110c and obtain a set of predicted outputs (e.g., predicted occupancy status). The analysis server 110a can then compare the predicted data with the ground truth data to determine the differences and train the AI model(s) 110c by adjusting the internal weights and parameters of the AI model 110c proportional to the determined differences according to a loss function. The analysis server 110a can train the AI model(s) 110c in a similar manner until the predictions of the trained AI model 110c are accurate to a certain threshold (e.g., recall or precision).
[0090] Additionally or alternatively, the analytics server 110a may use an unsupervised method in which the training data set is not labeled. Since labeling the data within the training data set can be time-consuming and may require excessive computing power, the analytics server 110a may utilize unsupervised training techniques to train the AI model 110c. In some embodiments, the analytics server 110a may use the methods discussed herein to automatically label the data instead of the unsupervised method.
[0091] After the AI model 110c is trained, the ego 140 may use it to predict occupancy data for the surroundings of one or more egos 140. For example, the AI model(s) 110c may divide the surroundings of the ego into different voxels and predict the occupancy status for each voxel. In some embodiments, the AI model(s) 110c (or the analytics server 110a that uses the data predicted by the AI model 110c) may generate an occupancy map or occupancy network representing the surroundings of one or more egos 140 at any given time.
[0092] In another example of how the AI model(s) 110c may be used, after training the AI model(s) 110c, the analytics server 110a (or the local chip of the ego 140) may collect data from the ego (e.g., one or more egos among the egos 140) to predict an occupancy data set for one or more egos 140. This example describes how the AI model(s) 110c may be used to predict occupancy data for one or more egos 140 in real-time or near real-time. This configuration may have a processor that executes the AI model, such as the analytics server 110a. However, one or more actions may be performed locally, for example, via a chip located within one or more egos 140. In operation, the AI model(s) 110c may be executed locally via the ego 140 such that the results may be used for autonomous navigation of the ego itself.
[0093] The processor may input image data of the space around the ego object 140 into the AI model 110c using a camera of the ego object 140. The processor may collect and / or analyze data received from various cameras (e.g., outward-facing cameras) of one or more egos 140. In another example, the processor may collect and aggregate footage recorded by one or more cameras of the ego 140. The processor may then send the footage to the AI model(s) 110c trained using the methods discussed herein.
[0094] The processor may predict the occupancy attributes of multiple voxels by executing the AI model 110c. The AI model(s) 110c may use the methods discussed herein to predict the occupancy status for different voxels surrounding one or more egos 140 using the received image data.
[0095] The processor can generate a data set based on multiple voxels and their corresponding occupancy attributes. The analysis server 110a can generate a data set including its occupancy status according to the corresponding coordinate values of different voxels. The data set can be a queryable data set available for sending the predicted occupancy status to different software modules.
[0096] In operation, one or more autosomes 140 can collect image data from their cameras and send the image data to the processor (locally placed on one or more autosomes 140) and / or the analysis server 110a, as depicted by the data stream 172. The processor can then execute the AI model(s) 110c to predict the occupancy data for one or more autosomes 140. If the prediction is performed by the analysis server 110a, then the occupancy data can be sent to one or more autosomes 140 using the data stream 174. If the processor is locally placed within one or more autosomes 140, then the occupancy data is sent to the autosome computing device 141 ( Figure 1A not shown in the figure).
[0097] Using the methods discussed herein, the training of the AI model(s) 110c can be performed such that the execution of the AI model(s) 110c can be locally executed (at inference time) on any autosome in the autosome 140. The collected data (e.g., navigation data collected during the navigation of the autosome 140, such as image data of the journey) can then be fed back into the AI model(s) 110c such that additional data can improve the AI model(s) 110c.
[0098] Figure 2 FIG. illustrates a flowchart of a method 200 executed in an AI-enabled visual data analysis system according to an embodiment. The method 200 can include step 210 to step 270. However, other embodiments can include additional or alternative steps, or one or more steps can be omitted. The method 200 is executed by an analysis server (e.g., a computer similar to the analysis server 110a). However, one or more steps of the method 200 can be executed by any number of computing devices (e.g., the processors of the autosome 140 and / or the autosome computing device 141) operating in the Figures 1A - 1C distributed computing system described in the figure. For example, one or more computing devices of the autosome can locally execute Figure 2 some or all of the steps described in the figure.
[0099] Using the methods discussed herein, analytics server 110a can collect data from the vehicles 140 and generate an initial inference that indicates initial labels for the various features included in the data. For example, the analytics server can collect the camera feeds of each vehicle and determine indications of the various features depicted within the camera feeds (such as trees, buildings, traffic lights, or traffic signs). The initial inference can be displayed on a platform where a human reviewer can confirm / validate the initial inference in view of the received camera footage. When the initial inference is validated, analytics server 110a can automatically label new footage received from the vehicles 140. The labeled data can then be sent to the AI models 110c.
[0100] Figure 2 A flowchart of a method that can be used to automatically label data for training one or more artificial intelligence models, such as the AI models 110c, is illustrated. Using the methods and systems discussed herein, an analytics server can ingest image data (such as a camera feed from the vehicle's surroundings) and automatically label the various features depicted within the camera feed with little to no human intervention.
[0101] Method 200 is described as being performed by an analytics server. However, one or more of the steps of method 200 can be performed by other processors. For example, step 210 can be performed locally by the vehicle computing device. Then, the other steps can be performed by a central processor (such as in the cloud).
[0102] In step 210, the analytics server can retrieve navigation data and image data from a collection of vehicles navigating within an environment that includes at least one feature. The analytics server can communicate with the vehicle computing devices. As discussed herein, the vehicle computing devices can communicate with the various sensors of the vehicle and collect sensor data. The vehicle computing devices can then send the sensor data to the analytics server.
[0103] As used herein, navigation data can include any data related to environmental navigation collected and / or retrieved by the vehicle (either autonomously or via a human operator). As discussed herein, vehicles can rely on various sensors and technologies to collect comprehensive navigation data such that they can navigate autonomously within and through various environments. Thus, vehicles can collect a wide variety of information from within the environments they navigate. Thus, navigation data can include data collected by Figures 1A - 1CAny data collected by any of the sensors discussed in . Additionally, the navigation data can include any data that is extracted or analyzed using any of the sensor data, including high-definition maps, trajectory information, etc. Non-limiting examples of navigation data can include visual inertial odometry (VIO), inertial measurement unit (IMU) data, and / or any data that can indicate the position and trajectory of the ego.
[0104] In some embodiments, the navigation data can be anonymized. Thus, the analytics server may not receive an indication of which data set / data point belongs to which ego within the ego set. Anonymization can be performed locally on the ego, e.g., via the ego computing device. Alternatively, anonymization can be performed by another processor before the analytics server receives the data.
[0105] In some embodiments, the ego processor / computing device can send only data strings without sending any ego identification data that would allow the analytics server and / or any other processor to determine which ego generated which data set. As a result, the analytics server can simply receive the image data (camera feed) of the ego, as well as the VIO data and ego IMU data captured by one or more sensors of the ego.
[0106] The analytics server can communicate with the processors of each ego in the set of egos navigating within various environments. Then, the analytics server can collect navigation data (in real-time, near real-time, or at various other frequencies) from that set of egos.
[0107] In addition to retrieving the navigation data, the analytics server can also retrieve image data (e.g., camera feed or video clip) of the set of egos as they navigate within different environments. The image data can include various features located within the environment. As used herein, features within the environment can refer to any physical item located within the environment in which one or more egos are navigating. Thus, the features can correspond to natural or man-made objects, whether traffic-related or not. Non-limiting examples of features can include lane lines or other traffic markings, road / traffic signs, traffic lights, sidewalk markings, buildings, etc.
[0108] Then, the analytics server can aggregate the data and preprocess the data (e.g., deduplicate the data and / or denoise the data). Additionally, the analytics server can analyze the received raw data to identify one or more attributes of the navigation itself. For example, the navigation data can be analyzed to determine the trajectory of the ego. As described herein, the aggregated data can be used to generate a 3D model of the environment itself.
[0109] Now refer to Figure 3, the data 300 visually represents navigation and image data retrieved from the ego while the ego navigates within the environment. The data 300 may include image data 302, 304, 306, 308, 312, 314, 316, and 318 (collectively referred to as the camera feed 301). The camera feed 301 may include image data captured by each of eight cameras of the ego as depicted by Figure 1C . Thus, as the ego navigates within the environment, eight different cameras collect image data of the ego's surrounding environment (e.g., the environment). The camera feed 301 may depict various features located within the environment. For example, the image data 302 depicts various lane lines (e.g., dashed lines dividing four lanes) and trees. The image data 304 depicts the same lane lines and trees from a different angle. The image data 306 depicts the same lane lines from yet another angle. Additionally, the image 306 also depicts a building on the other side of the street. The image data 308, 312, 318, 316, and 314 depict the same lane lines. However, some of these image data also depict additional features, such as traffic lights depicted in the image data 314, 308, and / or 312.
[0110] The navigation data 310 represents the trajectory of the ego from which the Figure 3 depicted image data was collected. The trajectory may be a two-dimensional or three-dimensional trajectory of the ego that is calculated using sensor data retrieved from the ego. In some embodiments, various navigation data may be used to determine the trajectory of the ego.
[0111] As discussed herein, the ego may be equipped with various position tracking data. Using this data, the ego's processor and / or the analysis server may generate a trajectory for the ego's travel path. Each image within the camera feed 301 may also include a timestamp that may correspond to the timestamp of the ego's trajectory calculated and depicted within the navigation data 310. Thus, the analysis server may identify up to eight images from different cameras of the ego at each timestamp and location within the ego's navigation within the environment.
[0112] Referring again to Figure 2 , at step 220, the analysis server may use the navigation data and image data of at least a subset of the ego set to generate a three-dimensional (3D) model of the environment, the 3D model including a virtual representation of at least one feature of the environment.
[0113] The analysis server can use the data retrieved in step 210 to generate a 3D model of the environment. The analysis server can first filter the navigation data and image data using the position / trajectory of each ego vehicle, such that the retrieved data is limited to a specific environment. The analysis server can then generate a 3D model of the environment. Each position within the 3D model can correspond to one or more images (or videos) of that position within the environment that have been captured from the cameras of one or more ego vehicles.
[0114] The 3D model can be similar to a high-definition map that includes various features depicted in the data retrieved from step 210. Before generating the 3D model, the analysis server can perform one or more computer modeling techniques to identify various features of the environment, such as road surfaces, objects, traffic features, etc. For example, using the navigation data and camera feeds received from the set of ego vehicles, the analysis server can perform occupancy or surface networks to determine the occupancy status of various plots within the environment. The analysis server can also perform various semantic analysis protocols, image segmentation protocols, object recognition protocols, and various other modeling techniques to identify various features located within the environment, such as buildings, lane markings, traffic lights, road signs, etc.
[0115] The ego computing device and / or the analysis server can be equipped with a VIO system that can retrieve the 3D trajectories of the ego vehicle and each camera within each ego vehicle. Using this data, the analysis server can generate (sometimes sparse) 3D structures at the time of capture of each camera (e.g., from the perspective of each camera). The analysis server can also generate a complete 3D (e.g., six degrees of freedom considering rotation and translation) using the camera feeds and the ego vehicle trajectories. Once this data has been retrieved from the ego vehicles and / or generated by the analysis server, the analysis server can identify multiple drives or navigation sessions (and their corresponding data) from similar environments (e.g., multiple ego vehicles navigating within the same neighborhood).
[0116] The analysis server can then use the data associated with different navigation sessions (VIO, odometry, and other data) to group similar navigation data. For example, the analysis server can use the retrieved image data and cluster various navigation sessions (camera feeds of various trips) based on their similarity (e.g., navigation within the same environment). The analysis server can then align the image data from different ego vehicles within the same trip cluster. That is, the 3D models generated based on each ego vehicle within the cluster can be aligned with other 3D models generated by other ego vehicles navigating within the same environment. Using all of the image data, the analysis server can reconstruct a 3D representation of the environment, referred to herein as the 3D model. The 3D model can also include a mesh surface representation of the driving surface and a representation of various vertical structures / features (such as buildings or signs).
[0117] To generate a 3D model, the analysis server can first filter various image data (navigation clips or camera feeds) received from different vehicles. Even within a vehicle cluster and / or a navigation cluster, the analysis server can identify and eliminate overlapping image data. In this way, redundancy is eliminated. The analysis server can filter the image data into non-overlapping navigation clips. For example, if two video clips of driving within the same street and the same lane are identified, the analysis server can use only one of the video clips when generating the 3D model.
[0118] After the image data is filtered, the analysis server can perform a coarse alignment protocol. The analysis server can use VIO data associated with the image data of different vehicles to find similarities between different image features. Once the shared features in two video clips are identified, some initial visual alignment can be performed on the two video clips using the shared features. This alignment may be rough as it provides a preliminary alignment of the environments navigated by the two vehicles.
[0119] After the initial alignment of the video clips, the analysis server can identify various video clips having a common feature. For example, the analysis server can identify ten navigation sessions involving (passing through) a specific crosswalk. The trips may not originate from the same location and may not share the same destination. However, at least a portion of the trips can be used because these portions share data from the same environment. That is, the camera feeds of these trips include the same feature (the crosswalk).
[0120] Once the aligned navigation sessions are identified, the analysis server can perform a pairwise matching protocol. Then, the analysis server can compare the trips that have been roughly aligned and determine additional features common within the data captured by each vehicle. For example, the analysis server can compare different frames of the camera feeds of each of the two roughly aligned trips so that the analysis server can identify matching features. The analysis server can identify key points (e.g., image data having unique and distinctive textures) within the image data captured from each trip and attempt to match a key point within a frame captured by the camera of the first vehicle with another key point within a frame captured by the camera of the second vehicle. The pairwise matching protocol allows the analysis server to correct camera feeds from different vehicles that captured images of the same feature (e.g., the same traffic light) but from different angles. Although these images may not appear similar to each other (because they were captured from different angles), they may share key points that can be matched.
[0121] The analysis server can then perform various optimization protocols. In some embodiments, the analysis server can perform a pose graph optimization protocol. The analysis server can use the pose graph optimization protocol to optimize the trajectories for different trips.
[0122] In some embodiments, the trajectories of the two avatars may not match because each navigation session is different. As a result, the 3D structures seen by each avatar are slightly different (even though they are views of the same structure). The optimization protocol executed by the analysis server reduces / minimizes these differences. When adjusting the six degrees of freedom pose of the camera, the analysis server can determine the projections of features near the avatar. For example, how a building is depicted within the image can vary (after adjustment). During optimization, this variation can be minimized / reduced. Via optimization, the analysis server can use different camera feeds from different avatars captured by cameras pointing in slightly different directions, as long as each camera feed includes the same key features of the same object within the physical environment. As a result, an aggregated camera feed can be used to generate a 3D model of the object. In some embodiments, the analysis server can use a large-scale non-linear least squares optimization protocol.
[0123] In some embodiments, by adjusting the 3D pose of each camera, the analysis server can align the rays for the cameras. As used herein, a ray refers to a beam of light emitted from the camera center passing through a detection (or observation) point. By adjusting the 3D pose of the camera, the camera ray can be moved. The analysis server can identify all the rays that include the corresponding 3D points. For example, a feature of the environment (such as a traffic light) can be selected, and all the camera rays pointing to the traffic light (at some points) can be identified, and their corresponding rays can be adjusted and optimized to determine the 3D coordinates of the traffic light.
[0124] In some embodiments, the analysis server can execute a bundle adjustment protocol to optimize the 3D pose of each camera and the 3D position of the key features common within the image data received from different avatars.
[0125] Now referring to Figure 4 , a non-limiting example of a 3D model and its corresponding camera feeds is illustrated. As depicted, the image data 402 to 410 represent camera feeds captured by the cameras of an avatar navigating within a street. Using the camera feeds in combination with other navigation data received from the avatar, the analysis server can generate a 3D model 412. The 3D model 412 can indicate the position of the avatar (414) driving within the environment 416. The environment 416 can be a 3D representation that includes features captured as a result of analyzing the avatar's camera feeds and navigation data. Thus, the environment 416 is similar to the environment within which the avatar is navigating. For example, the sidewalk 418 corresponds to the sidewalk seen in the image data 408. The model 412 can include all the features identified within the environment, such as traffic lights, road signs, etc. Additionally, the model 412 can include a mesh surface for the street on which the avatar is navigating.
[0126] In some embodiments, the analysis server may use an ensemble of self-drives driving in various environments and aggregate each model generated for each self-drive to generate an aggregated model. Now referring to Figure 5 , model 500 represents the aggregated model. Specifically, model 500 includes models generated using data retrieved from self-drives 502 to 510.
[0127] Returning to Figure 2 , at step 230, the analysis server may identify machine learning labels associated with at least one feature within the image data.
[0128] The analysis server may use the camera feed and / or 3D model to identify features within the area. As discussed herein, the analysis server may use various AI modeling and / or image recognition techniques to identify features included within the image data, such as traffic signs, traffic lights, lane markings, etc. Additionally or alternatively, a human reviewer may label the image data. For example, the human reviewer may view the camera feed and manually assign labels for different features depicted within the captured image (e.g., the camera feed).
[0129] In some embodiments, the analysis server may use a hybrid approach, where the analysis server may first execute an additional neural network (using the image data) to predict possible (e.g., initial estimates) labels for various features. For example, an image recognition protocol (e.g., a neural network) may generate an inference regarding the location of lane lines within the camera feed. Subsequently, the initial inference may be confirmed / validated by a human labeler. Moreover, the human labeler may also add, remove, and / or correct any labels generated as the initial inference. The interaction of the human labeler may also be monitored and used to improve the model used to generate the initial inference.
[0130] At step 240, the analysis server may receive second navigation data and second image data from a second self-drive not included within the ensemble of self-drives, the second self-drive navigation environment, the second image data including at least one feature.
[0131] The analysis server can receive new image data from a new ego (e.g., in the same area as the model discussed herein) navigating within the area. The analysis server can retrieve camera images from one or more of the ego's cameras. The ego discussed in step 240 may not be included in the egos that send their navigation and / or camera data to the analysis server in step 210. In some embodiments, the ego discussed with respect to step 240 may be a part of the ego discussed with respect to step 210. For example, the data received from the ego can be used to determine how to automatically label the camera feed within the area. As a result, the analysis server can use the automatic labeling paradigm discussed with respect to method 200 to automatically label the camera feed received from the same ego at a later time. Thus, the description of the ego herein is not intended to be limiting.
[0132] At step 250, the analysis server can automatically generate machine learning labels for at least one feature depicted within the second image data. Using the model, the analysis server can determine the labels associated with one or more features (step 240) received within the camera feed of the second ego. In some embodiments, the analysis server can use the navigation data (e.g., position data) of the ego to estimate the position and / or trajectory of the ego. For example, the analysis server can use the ego's VIO or IMU data to identify its travel trajectory and determine where the ego is navigating to / from.
[0133] Using the estimated position / trajectory, the analysis server can determine which model to use. The analysis server can align the camera feed of the ego with the various camera feeds received in step 210. As a result, the camera feeds can be aligned such that the labels of the features generated in steps 220 to 230 can be transferred to the features depicted within the camera feed received in step 250.
[0134] After identifying the model, the analysis server can align and compare the image data of the feature (received from the ego) with the description of the feature as indicated within the model. For example, if a video clip captured by the ego's camera depicts a feature, the analysis server can use the VIO, IMU, and other navigation data, as well as the video clip itself, to identify another video clip (captured in step 210) that includes the same feature. The analysis server can then align the video clips and determine the other features depicted within the video clips.
[0135] In some embodiments, the camera feed of the new ego can be compared with other image data captured by the set of egos to identify matching camera feeds. After aligning the camera feeds, the analysis server can pass the labels to the camera feed of the new ego.
[0136] Now referring to Figure 6, self-generated image data 602 to 610 (collectively referred to as camera feed 600). Using method 200, the analysis server has generated a model corresponding to the area in which the self is navigating. Using the self's position / trajectory data (or in some embodiments using an image recognition protocol performed using the camera feed), the analysis server can identify the model used for the self, where the model identifies various features of the environment. For example, the analysis server identifies model 612 and determines that the self is navigating within the estimated area of 614.
[0137] Using the identified model 612, the analysis server can determine one or more features depicted within camera feed 600. For example, image data 602 corresponds to the self's front camera and includes features of the street in which the self is navigating (feature 616). Using model 612, the analysis server determines that feature 616 is a crosswalk 618 included within model 612. The analysis server can also identify existing camera lenses of the same sidewalk (retrieved from one or more selves that previously navigated in the same street). As a result, the analysis server automatically labels feature 616 as a crosswalk (e.g., passing the label from a previous self to camera feed 600).
[0138] At step 260, the analysis server can optionally use the predicted / computed machine learning labels to train an AI model. The machine learning labels identified using method 200 (step 250) can be added to the camera feed received from a new self (step 240) and sent for training purposes to the AI model. Then, the camera feed and its labels can be ingested by the AI model (such as the occupancy network or occupancy detection model discussed herein) for training or retraining purposes.
[0139] Using method 200, the analysis server can automatically label various camera feeds captured from different selves. For example, and now referring to Figure 7, the different camera feeds depicted correspond to the same environment. However, each camera feed corresponds to a different condition (whether a weather condition or some other condition). For example, camera feed 700 may correspond to a dark condition; camera feed 702 may correspond to a foggy condition; camera feed 704 may correspond to an occluded condition (where an object at least partially obscures one or more features of the area); and camera feed 706 corresponds to a condition of rainfall. Although the features depicted are the same, the visual attributes of each feature may vary depending on the weather condition. For example, an image of the same feature (captured in different conditions) may look slightly different. Using method 200, the processor can determine (using the generated model) an indication of the feature and then automatically label the feature as it appears in different camera feeds (e.g., in different conditions or from different angles). In some embodiments, the camera feeds may not include overlapping elements. In some embodiments, Figure 3 The data depicted therein may represent an example of automatic labeling by automatically labeling in a challenging condition by passing automatic labeling after registering clips (drivers) from different conditions to a common 3D model or environment in good condition.
[0140] Using the methods and systems discussed herein, an analysis server can match / align a camera feed of an environment in a particular weather condition (e.g., fog or rainfall) with a camera feed of the same environment in a sunny condition. Subsequently, key features can be extracted, and the labels generated for the camera feed in sunny conditions can be transferred to the camera feed in fog or rainfall conditions.
[0141] In some embodiments, the 3D model can be enriched with other location data to transform it into an HD map. Then, the HD map can be used to localize the ego using visual data retrieved from the ego. For example, using a camera feed received from the ego, an analysis server (or a local processor of the ego) can match the camera feed to a particular location within the 3D model. This matching can be done by aligning key features of the camera feed (key features indicating a particular structure) with another key feature previously captured by another ego. As a result, using the camera feed and / or other navigation data, the position or trajectory of the ego can be determined, thereby using the camera feed of the ego and the 3D model to localize the ego.
[0142] The various illustrative logical blocks, modules, circuits, and algorithmic steps described in connection with the embodiments disclosed herein can be implemented as electronic hardware, computer software, or combinations of both. To clearly illustrate this interchangeability of hardware and software, various illustrative components, blocks, modules, circuits, and steps have been described generally in terms of their functionality. Whether such functionality is implemented as hardware or software depends upon the particular application and design constraints imposed on the overall system. Skilled artisans may implement the described functionality in varying ways for each particular application, but such implementation decisions should not be interpreted as causing a departure from the scope of the present disclosure or the claims.
[0143] Embodiments implemented in computer software can be implemented in software, firmware, middleware, microcode, hardware description language, or any combination thereof. Code segments or machine-executable instructions can represent a procedure, a function, a subprogram, a program, a routine, a subroutine, a module, a software package, a class, or any combination of instructions, data structures, or program statements. A code segment can be coupled to another code segment or a hardware circuit by passing and / or receiving information, data, arguments, parameters, or memory contents. Information, arguments, parameters, data, etc. can be passed, forwarded, or transmitted via any suitable means including memory sharing, message passing, token passing, network transmission, etc.
[0144] The actual software code or specialized control hardware used to implement these systems and methods does not limit the claimed features or the present disclosure. Accordingly, the operation and behavior of the systems and methods are described without reference to the specific software code, it being understood that the software and control hardware can be designed to implement the systems and methods based on the description herein.
[0145] When implemented in software, functions can be stored as one or more instructions or code on a non-transitory computer-readable or processor-readable storage medium. The steps of the methods or algorithms disclosed herein can be implemented in processor-executable software modules that can reside on a computer-readable or processor-readable storage medium. Non-transitory computer-readable or processor-readable media include both computer storage media and tangible storage media, and tangible storage media facilitate the transfer of a computer program from one place to another. The non-transitory processor-readable storage medium can be any available medium accessible by a computer. By way of example and not limitation, such non-transitory processor-readable media can include RAM, ROM, EEPROM, CD-ROM or other optical disk storage, magnetic disk storage or other magnetic storage devices, or any other tangible storage medium that can be used to store the desired program code in the form of instructions or data structures and that can be accessed by a computer or a processor. As used herein, disk and optical disk include compact disc (CD), laser disc, optical disc, digital versatile disc (DVD), Blu-ray disc, and floppy disk, where "disk" generally magnetically reproduces data, while "optical disc" optically reproduces data with a laser. Combinations of the above should also be included within the scope of computer-readable media. Additionally, operations of a method or algorithm can reside as code and / or instructions in one or any combination or collection on a non-transitory processor-readable medium and / or a computer-readable medium, which can be incorporated into a computer program product.
[0146] The foregoing description of the disclosed embodiments is provided to enable a person skilled in the art to make or use the embodiments described herein and their variations. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the principles defined herein can be applied to other embodiments without departing from the spirit or scope of the subject matter disclosed herein. Thus, the disclosure is not intended to be limited to the embodiments shown herein but is to be accorded the widest scope consistent with the following claims and the principles and novel features disclosed herein.
[0147] Although various aspects and embodiments have been disclosed, other aspects and embodiments are also contemplated. The various aspects and embodiments disclosed are for illustrative purposes and not intended to be limiting, and the true scope and spirit are indicated by the following claims.
Claims
1. A method, comprising: retrieving, by a processor, a set of navigation data and image data from an ego set navigating within an environment including at least one feature; generating, by the processor, a three-dimensional (3D) model of the environment using the navigation data and image data of at least a subset of the ego set, the 3D model including a virtual representation of the at least one feature of the environment; identifying, by the processor, machine learning labels associated with the at least one feature within the image data; receiving, by the processor, second navigation data and second image data from a second ego not included in the ego set, the second ego navigating within the environment, the second image data including the at least one feature; and automatically generating, by the processor, machine learning labels for the at least one feature depicted within the second image data.
2. The method according to claim 1, further comprising: filtering, by the processor, the set of navigation data and image data into a subset of the set of navigation data and image data according to the trajectory of each ego within the ego set.
3. The method according to claim 1, further comprising: locating, by the processor, the second ego using the second navigation data or the second image data according to the 3D model.
4. The method according to claim 1, wherein the at least one feature is at least one of a drivable surface, a traffic sign, a traffic light, a lane line, or a 3D structure.
5. The method according to claim 1, further comprising: sending, by the processor, the machine learning labels and the second image data to an artificial intelligence model.
6. The method according to claim 5, wherein the artificial intelligence model is an occupancy detection model.
7. The method according to claim 1, further comprising: executing, by the processor, an artificial intelligence model to predict the machine learning labels associated with the at least one feature within the image data.
8. The method according to claim 7, further comprising: receiving, by the processor, verification from a human reviewer regarding the machine learning labels predicted by the artificial intelligence model and associated with the at least one feature within the image data.
9. A computer-readable medium including an instruction set that, when executed, causes a processor to: retrieve a set of navigation data and image data from an ego set navigating within an environment including at least one feature; generate a three-dimensional (3D) model of the environment using the navigation data and image data of at least a subset of the ego set, the 3D model including a virtual representation of the at least one feature of the environment; identify machine learning labels associated with the at least one feature within the image data; receive second navigation data and second image data from a second ego not included in the ego set, the second ego navigating within the environment, the second image data including the at least one feature; and automatically generate machine learning labels for the at least one feature depicted within the second image data.
10. The computer-readable medium according to claim 9, wherein the instruction set further causes the processor to: Filter the set of navigation data and image data into a subset of the set of navigation data and image data based on the trajectories of each ego in the ego set.
11. The computer-readable medium according to claim 9, wherein the instruction set further causes the processor to: Locate the second ego using the second navigation data or the second image data based on the 3D model.
12. The computer-readable medium according to claim 9, wherein the at least one feature is at least one of a drivable surface, a traffic sign, a traffic light, a lane line, or a 3D structure.
13. The computer-readable medium according to claim 9, wherein the instruction set further causes the processor to: Send the machine learning label and the second image data to an artificial intelligence model.
14. The computer-readable medium according to claim 13, wherein the artificial intelligence model is an occupancy detection model.
15. The computer-readable medium according to claim 9, wherein the instruction set further causes the processor to: Execute the artificial intelligence model to predict the machine learning label associated with the at least one feature within the image data.
16. The computer-readable medium according to claim 15, wherein the instruction set further causes the processor to: Receive verification from a human reviewer regarding the machine learning label predicted by the artificial intelligence model and associated with the at least one feature within the image data.
17. A system, comprising: An ego set; and A processor in communication with the ego set, the processor being configured to: Retrieve a set of navigation data and image data from an ego set navigating within an environment including at least one feature; Generate a three-dimensional (3D) model of the environment using the navigation data and image data of at least a subset of the ego set, the 3D model including a virtual representation of the at least one feature of the environment; Identify a machine learning label associated with the at least one feature within the image data; Receive second navigation data and second image data from a second ego not included in the ego set, the second ego navigating within the environment, the second image data including the at least one feature; and Automatically generate a machine learning label for the at least one feature depicted in the second image data.
18. The system according to claim 17, wherein the processor is further configured to: Filter the set of navigation data and image data into a subset of the set of navigation data and image data based on the trajectories of each ego in the ego set.
19. The system according to claim 17, wherein the processor is further configured to: Locate the second ego using the second navigation data or the second image data based on the 3D model.
20. The system according to claim 17, wherein the at least one feature is at least one of a drivable surface, a traffic sign, a traffic light, a lane line, or a 3D structure.