Automatic operation domain tag generation
By identifying and generating training data related to the operational domain, the problem of incomplete training with real-world data is solved, the detection capability of machine learning models is improved, and the efficiency and accuracy of training data are enhanced.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-03
- Publication Date
- 2026-03-10
AI Technical Summary
Existing technologies struggle to effectively train machine learning models using real-world data, particularly due to a lack of semantic information and scene understanding, resulting in incomplete training data and difficulty in accurately identifying specific features and scenes.
By identifying and generating training data related to the operating domain, and comparing sensor data with annotated datasets, frame data related to specific operating domains can be identified and selected as training inputs, including features and scenes of interest, thereby improving the efficiency and accuracy of training data.
It improves the efficiency of converting real-world data into training data, ensures that the training data is more in line with the needs of real-world scenarios, and enhances the detection capabilities of machine learning models.
Smart Images

Figure CN121640116A_ABST
Abstract
Description
BACKGROUND
[0001] A machine learning (ML) model can be trained using a training dataset or training data, where the training dataset can include input features and corresponding target labels. The ML model can be trained to make predictions or classifications by finding patterns and / or relationships between the input features and the target labels. The training dataset can include input features and corresponding target labels. The ML model can learn from the training dataset the patterns and / or relationships that the ML model searches for in the input features.
[0002] For example, the input features can include certain objects, such as cars. The training dataset can include images (that include cars in various settings and / or orientations) and corresponding target labels (e.g., locating and identifying cars in the images). The ML model can learn, based at least on the training dataset, how to detect and / or identify cars in other images. Thus, the training data can characterize and / or shape the ML model. For example, the training data can affect the performance and / or generalization of the ML model. In certain instances, training data that reflects real-world values or scenarios on which the ML model can be implemented can help the ML model to be more accurate (e.g., accurately identify target labels based on provided input features).
[0003] Some methods for generating training data can include generating training data based at least on real-world data obtained using one or more sensors. For example, real-world data can be obtained using one or more cameras. However, such real-world data cannot be used directly as training data because the real-world data lacks semantic information about one or more features that can be present in the real-world data. Semantic information can refer to information about one or more features beyond literal representations. For example, semantic information can include textual data (e.g., descriptive words or sentences), image data (e.g., feature identification and relationships between features), and the like.
[0004] Furthermore, the real-world data can not include portions of a real-world scene corresponding to the real-world data, such that an understanding of the real-world scene can be incomplete. For example, the real-world data can only depict a certain portion of a feature, such that it can be difficult to fully explain characteristics of the feature, such as size, shape, orientation, and the like. In such instances, it can be difficult to distinguish a particular feature that can be used to train a particular ML model. SUMMARY
[0005] According to one or more embodiments of the present disclosure, a data record can be obtained. The data record can correspond to sensor data that includes frame data corresponding to one or more frames that depict a scene as represented by the frame data. The frame data of the one or more frames can be compared to an annotated data set that can include known features and annotations corresponding to the known features. In some embodiments, one or more features in the one or more frames can be identified based at least on the comparison between the frame data and the annotated data set. In these and other embodiments, a subset of the one or more frames can be determined that includes one or more features associated with one or more operational domains. Additionally, the subset of frames can be provided as training data to a detection model.
[0006] Embodiments of the present disclosure can improve training of certain ML models by identifying and / or generating suitable training data. For example, training data can be identified and / or generated using real-world data that is obtained using one or more sensors. One or more operational domains can be identified that correspond to the real-world data. In certain instances, an operational domain can refer to certain features of interest that can be useful to train a ML model for detection. In certain instances, an operational domain can include particular features based at least on a purpose or goal of a certain ML model. Additionally or alternatively, an operational domain can include a scene of interest, where the scene can include real-world and / or hypothetical situations and / or conditions that can exist in real-world data.
[0007] Embodiments of the present disclosure can more efficiently identify and / or generate training data for ML models than some other conventional methods. For example, some conventional methods can only include identifying patterns and features in real-world data. Such methods can not accurately understand context, relationships, and nuances in the real-world data and patterns associated therewith. BRIEF DESCRIPTION OF DRAWINGS
[0008] The present systems and methods for generating training data for ML models are described in detail below with reference to the attached drawing figures, wherein:
[0009] Figure 1 An example system configured to generate training data for ML models according to one or more embodiments of the present disclosure is shown;
[0010] Figures 2A-2B An example road map corresponding to an annotated data set according to one or more embodiments of the present disclosure is shown;
[0011] Figures 2C-2D An example view of a plurality of image frames according to one or more embodiments of the present disclosure is shown;
[0012] Figure 3 a flowchart showing a method for generating training data for an ML model is shown in accordance with one or more embodiments of the present disclosure;
[0013] Figure 4A is an illustration of an example autonomous vehicle in accordance with one or more embodiments of the present disclosure;
[0014] Figure 4B is an example of a camera position and field of view of an example autonomous vehicle in accordance with one or more embodiments of the present disclosure; Figure 4A
[0015] Figure 4C is an example of a camera position and field of view of an example autonomous vehicle in accordance with one or more embodiments of the present disclosure; Figure 4A is a block diagram of an example system architecture of an example autonomous vehicle in accordance with one or more embodiments of the present disclosure;
[0016] Figure 4D is a block diagram of an example computing device suitable for implementing one or more embodiments of the present disclosure; and Figure 4A is a system diagram of communication between a cloud-based server and an example autonomous vehicle in accordance with one or more embodiments of the present disclosure;
[0017] Figure 5 is a block diagram of an example computing device suitable for implementing one or more embodiments of the present disclosure; and
[0018] Figure 6 is a block diagram of an example data center suitable for implementing one or more embodiments of the present disclosure. DETAILED DESCRIPTION
[0019] One or more embodiments of the present disclosure can relate to improving training data to be used to train an ML model based on real-world data. In some embodiments, real-world data or sensor data can be obtained using one or more sensors (e.g., cameras, LiDAR sensors, RADAR sensors, ultrasonic sensors, etc.). For example, the sensor data can be associated with a recording obtained using the one or more sensors. In these and other embodiments, the sensor data can include information describing the recording. In some embodiments, the recording can include one or more frames, and the one or more frames can represent individual images and / or snapshots captured in order during obtaining the recording. For example, the one or more frames can be associated with different points in time and / or locations. In these and other embodiments, the one or more frames can be associated with corresponding frame data, which is part of the sensor data. In these and other embodiments, the frame data can include information for the individual images and / or snapshots.
[0020] In some embodiments, the system can be configured to analyze and / or process real-world data to identify certain frames (and corresponding frame data) that can be used to train ML models with respect to certain operational domains. For example, one or more frames that depict or correspond to operational domains that correspond to the purpose or goal of the ML model can be identified and curated to provide as training data to the ML model.
[0021] In some embodiments, certain images and / or frames of real-world data can be selected or identified based at least on images and / or frames that correspond to one or more particular operational domains to train ML models with respect to the one or more particular operational domains. For example, one or more particular operational domains can be identified to select frames that are suitable to train ML models with respect to the identified particular operational domains.
[0022] In some embodiments, one or more features associated with a frame can be identified. The identified features can be used to identify that the frame corresponds to certain operational domains. In some cases, the one or more features can be identified based at least on semantic labels. The semantic labels can refer to annotations and / or descriptive tags associated with the operational domains. The semantic labels can help to understand and / or organize operational domains in real-world data.
[0023] In some embodiments, one or more of the operational domains can refer to certain features of interest that can be useful to train ML models to detect. In some cases, the operational domains can include particular features based at least on the purpose or goal of a certain ML model. For example, a certain ML model can be designated to detect features that can affect the operation of a machine. For example, the machine can include a vehicle, and one or more of the operational domains can include features that can affect the vehicle’s travel, such as roads, lanes, lane lines, turning lanes, or obstacles (e.g., other vehicles, pedestrians, buildings, etc.).
[0024] In these and other embodiments, the operational domains (e.g., certain features of interest) can be identified by comparing frame data of one or more frames to an annotated dataset. The annotated dataset can include known features, labels, and / or characteristics of different regions or environments. In these and other embodiments, the annotated dataset can be obtained from various sources, such as ground truth maps.
[0025] Additionally or alternatively, the operational domain can include a scenario of interest. The scenario can include real-world and / or hypothetical situations and / or conditions that can exist in real-world data. For example, the operational domain can include weather conditions, traffic, a roadway (e.g., a highway exit), a curved road, an uneven road, etc. For example, a certain image frame can include a feature such as a highway exit. The certain image frame can be determined to be associated with a particular scenario that includes a highway exit. In some embodiments, one or more image frames can include multiple scenarios. For example, a frame can include and / or depict a highway exit on an overcast day, where the highway exit and the overcast day are different features or conditions that together can be described as a particular operational domain.
[0026] In these and other embodiments, the operational domain (e.g., scenario) can be identified by performing additional processes and / or operations. For example, certain complex features, such as a curved road, can be identified with further operations. For example, a curved road can be identified by calculating a curvature of a road segment along a route.
[0027] Compared to some conventional approaches, one or more embodiments of the present disclosure can help improve efficiency of converting real-world data into training data. For example, some approaches of generating training data can include identifying one or more features in raw data (e.g., real-world data). For example, the raw data can be obtained from different sources, such as a database, a sensor, etc. One or more features in an image can be identified, annotated, and / or labeled. In certain cases, one or more features can be manually annotated. For example, a human operator can draw a bounding shape around a feature and label and / or annotate the bounding shape to describe the feature corresponding to the bounding shape.
[0028] In certain cases, one or more features can be annotated using different tools (e.g., software platforms) that can assist in drawing a bounding shape and associating a label with the bounding shape. However, such conventional approaches can be inefficient. For example, such approaches can have scalability issues. Manually annotating and / or labeling a large amount of raw data is time-consuming, labor-intensive, and / or costly. Additionally, such approaches can include limited diversity. For example, a manually annotated training dataset can lack diversity with respect to scenarios, viewpoints, lighting conditions, and / or variations that a machine can encounter in real-world applications. For example, the raw data can be obtained as one or more sensors travel along a route or around a certain location. The one or more sensors can maintain the same viewpoint throughout, thereby limiting the data that can be collected. For example, the collected data can be limited with respect to various scenarios and / or conditions.
[0029] One or more embodiments disclosed herein can relate to generating training data for training ML models associated with a self-machine and / or components of one or more self-machines, which can include any applicable machine or system capable of performing one or more autonomous or semi-autonomous operations. Example self-machines can include, but are not limited to, vehicles (land, sea, space, and / or air), robots, robotic platforms, etc. For example, self-machine computing applications can include one or more applications that can be performed by an autonomous vehicle or a semi-autonomous vehicle, such as, for example, with respect to Figures 4A-4D An example autonomous vehicle 400 (or, referred to herein as “vehicle 400” or “self-machine 400”) is described. In this disclosure, references to “autonomous vehicle” or “semi-autonomous vehicle” can include any vehicle that can be configured to perform one or more autonomous or semi-autonomous navigation or driving operations. As such, such vehicles can also include vehicles that require an operator or in which an operator can also make such operations.
[0030] The systems and methods described herein can be used by, but are not limited to, non-autonomous vehicles or machines, semi-autonomous vehicles or machines (e.g., in one or more adaptive driver assistance systems (ADAS)), autonomous vehicles or machines, manned and unmanned robots or robotic platforms, warehouse vehicles, off-road vehicles, vehicles coupled to one or more trailers, spacecraft, ships, shuttles, emergency response vehicles, motorcycles, electric or motorized bicycles, aircraft, engineering vehicles, underwater vehicles, drones, and / or other vehicle types. Further, the systems and methods described herein can be used for various purposes, such as, by way of example and not limitation, machine control, machine motion, machine driving, synthetic data generation, model training, perception, augmented reality, virtual reality, mixed reality, robotics, safety and supervision, simulation and digital twin, autonomous or semi-autonomous machine applications, deep learning, environment simulation, object or actor simulation and / or digital twin, generative AI, data center processing, conversational AI (such as by leveraging one or more language models (such as one or more large language models (LLMs))), light transport simulation (e.g., ray tracing, path tracing, etc.), collaborative content creation for 3D assets, cloud computing, and / or any other suitable application.
[0031] The disclosed embodiments can be included in various different systems, such as automotive systems (e.g., control systems for autonomous or semi-autonomous machines, perception systems for autonomous or semi-autonomous machines), systems implemented using robots, aerial systems, medical systems, boating systems, smart area monitoring systems, systems for performing deep learning operations, systems for performing simulation operations, systems for performing digital twin operations, systems implemented using edge devices, systems including one or more virtual machines (VMs), systems for performing synthetic data generation operations, systems implemented at least partially in data centers, systems for performing conversational AI operations (e.g., systems implementing one or more language models (such as large language models (LLMs))), systems for performing one or more generative AI operations, systems for hosting real-time streaming applications, systems for presenting one or more of virtual reality content, augmented reality content, or mixed reality content, systems for performing optical transport simulation, systems for performing collaborative content creation of 3D assets, systems implemented at least partially using cloud computing resources, and / or other types of systems.
[0032] Embodiments of the present disclosure will be explained with reference to the drawings. It should be understood that these drawings are diagrams and illustrations of examples of the embodiments, and are not restrictive, and are not necessarily drawn to scale. In the drawings, like numbers represent similar structures and functions unless otherwise indicated.
[0033] With respect to Figure 1 , Figure 1 An example system 100 configured to generate training data for ML models is shown in accordance with one or more embodiments of the present disclosure. In some embodiments, system 100 can be implemented with respect to a machine. For example, system 100 can be implemented with respect to a vehicle 400. For example, system 100 can be configured to determine navigation operations for vehicle 400. Figures 4A-4D
[0034] In some embodiments, system 100 can include image processing module 104. In some embodiments, image processing module 104 can include code and routines configured to allow a computing system to perform one or more operations. Additionally or alternatively, image processing module 104 can be implemented using hardware including one or more processors, central processing units (CPUs), graphics processing units (GPUs), data processing units (DPUs), parallel processing units (PPUs), microprocessors (e.g., to perform or control performance of one or more operations), field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), accelerators (e.g., deep learning accelerators (DLAs)), programmable visual accelerators (including one or more direct memory access (DMA) systems and / or vector processing units (VPUs)), and / or other processor types. In these and other embodiments, image processing module 104 can be implemented using a combination of hardware and software. In this disclosure, operations described as being performed by image processing module 104 can include operations that the respective module can instruct a corresponding computing system to perform. In these or other embodiments, image processing module 104 can be implemented by one or more computing devices, such as the computing devices described in further detail with respect to Figures 4A-4D 、 Figure 5 and / or Figure 6 .
[0035] In some embodiments, image processing module 104 can be configured to obtain sensor data 102 and annotated dataset 106. In these and other embodiments, image processing module 104 can be configured to identify and / or select one or more image frames 108 from sensor data 102 based at least on annotated dataset 106.
[0036] In some embodiments, sensor data 102 can be obtained using one or more sensors (e.g., cameras). For example, the one or more sensors can be configured to obtain a recording of a real-world environment. In some cases, the one or more sensors can be associated with a machine, such as a vehicle. The recording can be associated with a real-world environment in which the vehicle travels. In these and other embodiments, sensor data 102 can correspond to information and / or data describing the recording. For example, sensor data 102 can include a data record associated with the recording.
[0037] In some embodiments, a data record can include one or more frames. For example, the one or more frames can represent individual snapshots and / or images included in the data record. The number of frames can vary based at least on the type and / or settings of the sensor. For example, one or more sensors can be configured to record a certain number of frames per second (FPS). The FPS and the length of the recording can define, accordingly, the total number of frames included in a certain data set of sensor data 102 corresponding to a certain data record.
[0038] In some embodiments, the sensor data 102 can be stored in a data store such that the sensor data 102 can be retrieved by the image processing module 104. In some embodiments, the one or more frames can be associated with corresponding frame data included in the sensor data 102. For example, the respective portion of the sensor data 102 corresponding to individual frames can be referred to as “frame data.” In these and other embodiments, the one or more frames can depict one or more features. For example, the one or more frames can depict objects (e.g., static and / or dynamic) and / or features (e.g., lane lines, images, text, patterns, textures, etc.) in a scene. In this disclosure, features depicted in a frame can also be referred to as being included in the frame.
[0039] In some embodiments, the image processing module 104 can be configured to determine image frames 108 that include input features that can be used for an ML model. In these and other embodiments, the input features can include and / or correspond to an operational domain for which the ML model can be trained. For example, in some embodiments, one or more operational domains can include certain features of interest that can be useful for training an ML model to detect. In certain instances, the operational domain can include particular features based at least on the purpose or goal of a certain ML model. For example, a certain ML model can be associated with a machine (e.g., a vehicle). The operational domain can include certain features that can affect the operation of the vehicle, such as roads, lanes, lane lines, turning lanes, or obstacles (e.g., other vehicles, pedestrians, buildings, etc.).
[0040] Additionally or alternatively, the operational domain can include a scene of interest. The scene can include real-world and / or hypothetical situations and / or conditions that can exist in real-world data. For example, the operational domain can include weather conditions, traffic, roadways (e.g., highway exits), curved roads, uneven roads, etc. For example, a certain image frame can include a feature such as a highway exit. The certain image frame can be determined to be associated with a particular scene that includes the highway exit. In some embodiments, one or more image frames can include multiple scenes. For example, a frame can include and / or depict a highway exit on an overcast day, where the highway exit and the overcast day are different features or conditions that can together be described as a particular operational domain.
[0041] In some embodiments, the image processing module 104 can be configured to identify one or more features in the sensor data 102 based at least on the annotated dataset 106. For example, the image processing module 104 can identify different objects and determine what the objects represent by comparing the image data 102 to the annotated dataset 106. The annotated dataset 106 can include a reference dataset that can be used as a reference for creating and annotating features in the image data 102. Further, the annotated dataset 106 can be used as ground truth data for identifying what the features represent. For example, the annotated dataset 106 can include a set of known features and corresponding annotations. In these and other embodiments, the image frames 108 generated by the image processing module 104 can include image data 102 with one or more features identified and annotated.
[0042] For example, the annotated dataset 106 can include a road map. For example, the road map can show a map or view of a setting or environment that can correspond to an area depicted in the sensor data 102. In some cases, the road map can include a plurality of directed road segments. The directed road segments can represent a direction of traffic flow in a corresponding road and / or lane. Further, the directed road segments can include ground truth labels for features (e.g., static, dynamic, etc.) within the directed road segments. In some embodiments, the labels can include a bounding shape that can be configured to encapsulate and / or surround an object and / or area.
[0043] For example, Figure 2A An example road map 200 is shown in accordance with one or more embodiments of the present disclosure. In some embodiments, the road map 200 can include one or more directed road segments 203 represented using corresponding arrows indicating a direction. The one or more directed road segments 203 can show a direction of traffic flow in the road map 200. In these and other embodiments, an individual directed road segment in the one or more directed road segments 203 can represent a directional flow corresponding to a particular segment and / or an area corresponding to the individual directed road segment. In some embodiments, the road map 200 can include known features and corresponding labels. For example, the road map 200 can include a crosswalk 205 and a label corresponding to the crosswalk 205. For example, a bounding shape can surround the crosswalk 205. Additionally or alternatively, the road map 200 can include more complex features, such as a curved road.
[0044] For example, Figure 2BA road map 210 is shown, illustrating a curved road 213 from different perspectives (e.g., map view, vehicle view). In some cases, a curved road may refer to a path that the machine can take to turn in different directions (e.g., left turn, right turn, U-turn, etc.). In some cases, the curved road 213 may not be associated with physical lane lines and / or directional road segments. For example, an intersection through which the machine turns may not include lane lines guiding the turn. In this case, the curved road 213 may be determined at least based on the curvature angle between roads. For example, the curved road 213 may illustrate how a vehicle can travel from a first set of directional road segments to a second set of directional road segments separated by intersections.
[0045] Back Figure 1 In some embodiments, the image processing module 204 may be based at least on a road map (e.g., Figure 2A Road map 200 and / or Figure 2B Image frame 108 is determined by the road map 210 and / or the labels included in the road map. In these and other embodiments, the road map may correspond to an annotated dataset 106. In some cases, the annotated dataset 106 or the corresponding road map may be obtained from a ground truth map. The ground truth map may include a set of reliable and known information that can be used to evaluate the correctness of the results of the ML model. The ground truth map or ground truth data may be obtained from multiple sources. For example, the ground truth map may be obtained from high-definition maps (HD maps), navigation maps, standard-definition maps (SD maps), etc., which correspond to different areas or scenes captured using one or more sensors of a machine. In these and other embodiments, the road map and corresponding labels can be generated by comparing sensor data 102 with the ground truth map.
[0046] In some embodiments, labels may correspond to static and / or dynamic features. Static features may include fixed and / or permanent elements or objects. Such static features may include features that do not typically change significantly in location and / or characteristics. For example, static features may include landmarks, buildings, structures, roads, lane markings on roads, traffic signs, etc. Dynamic features may include elements and / or objects whose location, characteristics, and / or state may change over time. For example, dynamic features may include vehicles, machines, pedestrians, weather information, etc.
[0047] In some embodiments, an annotated dataset 106, including one or more ground truth maps, can be used to identify and label static and dynamic features within sensor data 102. In some embodiments, labels generated at least based on the annotated dataset 106 can be filtered based on one or more rules to identify specific features for a specific purpose. For example, in the case of road maps and / or labels for vehicles, rules can be applied to identify specific labels that may be suitable for and / or useful to the vehicle. For example, rules may include finding associated lanes, finding the leftmost boundary, etc. Features identified in image data 102 by image processing module 104 at least based on the annotated data 106 can be represented using corresponding bounding shapes in image frame 108.
[0048] For example, Figure 2C Multiple image frames are shown (e.g.) Figure 1 View 220 of image frame 108. For example, view 220 shows a first frame 222a, a second frame 222b, a third frame 222c, a fourth frame 222d, a fifth frame 222e, and a sixth frame 222f (collectively, “frame 222”). Frame 222 may show a sequence of snapshots and / or images obtained by one or more sensors while the machine (e.g., a vehicle) is in motion. In some embodiments, frame 222 may depict the detection and / or identification of features such as a pedestrian crossing 225. In this case, a corresponding enclosing shape can be used to identify the pedestrian crossing 225. In these and other embodiments, frame 222 may show different portions and / or views of the pedestrian crossing 225 as the machine moves. For example, the third frame 222c shows the pedestrian crossing 225 in a closer view than the first frame 222a. In these and other embodiments, it may be based at least on a road map (e.g., Figure 2A Road maps 200) and / or annotated datasets (e.g., Figure 1 The annotated dataset 106 was used to identify pedestrian crossings 225.
[0049] In some embodiments, the identification and / or labeling of one or more features can be accomplished by performing additional processes and / or operations. For example, further operations can be used to identify certain complex features (such as a curved road). For instance, a curved road can be identified by calculating the curvature of a road segment along the route.
[0050] For example, Figure 2DView 230 with multiple frames is shown, such as first frame 232a, second frame 232b, third frame 232c, fourth frame 232d, fifth frame 232e, and sixth frame 232f (collectively, “Frame 232”). Frame 232 may show a sequence of snapshots and / or images obtained by one or more sensors while a machine (e.g., a vehicle) is in motion. In some cases, Frame 232 may identify and / or show a curved road 235. For example, a curved road 235 may be determined based at least on the curvature of the lines and / or road depicted in Frame 232.
[0051] Back Figure 1 In some embodiments, image frames 108 including one or more features identified by image processing module 104 may be obtained by selection module 110. In these and other embodiments, selection module 110 may be configured to identify which operational domains correspond to which frames in image frames 108. For example, selection module 110 may analyze annotated features in image frames 108 to determine which individual image frames 180 include features corresponding to operational domains.
[0052] Alternatively, selection module 110 may determine whether image frame 180 includes a domain of interest to be used for training the ML model. For example, the individual image frames identified as including a domain of interest may be further analyzed, at least based on the type of domain of interest included. For example, selection module 110 may determine selected frame 112 from image frame 108 that includes a domain of interest. For example, a particular ML model may be trained to detect features in rainy weather. In this case, selection module 110 may select individual frames that include a domain of interest associated with rainy weather.
[0053] In some embodiments, the selected frame 112 may be provided to the ML model as training data. For example, the selected frame 112 may correspond to one or more operation domains, which may be provided to the ML model as training input data to train the ML model with respect to certain operation domains. Furthermore, the selected frame 112 may illustrate the operation domains in a manner suitable for use as training data. In these and other embodiments, one or more individual frames of image frame 108 may be filtered out because they are not suitable as training data.
[0054] In some embodiments, the selection module 110 may be configured to further filter the selected frame 112 based at least on the location of the operation domain. For example, the location of the operation domain may be compared with the selected frame 112 to determine whether the selected frame 112 covers the location of the operation domain.
[0055] In some embodiments, the selection module 110 may be configured to project a three-dimensional (3D) operational domain detected in image frame 108 onto a two-dimensional (2D) coordinate system to determine the position of the operational domain. For example, the position of the operational domain relative to a machine and / or sensor may be determined. In some cases, one or more operational domains may be placed and / or located on a particular coordinate system. For example, the operational domain may correspond to a static feature, such as a traffic sign. The position of a traffic sign relative to a particular coordinate system can be determined by identifying a set of coordinates corresponding to the position of the traffic sign relative to the machine and / or sensor. For example, this set of coordinates can be identified by determining the distance and angle between the machine and the traffic sign. As another example, in the case where the operational domain corresponds to a scenario such as a highway exit, the position of the operational domain may be represented as a set of coordinates and / or a region within a particular coordinate system that includes the highway exit.
[0056] In some cases, a coordinate system may be associated with and / or related to one or more sensors that are being used to acquire sensor data 102. For example, in the case of acquiring collected map data using a camera, a particular coordinate system may include camera coordinates, where the coordinates are constructed or generated relative to the camera's position.
[0057] In some embodiments, selection module 110 may generate one or more positioning poses corresponding to the operating domain. Positioning poses may include estimates of the position and / or orientation (e.g., orientation) of the machine and / or one or more sensors. In some cases, a particular coordinate system may be divided into one or more tiles. In this case, positioning poses may be generated relative to the tiles. For example, the position of the operating domain may be determined at least based on the position of the tiles associated with the operating domain.
[0058] In these and other embodiments, the position of a specific operational domain relative to the camera coordinate system can be determined at least based on the positioning pose. For example, the position of a specific operational domain can be determined relative to different positioning poses. For example, different distances from different positioning poses can be combined (e.g., averaged and / or fused) to determine the position of a specific operational domain. Furthermore or alternatively, certain positioning poses associated with the operational domain can be assigned greater weight (e.g., have a greater impact on the determination) than other positioning poses, at least based on the distances between certain positioning poses and the operational domain. For example, different operational domains can be projected and / or displayed more accurately at certain distances. For example, signs (e.g., traffic signs) can be perceived more accurately, and / or the entire sign can be visible from 60 meters away. In this case, the positioning pose at approximately 60 meters away can be assigned greater weight than other positioning poses at different distances when determining the position of the sign.
[0059] In some embodiments, the validity of the selected frame 112 (e.g., suitability as training data) can be determined at least based on the location of the determined operational domain. In some embodiments, the location of the operational domain can be compared to the field of view (FOV) of one or more sensors. For example, a frustum of view for one or more sensors can be defined. The frustum of view can include geometry representing the view volume or FOV of one or more sensors. For example, for a camera, the frustum of view can be defined as what is visible and located within the camera's field of view. In some cases, the frustum of view can include multiple planes, such as a near plane (e.g., the plane closest to the camera) and a far plane (e.g., the plane farthest from the camera). The near and far planes can define the extent of the frustum (e.g., the extent of the frustum is from the near plane to the far plane).
[0060] In these and other embodiments, the selection module 110 can determine individual image frames from the selected frame 112 that include the operational domain located within the view frustum. For example, individual image frames in the selected frame 112 that do not correctly project the operational domain can be filtered out, so that these individual image frames are not used to train the ML model. In these and other embodiments, incorrect projection and / or inaccuracy may be caused by different types of errors, such as systematic errors and random errors.
[0061] In some cases, systematic errors can include consistent, repeatable errors caused by factors that can affect the entire process of feature detection performed by the ML model. For example, systematic errors can include intrinsic and / or extrinsic errors. Intrinsic errors can include errors caused by the intrinsic parameters of the sensor (e.g., camera), such as focal length, principal point, and / or lens distortion. When such intrinsic parameters are estimated during camera calibration, the system may suffer errors when projecting the operating domain into the frame. Extrinsic errors can include errors caused by the sensor's positioning and / or orientation relative to the machine associated with the sensor.
[0062] Random errors can include errors that may be caused by unpredictable factors. For example, random errors can include inconsistent or non-repeatable errors. For instance, random errors can be caused by dynamic changes in the environment, sensor noise, occlusion, lighting conditions, etc. Random errors can affect the quality of projection (e.g., depicting 3D objects in a 2D image frame).
[0063] In these and other embodiments, the selection module 110 can filter the selected frames 112 based at least on different types of errors. The unfiltered frames in the selected frames 112 can be provided as training data to the ML model. In some embodiments, the selected frames 112 can be converted and / or organized into different formats that may be suitable for different ML models. For example, the training data can be converted to CSV, JSON, a custom binary format, etc.
[0064] Can be Figure 1 Modifications, additions, or omissions may be made without departing from the scope of this disclosure. For example, system 100 may include more or fewer elements than those shown and described in this disclosure.
[0065] Figure 3 This is a flowchart illustrating a method 300 for generating training data for an ML model according to one or more embodiments of the present disclosure. In some embodiments, it may be related to... Figure 1 System 100 performs one or more operations of method 300. One or more operations of method 300 can be performed by any suitable system, apparatus, or device, such as system 100 of this disclosure. Figures 4A-4D Describing one or more autonomous vehicle systems, about Figure 5 Describe one or more computing devices and / or about Figure 6 Describe one or more data systems.
[0066] Method 300 may include one or more boxes. Although shown as discrete boxes, the operations associated with one or more boxes of method 300 may be divided into additional boxes, combined into fewer boxes, or eliminated, depending on the specific implementation.
[0067] Method 300 may include block 302. At block 302, a data record comprising one or more frames may be obtained. Additionally or alternatively, the one or more frames may include corresponding frame data. In some embodiments, one or more sensors may be used to obtain the data record. For example, one or more cameras associated with a machine may be used to obtain the data record. Additionally or alternatively, the data record may be obtained from a data storage that stores previously obtained data records. In some embodiments, this disclosure relates to... Figure 1 The sensor data 102 described may include or may be an example of the data record obtained at box 302.
[0068] At box 304, frame data of one or more frames can be compared to an annotated dataset, which includes known features and annotations corresponding to the known features. In some embodiments, the annotations may include labels corresponding to the known features. For example, the annotations may label what the known features represent. In some cases, the annotations may include enclosing shapes corresponding to the known features. For example, the enclosing shape may enclose the known features to represent the general boundaries of the known features. Additionally or alternatively, the annotations may include polylines corresponding to the known features. A polyline may include a set of points and / or vertices connected by line segments that typically delineate or trace linear features.
[0069] At box 306, one or more features in one or more frames can be identified, at least based on a comparison between frame data and an annotated dataset. For example, frame data can be analyzed to determine whether the corresponding frame includes any known features. The annotated dataset can be used as reference data to determine the presence of known features in one or more frames. In some embodiments, the identified features in the frames can be annotated to indicate the identity of the features. For example, the identified features can be labeled and / or a bounding shape corresponding to the identified features can be generated.
[0070] At box 308, a subset of one or more frames comprising one or more features associated with one or more operational domains can be determined. In some embodiments, an operational domain may include certain features of interest that may be useful for training an ML model. In some cases, an operational domain may include specific features based at least on the purpose or objective of an ML model. For example, an ML model may be associated with a machine (e.g., a vehicle). An operational domain may include certain features that can affect the operation of a vehicle, such as roads, lanes, lane markings, turning lanes, or obstacles (e.g., other vehicles, pedestrians, buildings, etc.).
[0071] Alternatively, the operational domain may include a scenario of interest. A scenario may include real-world and / or hypothetical situations and / or conditions that may exist in real-world data. For example, the operational domain may include weather conditions, traffic, access routes (e.g., highway exits), winding roads, uneven roads, etc.
[0072] In some embodiments, an operation domain can be described using a corresponding operation domain definition. In some embodiments, operation domain definitions can describe multiple operation domains together. For example, a first operation domain (e.g., vehicle) can be obtained together with a second operation domain (e.g., rainy day). In this case, the operation domain definition can define the operation domain as a vehicle in rainy weather.
[0073] At box 310, a subset of frames can be provided as training data to the detection model. For example, a subset of frames can be used to familiarize the detection model with the operational domains present in the subset. For example, the subset of frames can provide the detection model with instances that include different features and / or conditions (e.g., rainy days). In this case, the detection model can be trained to recognize vehicles with different features (e.g., vehicles under different weather conditions (e.g., rainy days)).
[0074] Method 300 may be modified, added to, or omitted without departing from the scope of this disclosure. For example, the operations of method 300 may be performed in a different order. Furthermore, or alternatively, two or more operations may be performed simultaneously. In addition, the operations and actions outlined are provided as examples only, and some operations and actions may be optional, combined into fewer operations and actions, or extended into additional operations and actions without departing from the essence of the described embodiments.
[0075] For example, after determining the frame subset at box 308, individual frames from the frame subset that include one or more errors can be filtered or removed. For example, individual frames from the frame subset that include one or more intrinsic errors, extrinsic errors, or random errors can be filtered.
[0076] In some embodiments, one or more frames that include operational domains that cannot be correctly projected within the frame boundaries can be removed from a subset of frames. For example, a 3D operational domain may not be correctly represented in a 2D frame due to one or more errors. In these and other embodiments, frame data corresponding to the frames can be analyzed to determine whether the 3D operational domain is correctly represented. In some cases, the operational domain can be compared with ground truth data to determine the accuracy of the operational domain in the frame. Additionally or alternatively, error analysis can be performed, which involves identifying common patterns or conditions present in incorrect projections. For example, frame data that includes patterns or conditions commonly present in frames containing errors can be removed from a subset of frames. In some cases, physical and manual verification can be performed. For example, the actual distance between one or more sensors and the operational domain can be measured and compared with the projected position of the operational domain in the frame to determine the accuracy of the frame.
[0077] In some cases, the error verification process can differ at least based on the type of feature. For example, when performing an error verification process for dynamic features, the projected positions of the dynamic features on multiple consecutive frames can be analyzed to determine whether the projected positions are consistent with the movement (e.g., the direction of movement) of the dynamic features.
[0078] Example autonomous vehicles
[0079] Figure 4AThis is an illustration of an example autonomous vehicle 400 according to some embodiments of the present disclosure. The autonomous vehicle 400 (or, alternatively, referred to herein as “vehicle 400”) may include, but is not limited to, passenger vehicles such as automobiles, trucks, buses, ambulances, shuttles, electric or motorized bicycles, motorcycles, fire trucks, police cars, ambulances, boats, engineering vehicles, underwater vessels, drones, and / or other types of vehicles (e.g., driverless and / or capable of accommodating one or more passengers). Autonomous vehicles are generally described according to the level of automation defined by the National Highway Traffic Safety Administration (NHTSA), a division of the U.S. Department of Transportation, and the Society of Automotive Engineers (SAE) “Taxonomy and Definitions for Terms Related to Driving Automation Systems for On-Road Motor Vehicles” (Standard No. J3016-201806, published June 15, 2018; Standard No. J3016-201609, published September 30, 2016; and previous and future versions of this standard). Vehicle 400 is capable of performing one or more functions that meet Level 3-5 of autonomous driving standards. Vehicle 400 is also capable of performing one or more functions that meet Level 1-5 of automated driving standards. For example, depending on the embodiment, vehicle 400 is capable of driver assistance (Level 1), partial automation (Level 2), conditional automation (Level 3), high automation (Level 4), and / or full automation (Level 5). The term "autonomous" as used herein can include any and / or all types of autonomy for 400 or other machines, such as full autonomy, high autonomy, conditional autonomy, partial autonomy, assisted autonomy, semi-autonomy, primary autonomy, or other names.
[0080] Vehicle 400 may include components such as chassis, body, wheels (e.g., 2, 4, 6, 8, 18, etc.), tires, axles, and other vehicle components. Vehicle 400 may include a propulsion system 450, such as an internal combustion engine, a hybrid power plant, an all-electric motor, and / or another type of propulsion system. Propulsion system 450 may be connected to the drivetrain of vehicle 400, which may include a transmission, to enable propulsion of vehicle 400. Propulsion system 450 may be controlled in response to receiving a signal from throttle / accelerator 452.
[0081] A steering system 454, which may include a steering wheel, can be used to steer the vehicle 400 (e.g., along a desired path or route) when the propulsion system 450 is operating (e.g., when the vehicle is in motion). The steering system 454 may receive signals from the steering actuator 456. For fully automatic (level 5) functions, the steering wheel may be optional.
[0082] The brake sensor system 446 can be used to operate the vehicle brakes in response to receiving signals from the brake actuator 448 and / or the brake sensor.
[0083] It may include one or more CPUs, System-on-a-Chip (SoC) 404 ( Figure 4C One or more controllers 436, including one or more GPUs, may provide signals (e.g., signals representing commands) to one or more components and / or systems of vehicle 400. For example, one or more controllers may send signals to operate vehicle brakes via one or more brake actuators 448, to operate steering system 454 via one or more steering actuators 456, and / or to operate propulsion system 450 via one or more throttles / accelerators 452. One or more controllers 436 may include one or more onboard (e.g., integrated) computing devices (e.g., supercomputers) that process sensor signals and output operating commands (e.g., signals representing commands) to enable autonomous driving and / or assist a human driver in driving vehicle 400. One or more controllers 436 may include a first controller 436 for autonomous driving functions, a second controller 436 for functional safety functions, a third controller 436 for artificial intelligence functions (e.g., computer vision), a fourth controller 436 for infotainment functions, a fifth controller 436 for redundancy in emergency situations, and / or other controllers. In some examples, a single controller 436 can handle two or more of the functions described above, and two or more controllers 436 can handle a single function, and / or any combination thereof.
[0084] One or more controllers 436 may provide signals for controlling one or more components and / or systems of vehicle 400 in response to sensor data (e.g., sensor inputs) received from one or more sensors. Sensor data may be received from, for example, but not limited to, global navigation satellite system sensors 458 (e.g., GPS sensors), RADAR sensors 460, ultrasonic sensors 462, LIDAR sensors 464, inertial measurement unit (IMU) sensors 466 (e.g., accelerometers, gyroscopes, magnetic compasses, magnetometers, etc.), microphones 796, stereo cameras 468, wide-angle cameras 470 (e.g., fisheye cameras), infrared cameras 472, surround cameras 474 (e.g., 360-degree cameras), long-range and / or medium-range cameras 498, speed sensors 444 (e.g., for measuring the rate of vehicle 400), vibration sensors 442, steering sensors 440, braking sensors (e.g., as part of braking sensor system 446), and / or other sensor types.
[0085] One or more of the controllers 436 may receive inputs (e.g., represented by input data) from the instrument cluster 432 of the vehicle 400 and provide outputs (e.g., represented by output data, display data, etc.) via a human-machine interface (HMI) display 434, an auditory signaling device, a speaker, and / or via other components of the vehicle 400. These outputs may include information such as vehicle speed, rate, time, map data (e.g., [missing information]). Figure 4C Information such as the HD map 422, location data (e.g., the location of vehicle 400 on the map), direction, and the location of other vehicles (e.g., occupying a grid), as well as information about objects and their states perceived by controller 436, etc. For example, HMI display 434 may display information about the existence of one or more objects (e.g., street signs, warning signs, traffic light changes, etc.) and / or information about driving maneuvers that the vehicle has made, is making, or will make (e.g., changing lanes now, leaving 34B in two miles, etc.).
[0086] Vehicle 400 further includes a network interface 424, which can communicate via one or more networks using one or more wireless antennas 426 and / or a modem. For example, network interface 424 may be able to communicate via LTE, WCDMA, UMTS, GSM, CDMA2000, etc. One or more wireless antennas 426 may also enable communication between objects in the environment (e.g., vehicles, mobile devices, etc.) using one or more local area networks such as Bluetooth, Bluetooth LE, Z-Wave, ZigBee, etc. and / or one or more low-power wide area networks (LPWAN) such as LoRaWAN, SigFox, etc.
[0087] Figure 4B For use in accordance with some embodiments of this disclosure Figure 4A This is an example of the camera position and field of view of an autonomous vehicle 400. The camera and its respective field of view are an example embodiment and are not intended to be limiting. For example, additional and / or replaceable cameras may be included, and / or these cameras may be located at different positions on the vehicle 400.
[0088] The camera type used for the camera may include, but is not limited to, a digital camera suitable for use with components and / or systems of vehicle 400. The camera may operate at Automotive Safety Integrity Level (ASIL) B and / or another ASIL. The camera type may have any image capture rate, such as 60 frames per second (fps), 120 fps, 240 fps, etc., depending on the embodiment. The camera may be able to use a rolling shutter, a global shutter, another type of shutter, or a combination thereof. In some examples, the color filter array may include a red-white-white-white (RCCC) color filter array, a red-white-white-blue (RCCB) color filter array, a red-blue-green-white (RBGC) color filter array, a Foveon X3 color filter array, a Bayer sensor (RGGB) color filter array, a monochrome sensor color filter array, and / or another type of color filter array. In some embodiments, a sharp-pixel camera, such as a camera with RCCC, RCCB, and / or RBGC color filter arrays, may be used in efforts to improve light sensitivity.
[0089] In some examples, one or more of the cameras can be used to perform advanced driver assistance system (ADAS) functions (e.g., as part of a redundant or fail-safe design). For example, a multi-function monocular camera can be installed to provide functions including lane departure warning, traffic sign assistance, and intelligent headlight control. One or more of the cameras (e.g., all cameras) can simultaneously record and provide image data (e.g., video).
[0090] One or more of the cameras can be mounted in mounting components such as custom-designed (3-D printed) parts to cut off stray light and reflections from inside the vehicle (e.g., reflections from the dashboard in the windshield mirror) that may interfere with the camera's image data capture capabilities. Regarding the wing mirror mounting components, the wing mirror components can be custom-3-D printed so that the camera mounting plate matches the shape of the wing mirror. In some examples, one or more cameras can be integrated into the wing mirror. For side-view cameras, one or more cameras can also be integrated into the four pillars at each corner of the cab.
[0091] A camera with a field of view that includes the environment in front of the vehicle 400 (e.g., a front-facing camera) can be used for surround view to help identify forward paths and obstacles, and, with the assistance of one or more controllers 436 and / or control SoCs, to provide information crucial for generating an occupancy grid and / or determining a preferred vehicle path. The front-facing camera can be used to perform many of the same ADAS functions as LiDAR, including emergency braking, pedestrian detection, and collision avoidance. The front-facing camera can also be used in ADAS functions and systems, including Lane Departure Warning (LDW), Autonomous Cruise Control (ACC), and / or other functions such as traffic sign recognition.
[0092] A variety of cameras can be used in front-facing configurations, including monocular camera platforms such as CMOS (Complementary Metal-Oxide-Semiconductor) color imagers. Another example could be a wide-angle camera 470, which can be used to perceive objects entering the field of view from the periphery (such as pedestrians, traffic at intersections, or bicycles). Although Figure 4B The middle image shows only one wide-angle camera, but any number of wide-angle cameras 470 can be present on the vehicle 400. Furthermore, remote cameras 498 (e.g., a pair of long-view stereo cameras) can be used for depth-based object detection, especially for objects for which neural networks have not yet been trained. Remote cameras 498 can also be used for object detection and classification, as well as basic object tracking.
[0093] One or more stereo cameras 468 may also be included in a front-mounted configuration. The stereo camera 468 may include an integrated control unit comprising a scalable processing unit that can provide a multi-core microprocessor and programmable logic (FPGA) with an integrated CAN or Ethernet interface on a single chip. Such a unit can be used to generate a 3D map of the vehicle environment, including distance estimates for all points in the image. Alternative stereo cameras 468 may include a compact stereo vision sensor that may include two camera lenses (one on each side) and an image processing chip capable of measuring the distance from the vehicle to a target object and using the generated information (e.g., metadata) to activate autonomous emergency braking and lane departure warning functions. Other types of stereo cameras 468 may be used in addition to those described herein, or alternatively.
[0094] Cameras with a field of view including the side portion of the vehicle 400 (e.g., side-view cameras) can be used for surround view, providing information for creating and updating occupancy grids and generating side-impact collision warnings. For example, surround camera 474 (e.g., ... Figure 4B The four surround cameras 474 shown can be mounted on the vehicle 400. The surround cameras 474 can include wide-angle cameras 470, fisheye cameras, 360-degree cameras, and / or the like. For example, four fisheye cameras can be positioned at the front, rear, and sides of the vehicle. In an alternative arrangement, the vehicle can use three surround cameras 474 (e.g., left, right, and rear) and can utilize one or more other cameras (e.g., forward-facing cameras) as a fourth surround-view camera.
[0095] A camera with a field of view that includes the environment behind the vehicle 400 (e.g., a rear-view camera) can be used for parking assistance, surround view, rear collision warning, and creating and updating occupancy grids. A wide variety of cameras can be used, including but not limited to those also suitable as front-facing cameras as described herein (e.g., long-range and / or mid-range camera 498, stereo camera 468, infrared camera 472, etc.).
[0096] Figure 4C For use in accordance with some embodiments of this disclosure Figure 4A The example autonomous vehicle 400 is illustrated in the block diagram of an example system architecture. It should be understood that this arrangement, and other arrangements described herein, are merely illustrative. Other arrangements and elements (e.g., machines, interfaces, functions, sequences, functional groupings, etc.) may be used in addition to or in place of those shown, and some elements may be omitted entirely. Furthermore, many of the elements described herein are functional entities, which may be implemented as discrete or distributed components or in combination with other components, and in any suitable combination and location. The various functions described herein as being performed by these entities can be implemented via hardware, firmware, and / or software. For example, the various functions can be implemented by a processor executing instructions stored in memory.
[0097] Figure 4C Each component, feature, and system in vehicle 400 is illustrated as being connected via bus 402. Bus 402 may include a Controller Area Network (CAN) data interface (or, alternatively, referred to herein as the "CAN bus"). CAN may be a network within vehicle 400 used to assist in the control of various features and functions of vehicle 400, such as the actuation of brakes, acceleration, braking, steering, windshield wipers, etc. The CAN bus may be configured to have dozens or even hundreds of nodes, each with its own unique identifier (e.g., CAN ID). The CAN bus can be read to find steering wheel angle, ground speed, engine speed per minute (RPM), button positions, and / or other vehicle status indicators. The CAN bus may be ASIL B compliant.
[0098] Although bus 402 is described herein as a CAN bus, this is not intended to be limiting. For example, FlexRay and / or Ethernet may be used in addition to or alternatively to a CAN bus. Furthermore, although bus 402 is represented by a single line, this is not intended to be limiting. For example, any number of buses 402 may exist, which may include one or more CAN buses, one or more FlexRay buses, one or more Ethernet buses, and / or one or more other types of buses using different protocols. In some examples, two or more buses 402 may be used to perform different functions and / or may be used for redundancy. For example, a first bus 402 may be used for a collision avoidance function, and a second bus 402 may be used for drive control. In any example, each bus 402 may communicate with any component of vehicle 400, and two or more buses 402 may communicate with the same component. In some examples, each SoC 404, each controller 436, and / or each computer within the vehicle may have access to the same input data (e.g., input from sensors of vehicle 400) and may be connected to a common bus such as a CAN bus.
[0099] Vehicle 400 may include one or more controllers 436, such as those described herein. Figure 4A The controllers described. Controller 436 can be used for a wide variety of functions. Controller 436 can be coupled to any other different components and systems of vehicle 400 and can be used for the control of vehicle 400, artificial intelligence of vehicle 400, infotainment and / or the like for vehicle 400.
[0100] Vehicle 400 may include one or more System-on-Chip (SoC) 404. SoC 404 may include a CPU 406, GPU 408, processor 410, cache 412, accelerator 414, data storage 416, and / or other components and features not shown. SoC 404 can be used to control vehicle 400 across a wide variety of platforms and systems. For example, one or more SoCs 404 may be combined with an HD map 422 in a system (e.g., the system of vehicle 400), the HD map being transmitted via a network interface 424 from one or more servers (e.g., [server name missing]). Figure 4D One or more servers (478) receive map refresh and / or updates.
[0101] CPU 406 may include a CPU cluster or CPU complex (or, alternatively, referred to herein as "CCPLEX"). CPU 406 may include multiple cores and / or L2 cache. For example, in some embodiments, CPU 406 may include eight cores in a coherent multiprocessor configuration. In some embodiments, CPU 406 may include four dual-core clusters, each cluster having a dedicated L2 cache (e.g., 2MB L2 cache). CPU 406 (e.g., CCPLEX) may be configured to support simultaneous cluster operation, such that any combination of clusters of CPU 406 can be active at any given time.
[0102] CPU 406 can implement power management capabilities including one or more of the following features: automatic clock gating of hardware blocks when idle to conserve dynamic power; clock gating of each core when the core is not actively executing instructions due to the execution of WFI / WFE instructions; independent power gating of each core; independent clock gating of each core cluster when all cores are clock-gated or power-gated; and / or independent power gating of each core cluster when all cores are power-gated. CPU 406 can further implement enhanced algorithms for managing power states, wherein allowed power states and desired wake-up times are specified, and the hardware / microcode determines the optimal power state to enter for the core, cluster, and CCPLEX. The processing core can support simplified power state entry sequences in software, with this work offloaded to the microcode.
[0103] GPU 408 may include an integrated GPU (or, alternatively, referred to herein as an "iGPU"). GPU 408 may be programmable and efficient for parallel workloads. In some examples, GPU 408 may use an enhanced tensor instruction set. GPU 408 may include one or more streaming microprocessors, wherein each streaming microprocessor may include an L1 cache (e.g., an L1 cache with at least 96KB of storage capacity), and two or more of these streaming microprocessors may share an L2 cache (e.g., an L2 cache with 512KB of storage capacity). In some embodiments, GPU 408 may include at least eight streaming microprocessors. GPU 408 may use a computation application programming interface (API). Furthermore, GPU 408 may use one or more parallel computing platforms and / or programming models (e.g., NVIDIA's CUDA).
[0104] In automotive and embedded applications, the GPU 408 can be power-optimized for optimal performance. For example, the GPU 408 can be fabricated on FinFETs. However, this is not intended to be limiting, and the GPU 408 can be fabricated using other semiconductor manufacturing processes. Each streaming microprocessor can combine several mixed-precision processing cores divided into multiple blocks. For example, and without limitation, 64 PF32 cores and 32 PF64 cores can be divided into four processing blocks. In such an example, each processing block can be allocated 16 FP32 cores, 8 FP64 cores, 16 INT32 cores, two mixed-precision NVIDIA Tensor cores for deep learning matrix arithmetic, an L0 instruction cache, a warp scheduler, dispatch units, and / or a 64KB register file. Furthermore, the streaming microprocessor can include independent parallel integer and floating-point data paths to provide efficient execution of workloads by leveraging the mixture of computation and addressing computation. The streaming microprocessor can include independent thread scheduling capabilities to allow for finer-grained synchronization and cooperation between parallel threads. Streaming microprocessors can include a combination of L1 data cache and shared memory units to improve performance while simplifying programming.
[0105] The GPU 408 may include, in some examples, a High Bandwidth Memory (HBM) and / or a 16GB HBM2 memory subsystem providing a peak memory bandwidth of approximately 900GB / s. In some examples, in addition to HBM memory or alternatively, Synchronous Graphics Random Access Memory (SGRAM), such as Generation 5 Graphics Double Data Rate Synchronous Random Access Memory (GDDR5), may be used.
[0106] The GPU 408 may include unified memory technology, which includes access counters to allow memory pages to be migrated more precisely to the processors that access them most frequently, thereby improving the efficiency of shared memory ranges between processors. In some examples, Address Translation Service (ATS) support can be used to allow the GPU 408 to directly access the CPU 406 page tables. In such examples, when the GPU 408 Memory Management Unit (MMU) experiences a miss, the address translation request can be transferred to the CPU 406. In response, the CPU 406 can look up the virtual-physical mapping for the address in its page tables and transfer the translation back to the GPU 408. In this way, unified memory technology can allow a single unified virtual address space for the memory of both the CPU 406 and the GPU 408, thereby simplifying GPU 408 programming and porting applications to the GPU 408.
[0107] In addition, the GPU 408 may include access counters that track how frequently the GPU 408 accesses the memory of other processors. These access counters help ensure that memory pages are moved to the physical memory of the processor that accesses those pages most frequently.
[0108] SoC 404 may include any number of caches 412, including those described herein. For example, cache 412 may include an L3 cache available to both CPU 406 and GPU 408 (e.g., it is connected to both CPU 406 and GPU 408). Cache 412 may include a write-back cache, which can track the state of rows, for example, using a cache coherence protocol (e.g., MEI, MESI, MSI, etc.). Depending on the embodiment, the L3 cache may include 4 MB or more, but a smaller cache size may also be used.
[0109] SoC 404 may include one or more arithmetic logic units (ALUs) that can be used to perform processing of any of a variety of tasks or operations relating to vehicle 400, such as processing a DNN. Furthermore, SoC 404 may include a floating-point unit (FPU) or other mathematical coprocessor or digital coprocessor type for performing mathematical operations within the system. For example, SoC 404 may include one or more FPUs integrated as execution units within CPU 406 and / or GPU 408.
[0110] SoC 404 may include one or more accelerators 414 (e.g., hardware accelerators, software accelerators, or combinations thereof). For example, SoC 404 may include a hardware acceleration cluster, which may include optimized hardware accelerators and / or large on-chip memory. This large on-chip memory (e.g., 4MB SRAM) can enable the hardware acceleration cluster to accelerate neural networks and other computations. The hardware acceleration cluster can be used to supplement GPU 408 and offload some tasks from GPU 408 (e.g., freeing up more cycles of GPU 408 to perform other tasks). As an example, accelerator 414 can be used for targeted workloads (e.g., perceptrons, convolutional neural networks (CNNs), etc.) that are stable enough to be easily controlled for acceleration. When used herein, the term "CNN" can include all types of CNNs, including region-based or region convolutional neural networks (RCNNs) and fast RCNNs (e.g., for object detection).
[0111] Accelerator 414 (e.g., a hardware acceleration cluster) may include a Deep Learning Accelerator (DLA). The DLA may include one or more Tensor Processing Units (TPUs) that can be configured to provide an additional 10 trillion operations per second for deep learning applications and inference. The TPU may be an accelerator configured to perform image processing functions (e.g., for CNNs, RCNNs, etc.) and optimized for performing image processing functions. The DLA may be further optimized for a specific set of neural network types and floating-point operations and inference. The DLA is designed to provide higher performance per millimeter than a general-purpose GPU and significantly outperform CPUs. The TPU can perform several functions, including single-instance convolution functions, support for INT8, INT16, and FP16 data types for both features and weights, and post-processor functions.
[0112] DLA can execute neural networks, especially CNNs, quickly and efficiently on processed or unprocessed data for any function across a wide variety of applications, such as, but not limited to: CNNs for object recognition and detection using data from camera sensors; CNNs for distance estimation using data from camera sensors; CNNs for emergency vehicle detection and recognition using data from microphones; CNNs for face recognition and vehicle owner recognition using data from camera sensors; and / or CNNs for safety and / or safety-related events.
[0113] The DLA can perform any function of the GPU 408, and by using inference accelerators, for example, a designer can make either the DLA or the GPU 408 target any function. For example, a designer can focus the CNN processing and floating-point operations on the DLA and leave other functions to the GPU 408 and / or other accelerators 414.
[0114] Accelerator 414 (e.g., a hardware acceleration cluster) may include a programmable vision accelerator (PVA), which may alternatively be referred to herein as a computer vision accelerator. The PVA may be designed and configured to accelerate computer vision algorithms for advanced driver assistance systems (ADAS), autonomous driving, and / or augmented reality (AR) and / or virtual reality (VR) applications. The PVA can provide a balance between performance and flexibility. For example, each PVA may include, for example, but not limited to, any number of reduced instruction set computer (RISC) cores, direct memory access (DMA), and / or any number of vector processors.
[0115] RISC cores can interact with image sensors (such as the image sensor of any camera described herein), image signal processors, and / or the like. Each of these RISC cores may include any amount of memory. Depending on the embodiment, the RISC core may use any of several protocols. In some examples, the RISC core may execute a real-time operating system (RTOS). RISC cores may be implemented using one or more integrated circuit devices, application-specific integrated circuits (ASICs), and / or memory devices. For example, a RISC core may include an instruction cache and / or tightly coupled RAM.
[0116] DMA enables PVA components to access system memory independently of the CPU 406. DMA can support any number of features used to provide optimizations to the PVA, including but not limited to support for multidimensional addressing and / or circular addressing. In some examples, DMA can support addressing in up to six or more dimensions, which can include block width, block height, block depth, horizontal block step, vertical block step, and / or depth step.
[0117] A vector processor can be a programmable processor designed to efficiently and flexibly execute programming for computer vision algorithms and provide signal processing capabilities. In some examples, a PVA may include a PVA core and two vector processing subsystem partitions. The PVA core may include a processor subsystem, one or more DMA engines (e.g., two DMA engines), and / or other peripherals. The vector processing subsystem may operate as the main processing engine of the PVA and may include a vector processing unit (VPU), an instruction cache, and / or vector memory (e.g., VMEM). The VPU core may include a digital signal processor, such as, for example, a Single Instruction Multiple Data (SIMD) or Very Long Instruction Word (VLIW) digital signal processor. The combination of SIMD and VLIW can enhance throughput and speed.
[0118] Each of the vector processors may include an instruction cache and may be coupled to dedicated memory. Consequently, in some examples, each of the vector processors may be configured to execute independently of other vector processors. In other examples, the vector processors included in a particular PVA may be configured to employ data parallelism. For example, in some embodiments, multiple vector processors included in a single PVA may execute the same computer vision algorithm, but on different regions of an image. In other examples, vector processors included in a particular PVA may execute different computer vision algorithms simultaneously on the same image, or even different algorithms on a sequence of images or portions of an image. Among other things, any number of PVAs may be included in a hardware-accelerated cluster, and any number of vector processors may be included in each of these PVAs. Furthermore, the PVA may include additional error-correcting code (ECC) memory to enhance overall system security.
[0119] Accelerator 414 (e.g., a hardware acceleration cluster) may include an on-chip computer vision network and SRAM to provide high-bandwidth, low-latency SRAM for accelerator 414. In some examples, on-chip memory may include at least 4MB of SRAM consisting of, for example, but not limited to, eight field-configurable memory blocks, accessible by both the PVA and DLA. Each pair of memory blocks may include an Advanced Peripheral Bus (APB) interface, configuration circuitry, a controller, and a multiplexer. Any type of memory may be used. The PVA and DLA may access memory via a backbone that provides high-speed memory access to the PVA and DLA. The backbone may include (e.g., using an APB) an on-chip computer vision network that interconnects the PVA and DLA to memory.
[0120] On-chip computer vision networks can include interfaces that ensure both the PVA and DLA provide ready and valid signals before transmitting any control signals / addresses / data. Such interfaces can provide separate phases and channels for transmitting control signals / addresses / data, as well as burst communication for continuous data transmission. This type of interface can conform to ISO 26262 or IEC 61508 standards, but other standards and protocols can also be used.
[0121] In some examples, SoC 404 may include, for example, a real-time ray tracing hardware accelerator as described in U.S. Patent Application No. 16 / 101,232, filed August 10, 2018. This real-time ray tracing hardware accelerator can be used to quickly and efficiently determine the location and extent of objects (e.g., within a world model) to generate real-time visualization simulations for RADAR signal interpretation, sound propagation synthesis and / or analysis, SONAR system simulation, general wave propagation simulation, comparison with LiDAR data for localization and / or other functional purposes, and / or other uses. In some embodiments, one or more Tree Traversal Units (TTUs) may be used to perform one or more ray tracing-related operations.
[0122] Accelerators 414 (e.g., hardware accelerator clusters) have broad applications in autonomous driving. PVAs can be programmable vision accelerators used in critical processing stages of ADAS and autonomous vehicles. The capabilities of PVAs are a good match for algorithmic domains requiring predictable processing, low power, and low latency. In other words, PVAs perform well in semi-dense or dense rule computation, even on small datasets requiring predictable runtimes with low latency and low power. Therefore, in the context of platforms for autonomous vehicles, PVAs are designed to run classical computer vision algorithms because they are efficient in object detection and integer mathematical operations.
[0123] For example, according to one embodiment of this technology, PVA is used to perform computer stereo vision. In some examples, semi-global matching-based algorithms may be used, but this is not intended to be limiting. Many applications for Level 3–5 autonomous driving require instantaneous motion estimation / stereo matching (e.g., from moving structures, pedestrian recognition, lane detection, etc.). PVA can perform computer stereo vision functions on input from two monocular cameras.
[0124] In some examples, PVA can be used to perform intensive optical flow, processing raw RADAR data (e.g., using 4D Fast Fourier Transform) to provide processed RADAR. In other examples, PVA is used for time-of-flight depth processing, which, for example, involves processing raw time-of-flight data to provide processed time-of-flight data.
[0125] DLA can be used to run any type of network to enhance control and driving safety, including, for example, neural networks that output a confidence metric for each object detection. Such a confidence value can be interpreted as a probability or as providing a relative “weight” for each detection compared to other detections. This confidence value allows the system to make further decisions about which detections should be considered true positives rather than false positives. For example, the system can set a threshold for the confidence and only consider detections exceeding the threshold as true positives. In an Automatic Emergency Braking (AEB) system, false positives can cause the vehicle to automatically perform emergency braking, which is clearly undesirable. Therefore, only the most confident detections should be considered as triggers for AEB. DLA can run a neural network to regress the confidence value. This neural network can take at least a subset of parameters as its input, such as bounding box dimensions, ground plane estimates (e.g., from another subsystem), inertial measurement unit (IMU) sensor 466 outputs related to vehicle orientation and distance, 3D position estimates of objects obtained from the neural network and / or other sensors (e.g., LiDAR sensor 464 or RADAR sensor 460), etc.
[0126] SoC 404 may include one or more data stores 416 (e.g., memory). Data stores 416 may be on-chip memory of SoC 404, which may store neural networks to be executed on the GPU and / or DLA. In some examples, for redundancy and security, data stores 416 may be large enough to store multiple instances of the neural network. Data stores 412 may include L2 or L3 caches 412. References to data stores 416 may include references to memory associated with PVA, DLA, and / or other accelerators 414 as described herein.
[0127] SoC 404 may include one or more processors 410 (e.g., embedded processors). Processor 410 may include a startup and power management processor, which may be a dedicated processor and subsystem for handling startup power and management functions, as well as safety implementation. The startup and power management processor may be part of the SoC 404 startup sequence and may provide runtime power management services. The startup power and management processor may provide clock and voltage programming, auxiliary system low-power state transitions, SoC 404 thermal and temperature sensor management, and / or SoC 404 power state management. Each temperature sensor may be implemented as a ring oscillator whose output frequency is proportional to the temperature, and SoC 404 may use the ring oscillator to detect the temperature of CPU 406, GPU 408, and / or accelerator 414. If it is determined that the temperature exceeds a threshold, the startup and power management processor may enter a temperature fault routine and place SoC 404 into a lower power state and / or place vehicle 400 into a driver-safe parking mode (e.g., safely stop vehicle 400).
[0128] Processor 410 may further include a set of embedded processors that can be used as an audio processing engine. The audio processing engine can be an audio subsystem that allows for full hardware support for multi-channel audio via multiple interfaces, as well as a wide and flexible range of audio I / O interfaces. In some examples, the audio processing engine is a dedicated processor core with a digital signal processor and dedicated RAM.
[0129] The processor 410 may further include an always-on-processor engine that can provide the necessary hardware features to support low-power sensor management and wake-up use cases. This always-on-processor engine may include a processor core, tightly coupled RAM, support for peripherals (such as timers and interrupt controllers), various I / O controller peripherals, and routing logic.
[0130] Processor 410 may further include a secure cluster engine, which includes a dedicated processor subsystem for handling security management of automotive applications. The secure cluster engine may include two or more processor cores, tightly coupled RAM, support for peripheral devices (e.g., timers, interrupt controllers, etc.), and / or routing logic. In secure mode, the two or more cores may operate in lockstep mode and function as a single core with comparison logic that detects any differences between their operations.
[0131] The processor 410 may further include a real-time camera engine, which may include a dedicated processor subsystem for handling real-time camera management.
[0132] The processor 410 may further include a high dynamic range signal processor, which may include an image signal processor, which is a hardware engine that is part of the camera processing pipeline.
[0133] Processor 410 may include a video image compositer, which may be (e.g., implemented on a microprocessor) a processing block, implementing video post-processing functions required by the video playback application to generate the final image for the player window. The video image compositer may perform lens distortion correction on the wide-angle camera 470, the surround camera 474, and / or the in-cabin monitoring camera sensor. The in-cabin monitoring camera sensor is preferably monitored by a neural network running on another instance of an advanced SoC, configured to recognize in-cabin events and respond accordingly. The in-cabin system may perform lip reading to activate mobile phone services and make calls, dictate emails, change vehicle destinations, activate or change the vehicle's infotainment system and settings, or provide voice-activated web browsing. Some functions are only available to the driver when the vehicle is operating in autonomous mode and are disabled in other situations.
[0134] Video image compositers can include enhanced temporal denoising for both spatial and temporal noise reduction. For example, in the case of motion in the video, denoising appropriately weights spatial information, reducing the weight of information provided by neighboring frames. In cases where the image or part of the image does not contain motion, the temporal denoising performed by the video image compositer can use information from previous images to reduce noise in the current image.
[0135] The video image compositer can also be configured to perform stereo correction on input stereo lens frames. When the operating system desktop is in use and the GPU 408 does not need to continuously render new surfaces, the video image compositer can be further used for user interface components. Even when the GPU 408 is powered on and active, performing 3D rendering, the video image compositer can be used to offload the GPU 408 to improve performance and responsiveness.
[0136] SoC 404 may further include a Mobile Industry Processor Interface (MIPI) camera serial interface, a high-speed interface, and / or a video input block that can be used for camera and related pixel input functions for receiving video and input from a camera. SoC 404 may further include an input / output controller that can be software-controlled and can be used to receive I / O signals not assigned to a specific role.
[0137] SoC 404 may further include a wide range of peripheral interfaces to enable communication with peripherals, audio codecs, power management and / or other devices. SoC 404 can be used to process data from cameras and sensors (e.g., LIDAR sensor 464, RADAR sensor 460, etc., which can be connected via Gigabit Multimedia Serial Link and Ethernet), data from bus 402 (e.g., vehicle 400 speed, steering wheel position, etc.), and data from GNSS sensor 458 (connected via Ethernet or CAN bus). SoC 404 may further include a dedicated high-performance, high-capacity memory controller, which may include its own DMA engine, and which can be used to free up CPU 406 from routine data management tasks.
[0138] SoC 404 can be an end-to-end platform with a flexible architecture spanning Levels 3-5 of automation, providing a comprehensive functional safety architecture that leverages and efficiently utilizes computer vision and ADAS technologies for diversity and redundancy, along with deep learning tools, to deliver a flexible and reliable driving software stack. SoC 404 can be faster, more reliable, and even more energy- and space-efficient than conventional systems. For example, when combined with CPU 406, GPU 408, and data storage 416, accelerator 414 can provide a fast and efficient platform for Level 3-5 autonomous vehicles.
[0139] Therefore, this technology offers capabilities and functionalities that cannot be achieved through conventional systems. For example, computer vision algorithms can be executed on CPUs, which can be configured using high-level programming languages such as C to execute a wide variety of processing algorithms across a diverse range of visual data. However, CPUs often cannot meet the performance requirements of many computer vision applications, such as those related to execution time and power consumption. In particular, many CPUs cannot execute complex object detection algorithms in real time, which is a requirement for automotive ADAS applications and practical Level 3-5 autonomous vehicles.
[0140] In contrast to conventional systems, the techniques described in this paper, by providing CPU complexes, GPU complexes, and hardware acceleration clusters, allow multiple neural networks to be executed simultaneously and / or sequentially, and the results combined to achieve Level 3–5 autonomous driving capabilities. For example, a CNN executed on a DLA or dGPU (e.g., GPU 420) could include text and word recognition, allowing a supercomputer to read and understand traffic signs, including those for which neural networks have not yet been specifically trained. The DLA could further include a neural network capable of recognizing, interpreting, and providing semantic understanding of the signs, and passing that semantic understanding to a path planning module running on the CPU complex.
[0141] As another example, multiple neural networks can operate simultaneously, as required for Level 3, 4, or 5 driving. For instance, a warning sign consisting of "Caution: Flashing lights indicate icy conditions" along with a light can be interpreted independently or jointly by several neural networks. The sign itself can be recognized as a traffic sign by a deployed first neural network (e.g., a trained neural network), and the text "Flashing lights indicate icy conditions" can be interpreted by a deployed second neural network that informs the vehicle's path planning software (preferably executing on a CPU complex) that icy conditions exist when the flashing lights are detected. The flashing lights can be identified by a deployed third neural network operating across multiple frames, informing the vehicle's path planning software of the presence (or absence) of the flashing lights. All three neural networks can operate simultaneously, for example, within a DLA and / or on a GPU 408.
[0142] In some examples, the CNN used for facial recognition and owner identification can use data from camera sensors to identify the presence of an authorized driver and / or owner of vehicle 400. A processing engine always on the sensors can be used to unlock the vehicle and turn on the lights when the owner approaches the driver's door, and in safe mode, to disable the vehicle when the owner leaves. In this way, SoC 404 provides security against theft and / or carjacking.
[0143] In another example, the CNN used for emergency vehicle detection and identification can use data from microphone 796 to detect and identify emergency vehicle siren. In contrast to conventional systems that use a general classifier to detect siren and manually extract features, SoC 404 uses a CNN to classify environmental and urban sounds as well as visual data. In a preferred embodiment, the CNN running on the DLA is trained to recognize the relative shut-off rate of emergency vehicles (e.g., by using the Doppler effect). The CNN can also be trained to recognize emergency vehicles specific to the localized area in which the vehicle operates, as identified by GNSS sensor 458. Thus, for example, when operating in Europe, the CNN will seek to detect European siren, and when operating in the United States, the CNN will seek to identify siren only in North America. Once an emergency vehicle is detected, with the assistance of ultrasonic sensor 462, the control program can be used to execute emergency vehicle safety routines, causing the vehicle to slow down, pull over to the side of the road, stop, and / or idle until the emergency vehicle passes.
[0144] The vehicle may include a CPU 418 (e.g., a discrete CPU or dCPU) that can be coupled to the SoC 404 via a high-speed interconnect (e.g., PCIe). The CPU 418 may include, for example, an x86 processor. The CPU 418 can be used to perform any of a wide variety of functions, including, for example, arbitrating the results of potential inconsistencies between ADAS sensors and the SoC 404, and / or monitoring the status and health of the controller 436 and / or the infotainment SoC 430.
[0145] Vehicle 400 may include a GPU 420 (e.g., a discrete GPU or dGPU) that can be coupled to SoC 404 via a high-speed interconnect (e.g., NVIDIA's NVLINK). GPU 420 may provide additional artificial intelligence capabilities, for example by executing redundant and / or different neural networks, and can be used to train and / or update neural networks based on inputs from sensors of vehicle 400 (e.g., sensor data).
[0146] Vehicle 400 may further include a network interface 424, which may include one or more wireless antennas 426 (e.g., one or more wireless antennas for different communication protocols, such as cellular antennas, Bluetooth antennas, etc.). Network interface 424 can be used to enable wireless connectivity via the Internet to the cloud (e.g., with server 478 and / or other network devices), with other vehicles, and / or with computing devices (e.g., a passenger's client device). For communication with other vehicles, a direct link can be established between the two vehicles, and / or an indirect link can be established (e.g., across networks and via the Internet). A direct link can be provided using a vehicle-to-vehicle communication link. The vehicle-to-vehicle communication link can provide vehicle 400 with information about vehicles approaching vehicle 400 (e.g., vehicles in front, to the side, and / or behind vehicle 400). This functionality can be part of vehicle 400's cooperative adaptive cruise control function.
[0147] Network interface 424 may include a SoC that provides modulation and demodulation functions and enables controller 436 to communicate over a wireless network. Network interface 424 may include an RF front-end for up-conversion from baseband to RF and down-conversion from RF to baseband. Frequency conversion can be performed using known processes and / or using a superheterodyne process. In some examples, the RF front-end functionality may be provided by a separate chip. The network interface may include wireless functions for communication via LTE, WCDMA, UMTS, GSM, CDMA2000, Bluetooth, Bluetooth LE, Wi-Fi, Z-Wave, ZigBee, LoRaWAN, and / or other wireless protocols.
[0148] Vehicle 400 may further include data storage 428, which may include off-chip (e.g., off-chip SoC 404) storage devices. Data storage 428 may include one or more storage elements, including RAM, SRAM, DRAM, VRAM, flash memory, hard disk, and / or other components and / or devices capable of storing at least one bit of data.
[0149] Vehicle 400 may further include a GNSS sensor 458. The GNSS sensor 458 (e.g., GPS, assisted GPS sensor, differential GPD (DGPS) sensor, etc.) is used for auxiliary mapping, sensing, occupancy grid generation, and / or path planning functions. Any number of GNSS sensors 458 can be used, including, for example, but not limited to, GPS using a USB connector with an Ethernet-to-serial (RS-232) bridge.
[0150] Vehicle 400 may further include a RADAR sensor 460. The RADAR sensor 460 can be used by vehicle 400 for remote vehicle detection even in dark and / or inclement weather conditions. The RADAR functional safety level may be ASIL B. The RADAR sensor 460 can use CAN and / or bus 402 (e.g., to transmit data generated by the RADAR sensor 460) for control and access to object tracking data, and in some examples, Ethernet access for accessing raw data. A wide variety of RADAR sensor types can be used. For example, and without limitation, the RADAR sensor 460 can be adapted for front, rear, and side RADAR use. In some examples, a pulse Doppler RADAR sensor is used.
[0151] The RADAR sensor 460 can include different configurations, such as long-range with a narrow field of view, short-range with a wide field of view, short-range side coverage, etc. In some examples, the long-range RADAR can be used for adaptive cruise control functions. The long-range RADAR system can provide a wide field of view (e.g., within 250m) achieved through two or more independent scans. The RADAR sensor 460 can help distinguish between stationary and moving objects and can be used by ADAS systems for emergency braking assistance and forward collision warning. The long-range RADAR sensor can include a single-site multi-mode RADAR with multiple (e.g., six or more) fixed RADAR antennas and high-speed CAN and FlexRay interfaces. In an example with six antennas, the four central antennas can create a focused beam pattern designed to record the vehicle 400's surroundings at higher rates with minimal traffic interference from adjacent lanes. The other two antennas can extend the field of view, enabling rapid detection of vehicles entering or leaving the vehicle 400's lane.
[0152] As an example, a mid-range RADAR system can include a range of up to 160m (front) or 80m (rear) and a field of view of up to 42 degrees (front) or 150 degrees (rear). Short-range RADAR systems can include, but are not limited to, RADAR sensors designed to be mounted at both ends of the rear bumper. When mounted at both ends of the rear bumper, such a RADAR sensor system can create two beams that continuously monitor blind spots behind and beside the vehicle.
[0153] Short-range RADAR systems can be used in ADAS systems for blind spot detection and / or lane change assistance.
[0154] Vehicle 400 may further include ultrasonic sensors 462. Ultrasonic sensors 462, which may be positioned at the front, rear, and / or sides of vehicle 400, can be used for parking assistance and / or creating and updating occupancy grids. A wide variety of ultrasonic sensors 462 can be used, and different ultrasonic sensors 462 can be used for different detection ranges (e.g., 2.5m, 4m). Ultrasonic sensors 462 can operate at functional safety level ASIL B.
[0155] Vehicle 400 may include a LIDAR sensor 464. The LIDAR sensor 464 may be used for object and pedestrian detection, emergency braking, collision avoidance, and / or other functions. The LIDAR sensor 464 may be of functional safety level ASIL B. In some examples, vehicle 400 may include multiple LIDAR sensors 464 (e.g., two, four, six, etc.) that can use Ethernet (e.g., to provide data to a Gigabit Ethernet switch).
[0156] In some examples, the LiDAR sensor 464 may be able to provide a list of objects and their distances within a 360-degree field of view. A commercially available LiDAR sensor 464 may have an advertising range of, for example, approximately 100m, with an accuracy of 2cm-3cm, and support for a 100Mbps Ethernet connection. In some examples, one or more non-protruding LiDAR sensors 464 may be used. In such examples, the LiDAR sensor 464 may be implemented as a small device that can be embedded in the front, rear, sides, and / or corners of a vehicle 400. In such examples, the LiDAR sensor 464 may provide a horizontal field of view of up to 120 degrees and a vertical field of view of 35 degrees, even for low-reflectivity objects, with a range of 200m. A front-mounted LiDAR sensor 464 may be configured for a horizontal field of view between 45 degrees and 135 degrees.
[0157] In some examples, LiDAR technologies such as 3D flash LiDAR can also be used. 3D flash LiDAR uses a flash of laser light as the emission source to illuminate the vehicle's surroundings up to approximately 200 meters. A flash LiDAR unit includes a receiver that records the laser pulse propagation time and reflected light on each pixel, which in turn corresponds to the range from the vehicle to the object. Flash LiDAR allows for the generation of highly accurate and distortion-free images of the surrounding environment using each laser flash. In some examples, four flash LiDAR sensors can be deployed, one on each side of the vehicle. Available 3D flash LiDAR systems include solid-state 3D staring array LiDAR cameras (e.g., non-browsing LiDAR devices) without moving parts other than a fan. Flash LiDAR devices can use 5 nanosecond Class I (eye-safe) laser pulses per frame and can capture reflected laser light in the form of a 3D range point cloud and co-registered intensity data. By using flash LiDAR, and because flash LiDAR is a solid-state device with no moving parts, the LiDAR sensor 464 is less susceptible to motion blur, vibration, and / or shock.
[0158] The vehicle may further include an IMU sensor 466. In some examples, the IMU sensor 466 may be located at the center of the rear axle of the vehicle 400. The IMU sensor 466 may include, for example, but not limited to, an accelerometer, a magnetometer, a gyroscope, a magnetic compass, and / or other sensor types. In some examples, such as in a six-axis application, the IMU sensor 466 may include an accelerometer and a gyroscope, while in a nine-axis application, the IMU sensor 466 may include an accelerometer, a gyroscope, and a magnetometer.
[0159] In some embodiments, the IMU sensor 466 can be implemented as a miniature, high-performance GPS-assisted inertial navigation system (GPS / INS) that combines a microelectromechanical system (MEMS) inertial sensor, a high-sensitivity GPS receiver, and an advanced Kalman filter algorithm to provide estimates of position, velocity, and attitude. Thus, in some examples, the IMU sensor 466 can enable the vehicle 400 to estimate heading by directly observing and correlating velocity changes from GPS to the IMU sensor 466 without input from a magnetic sensor. In some examples, the IMU sensor 466 and the GNSS sensor 458 can be combined into a single integrated unit.
[0160] The vehicle may include a microphone 496 placed in and / or around the vehicle 400. Among other things, the microphone 496 may be used for emergency vehicle detection and identification.
[0161] The vehicle may further include any number of camera types, including stereo camera 468, wide-angle camera 470, infrared camera 472, surround camera 474, long-range and / or mid-range camera 498, and / or other camera types. These cameras can be used to capture image data around the entire perimeter of the vehicle 400. The types of cameras used depend on the embodiment and the requirements of the vehicle 400, and any combination of camera types can be used to provide the necessary coverage around the vehicle 400. Furthermore, the number of cameras may vary depending on the embodiment. For example, the vehicle may include six cameras, seven cameras, ten cameras, twelve cameras, and / or another number of cameras. As an example and without limitation, these cameras may support Gigabit Multimedia Serial Link (GMSL) and / or Gigabit Ethernet. Each of the cameras is described herein with respect to... Figure 4A and Figure 4B It was described in more detail.
[0162] Vehicle 400 may further include vibration sensor 442. Vibration sensor 442 can measure vibrations of vehicle components such as axles. For example, changes in vibration can indicate changes in the road surface. In another example, when two or more vibration sensors 442 are used, differences between vibrations can be used to determine friction or slippage on the road surface (e.g., when there is a vibration difference between a power drive shaft and a free-rotating shaft).
[0163] Vehicle 400 may include ADAS system 438. In some examples, ADAS system 438 may include SoC. ADAS system 438 may include autonomous / adaptive / automatic cruise control (ACC), cooperative adaptive cruise control (CACC), forward collision warning (FCW), automatic emergency braking (AEB), lane departure warning (LDW), lane keeping assist (LKA), blind spot warning (BSW), rear cross traffic warning (RCTW), collision warning system (CWS), lane centering (LC) and / or other features and functions.
[0164] The ACC system can use a RADAR sensor 460, a LIDAR sensor 464, and / or a camera. The ACC system can include longitudinal ACC and / or lateral ACC. Longitudinal ACC monitors and controls the distance to vehicles immediately in front of vehicle 400 and automatically adjusts the vehicle speed to maintain a safe distance. Lateral ACC performs distance holding and, if necessary, advises vehicle 400 to change lanes. Lateral ACC is associated with other ADAS applications such as LCA and CWS.
[0165] CACC uses information from other vehicles, which can be received indirectly from other vehicles via a wireless link or network connection (e.g., via the Internet) through network interface 424 and / or wireless antenna 426. Direct links can be provided by vehicle-to-vehicle (V2V) communication links, while indirect links can be infrastructure-to-vehicle (I2V) communication links. Typically, the V2V communication concept provides information about vehicles immediately ahead (e.g., vehicles immediately in front of vehicle 400 and in the same lane), while the I2V communication concept provides information about traffic further ahead. A CACC system can include either or both I2V and V2V information sources. Given information about vehicles ahead of vehicle 400, CACC can be more reliable, and it has the potential to improve traffic flow and reduce road congestion.
[0166] The Forward-Looking Warning (FCW) system is designed to alert the driver to hazards, enabling the driver to take corrective action. The FCW system uses a front-facing camera and / or RADAR sensor 460 coupled to a dedicated processor, DSP, FPGA, and / or ASIC, which is electrically coupled to driver feedback such as a display, speaker, and / or vibrating components. The FCW system can provide warnings in the form of, for example, audible, visual, haptic, and / or rapid braking pulses.
[0167] An AEB (Autonomous Emergency Braking) system detects an impending forward collision with another vehicle or other object and can automatically apply the brakes if the driver does not take corrective action within a specified time or distance parameter. The AEB system can use a front-facing camera and / or RADAR sensor 460 coupled to a dedicated processor, DSP, FPGA, and / or ASIC. When the AEB system detects a hazard, it typically first alerts the driver to take corrective action to avoid a collision, and if the driver does not take corrective action, the AEB system can automatically apply the brakes to attempt to prevent or at least mitigate the effects of the predicted collision. The AEB system may include technologies such as dynamic brake support and / or collision proximity braking.
[0168] The Lane Departure Warning (LDW) system provides visual, auditory, and / or tactile warnings, such as steering wheel or seat vibrations, to alert the driver when the vehicle crosses a lane marking. When the driver indicates intentional lane departure, the LDW system is deactivated by activating a turn signal. The LDW system can utilize a front-facing camera coupled to a dedicated processor, DSP, FPGA, and / or ASIC, which is electrically coupled to driver feedback such as a display, speaker, and / or vibrating components.
[0169] The Lane Keeping Assist (LKA) system is a variant of the Lane Departure Warning (LDW) system. If vehicle 400 begins to leave its lane, the LKA system provides corrective steering input or braking to vehicle 400. The Blind Spot Warning (BSW) system detects and warns the driver of vehicles in the vehicle's blind spot. The BSW system can provide visual, audible, and / or tactile warnings to indicate that merging or changing lanes is unsafe. The system can provide additional warnings when the driver uses a turn signal. The BSW system can use one or more rear-facing cameras and / or one or more RADAR sensors.
[0170] RCTW systems can provide visual, auditory, and / or tactile notifications when an object is detected outside the range of a rear-view camera while the vehicle is reversing at 400 degrees. Some RCTW systems include AEB (Autonomous Emergency Braking) to ensure the application of the vehicle's brakes to avoid a collision. RCTW systems may use one or more rear-view RADAR sensors 460 coupled to a dedicated processor, DSP, FPGA, and / or ASIC, which is electrically coupled to driver feedback such as a display, speaker, and / or vibrating components.
[0171] Conventional ADAS systems can be prone to false positives, which can be annoying and distracting for the driver, but typically not catastrophic, as ADAS systems alert the driver and allow them to determine whether a safe condition truly exists and take appropriate action. However, in an autonomous vehicle 400, in the event of conflicting results, the vehicle 400 itself must decide whether to heed the results from the main computer or auxiliary computer (e.g., the first controller 436 or the second controller 436). For example, in some embodiments, ADAS system 438 may be a backup and / or auxiliary computer used to provide perception information to a backup computer rationality module. The backup computer rationality monitor may run redundant and diverse software on hardware components to detect faults in perception and dynamic driving tasks. Outputs from ADAS system 438 may be provided to a supervisory MCU. If outputs from the main computer and auxiliary computer conflict, the supervisory MCU must determine how to reconcile the conflict to ensure safe operation.
[0172] In some examples, the master computer can be configured to provide a confidence score to the supervisory MCU, indicating the master computer's confidence level in the selected result. If the confidence score exceeds a threshold, the supervisory MCU can follow the master computer's direction regardless of whether the auxiliary computer provides conflicting or inconsistent results. If the confidence score does not meet the threshold and the master and auxiliary computers indicate different results (e.g., conflict), the supervisory MCU can arbitrate between these computers to determine the appropriate result.
[0173] The supervisory MCU can be configured to run a neural network trained and configured to determine the conditions under which the auxiliary computer provides a false alarm based on outputs from both the host and auxiliary computers. Thus, the neural network in the supervisory MCU can learn when the output of the auxiliary computer can be trusted and when it cannot. For example, when the auxiliary computer is a RADAR-based FCW system, the neural network in the supervisory MCU can learn when the FCW system is identifying a metallic object that is not actually dangerous, such as a drain grid or manhole cover that triggers an alarm. Similarly, when the auxiliary computer is a camera-based LDW system, the neural network in the supervisory MCU can learn to ignore the LDW when a cyclist or pedestrian is present and lane departure is actually the safest strategy. In embodiments that include a neural network running on the supervisory MCU, the supervisory MCU may include at least one of a DLA or GPU suitable for running the neural network using associated memory. In a preferred embodiment, the supervisory MCU may include components of SoC 404 and / or be included as components of SoC 404.
[0174] In other examples, ADAS system 438 may include an auxiliary computer that performs ADAS functions using conventional computer vision rules. This allows the auxiliary computer to use classic computer vision rules (if-then), and the presence of neural networks in the supervising MCU can improve reliability, safety, and performance. For example, diverse implementations and intentional non-identity make the entire system more fault-tolerant, especially for failures caused by software (or software-hardware interface) functionality. For instance, if a software vulnerability or bug exists in the software running on the host computer and non-identical software code running on the auxiliary computer provides the same overall result, the supervising MCU can be more confident that the overall result is correct and that the vulnerability in the software or hardware on the host computer does not cause a substantial error.
[0175] In some examples, the output of ADAS system 438 can be fed to the perception block and / or the dynamic driving task block of the main computer. For example, if ADAS system 438 issues a forward collision warning because an object is immediately in front, the perception block can use this information when identifying the object. In other examples, the assistance computer can have its own neural network, which is trained and thus reduces the risk of false positives as described herein.
[0176] Vehicle 400 may further include an infotainment SoC 430 (e.g., an in-vehicle infotainment system (IVI)). Although illustrated and described as an SoC, the infotainment system may not be an SoC and may include two or more discrete components. The infotainment SoC 430 may include a combination of hardware and software that can be used to provide vehicle 400 with audio (e.g., music, personal digital assistant, navigation instructions, news, radio, etc.), video (e.g., TV, movies, streaming media, etc.), telephone (e.g., hands-free calling), network connectivity (e.g., LTE, Wi-Fi, etc.) and / or information services (e.g., navigation system, rear parking assistance, radio data system, vehicle-related information such as fuel level, total coverage distance, brake fuel level, fuel level, door opening / closing, air filter information, etc.). For example, the infotainment SoC 430 may include a radio, disc player, navigation system, video player, USB and Bluetooth connectivity, in-vehicle computer, in-vehicle entertainment, Wi-Fi, steering wheel audio controls, hands-free voice controls, head-up display (HUD), HMI display 434, telematics device, control panel (e.g., for controlling and / or interacting with various components, features, and / or systems) and / or other components. The infotainment SoC 430 may further be used to provide information (e.g., visual and / or auditory) to the vehicle's users, such as information from ADAS system 438, autonomous driving information such as planned vehicle maneuvers, trajectories, surrounding environment information (e.g., intersection information, vehicle information, road information, etc.), and / or other information.
[0177] The infotainment SoC 430 may include GPU functionality. The infotainment SoC 430 can communicate with other devices, systems, and / or components of the vehicle 400 via bus 402 (e.g., CAN bus, Ethernet, etc.). In some examples, the infotainment SoC 430 may be coupled to a supervisory MCU, allowing the GPU of the infotainment system to perform some autonomous driving functions in the event of a failure of the main controller 436 (e.g., the primary and / or backup computer of the vehicle 400). In such an example, the infotainment SoC 430 may place the vehicle 400 into a driver-safe parking mode as described herein.
[0178] Vehicle 400 may further include an instrument cluster 432 (e.g., a digital instrument panel, electronic instrument cluster, digital instrument panel, etc.). The instrument cluster 432 may include a controller and / or a supercomputer (e.g., a discrete controller or supercomputer). The instrument cluster 432 may include a set of instruments such as a speedometer, fuel level, oil pressure, tachometer, odometer, turn indicator, shift position indicator, seatbelt warning light, parking brake warning light, engine malfunction indicator, airbag (SRS) system information, lighting controls, safety system controls, navigation information, etc. In some examples, information may be displayed and / or shared between the infotainment SoC 430 and the instrument cluster 432. In other words, the instrument cluster 432 may be included as part of the infotainment SoC 430, or vice versa.
[0179] Figure 4D For cloud-based servers and according to some embodiments of this disclosure Figure 4A This is a system diagram illustrating communication between example autonomous vehicles 400. System 476 may include server 478, network 490, and vehicles including vehicle 400. Server 478 may include multiple GPUs 484(A)-484(H) (collectively referred to herein as GPU 484), PCIe switches 482(A)-482(H) (collectively referred to herein as PCIe switch 482), and / or CPUs 480(A)-480(B) (collectively referred to herein as CPU 480). GPUs 484, CPUs 480, and PCIe switches may be interconnected with high-speed interconnects and / or PCIe connections 486, such as, but not limited to, NVLink interface 488 developed by NVIDIA. In some examples, GPUs 484 are connected via NVLink and / or NVSwitch SoCs, and GPUs 484 and PCIe switches 482 are connected via PCIe interconnects. Although eight GPUs 484, two CPUs 480, and two PCIe switches are shown in the diagram, this is not intended to be limiting. Depending on the embodiment, each of the servers 478 may include any number of GPUs 484, CPUs 480, and / or PCIe switches. For example, each of the servers 478 may include eight, sixteen, thirty-two, and / or more GPUs 484.
[0180] Server 478 can receive image data from vehicles via network 490, representing images of unexpected or altered road conditions such as recently commenced roadworks. Server 478 can also transmit neural network 492, updated neural network 492, and / or map information 494, including information about traffic and road conditions, to vehicles via network 490. Updates to map information 494 may include updates to HD map 422, such as information about construction sites, potholes, bends, floods, or other obstacles. In some examples, neural network 492, updated neural network 492, and / or map information 494 may have been generated from new training and / or data received from any number of vehicles in the environment, and / or based on experience gained from training performed at a data center (e.g., using server 478 and / or other servers).
[0181] Server 478 can be used to train machine learning models (e.g., neural networks) based on training data. Training data can be generated by the vehicle and / or generated in a simulation (e.g., using a game engine). In some examples, the training data is labeled (e.g., where the neural network benefits from supervised learning) and / or undergoes other preprocessing, while in other examples, the training data is not labeled and / or preprocessed (e.g., where the neural network does not require supervised learning). Training can be performed according to any one or more classes of machine learning techniques, including but not limited to: supervised training, semi-supervised training, unsupervised training, self-learning, reinforcement learning, joint learning, transfer learning, feature learning (including principal component analysis and cluster analysis), multilinear subspace learning, manifold learning, representation learning (including alternative dictionary learning), rule-based machine learning, anomaly detection, and any variations or combinations thereof. Once the machine learning model is trained, it can be used by the vehicle (e.g., transmitted to the vehicle via network 490), and / or the machine learning model can be used by server 478 to remotely monitor the vehicle.
[0182] In some examples, server 478 can receive data from vehicles and apply that data to state-of-the-art real-time neural networks for real-time intelligent inference. Server 478 may include a deep learning supercomputer powered by GPU 484 and / or a dedicated AI computer, such as the DGX and DGX Station machines developed by NVIDIA. However, in some examples, server 478 may include a deep learning infrastructure in a data center that uses only CPU power.
[0183] The deep learning infrastructure of server 478 may be capable of rapid, real-time inference and can be used to assess and verify the health status of the processor, software, and / or associated hardware in vehicle 400. For example, the deep learning infrastructure may receive periodic updates from vehicle 400, such as image sequences and / or objects located in those image sequences by vehicle 400 (e.g., via computer vision and / or other machine learning object classification techniques). The deep learning infrastructure may run its own neural network to identify objects and compare them with objects identified by vehicle 400. If the results do not match and the infrastructure concludes that the AI in vehicle 400 has malfunctioned, then server 478 may transmit a signal to vehicle 400 instructing the vehicle 400's fail-safe computer to take control, notify passengers, and complete a safe stopping operation.
[0184] For inference, server 478 may include GPU 484 and one or more programmable inference accelerators (such as NVIDIA's TensorRT 3). The combination of a GPU-powered server and inference acceleration enables real-time response. In other examples, such as where performance is less critical, CPU, FPGA, and other processor-powered servers can be used for inference.
[0185] Example computing device
[0186] Figure 5 The block diagram is provided for an example computing device 500 suitable for implementing some embodiments of the present disclosure. The computing device 500 may include an interconnect system 502 directly or indirectly coupled to the following devices: memory 504, one or more central processing units (CPUs) 506, one or more graphics processing units (GPUs) 508, a communication interface 510, input / output (I / O) ports 512, input / output components 514, a power supply 516, one or more presentation components 518 (e.g., a display), and one or more logic units 520. In at least one embodiment, the computing device 500 may include one or more virtual machines (VMs), and / or any component thereof may include virtual components (e.g., virtual hardware components). For a non-limiting example, one or more GPUs 508 may include one or more vGPUs, one or more CPUs 506 may include one or more vCPUs, and / or one or more logic units 520 may include one or more virtual logic units. Therefore, computing device 500 may include discrete components (e.g., a complete GPU dedicated to computing device 500), virtual components (e.g., a portion of the GPU dedicated to computing device 500), or a combination thereof.
[0187] although Figure 5The various boxes are shown connected via an interconnect system 502 with wiring, but this is not intended to be limiting and is merely for clarity. For example, in some embodiments, a presentation component 518, such as a display device, can be considered an I / O component 514 (e.g., if the display is a touchscreen). As another example, CPU 506 and / or GPU 508 may include memory (e.g., memory 504 may represent a storage device other than the memory of GPU 508, CPU 506, and / or other components). In other words, Figure 5 The computing devices mentioned are merely illustrative. No distinction is made between categories such as "workstation," "server," "laptop," "desktop," "tablet," "client device," "mobile device," "handheld device," "game console," "electronic control unit (ECU)," "virtual reality system," and / or other device or system types, as all of these are considered within the same category. Figure 5 Within the scope of computing devices.
[0188] Interconnect system 502 may represent one or more links or buses, such as address buses, data buses, control buses, or combinations thereof. Interconnect system 502 may include one or more link or bus types, such as Industry Standard Architecture (ISA) bus, Extended Industry Standard Architecture (EISA) bus, Video Electronics Standards Association (VESA) bus, Peripheral Component Interconnect (PCI) bus, Peripheral Component Interconnect Fast (PCIe) bus, and / or another type of bus or link. In some embodiments, there is a direct connection between components. As an example, CPU 506 may be directly connected to memory 504. Furthermore, CPU 506 may be directly connected to GPU 508. In cases where there is a direct or point-to-point connection between components, interconnect system 502 may include a PCIe link to perform the connection. In these examples, a PCI bus is not required in computing device 500.
[0189] Memory 504 may include any medium of a wide variety of computer-readable media. Computer-readable media can be any available medium that can be accessed by computing device 500. Computer-readable media may include volatile and non-volatile media, as well as removable and non-removable media. For example and without limitation, computer-readable media may include computer storage media and communication media.
[0190] Computer storage media may include volatile and non-volatile media and / or removable and non-removable media, implemented in any way or by any method or technique for storing information such as computer-readable instructions, data structures, program modules, and / or other data types. For example, memory 504 may store computer-readable instructions (e.g., representing programs and / or program elements, such as an operating system). Computer storage media may include, but is not limited to, RAM, ROM, EEPROM, flash memory or other storage technologies, CD-ROM, digital versatile disc (DVD) or other optical disc storage devices, magnetic tape cassettes, magnetic tape, disk storage devices or other magnetic storage devices, or any other medium that can be used to store desired information and can be accessed by computing device 500. As used herein, computer storage media does not include the signal itself.
[0191] Computer storage media may contain computer-readable instructions, data structures, program modules, and / or other data types in modulated data signals such as carrier waves or other transmission mechanisms, and include any information transport medium. The term "modulated data signal" can refer to a signal whose characteristics are set or altered in a manner that encodes information into that signal. For example and without limitation, computer storage media may include wired media such as wired networks or direct wired connections, and wireless media such as sound, RF, infrared, and other wireless media. Any combination of the above should also be included within the scope of computer-readable media.
[0192] CPU 506 may be configured to execute at least some of computer-readable instructions to control one or more components of computing device 500 to perform one or more of the methods and / or processes described herein. Each of CPU 506 may include one or more cores (e.g., one, two, four, eight, twenty-eight, seventy-two, etc.) capable of processing a large number of software threads simultaneously. CPU 506 may include any type of processor and may include different types of processors depending on the type of computing device 500 implemented (e.g., processors with fewer cores for mobile devices and processors with more cores for servers). For example, depending on the type of computing device 500, the processor may be an advanced RISC mechanism (ARM) processor implemented using Reduced Instruction Set Computing (RISC) or an x86 processor implemented using Complex Instruction Set Computing (CISC). In addition to one or more microprocessors or supplementary coprocessors such as math coprocessors, computing device 500 may also include one or more CPUs 506.
[0193] In addition to or replacing CPU 506, GPU 508 may also be configured to execute at least some computer-readable instructions to control one or more components of computing device 500 to perform one or more of the methods and / or processes described herein. One or more GPUs 508 may be integrated GPUs (e.g., having one or more CPUs 506) and / or one or more GPUs 508 may be discrete GPUs. In embodiments, one or more GPUs 508 may be coprocessors of one or more CPUs 506. Computing device 500 may use GPU 508 to render graphics (e.g., 3D graphics) or perform general-purpose computing. For example, GPU 508 may be used for general-purpose computing on a GPU (GPGPU). GPU 508 may include hundreds or thousands of cores capable of processing hundreds or thousands of software threads simultaneously. GPU 508 may generate pixel data for outputting an image in response to rendering commands (e.g., rendering commands received via a host interface from CPU 506). GPU 508 may include graphics memory, such as display memory, for storing pixel data or any other suitable data (e.g., GPGPU data). Display memory may be included as part of memory 504. GPU 508 may include two or more GPUs operating in parallel (e.g., via links). The links may connect the GPUs directly (e.g., using NVLINK) or via a switch (e.g., using NVSwitch). When combined, each GPU 508 may generate different portions of pixel data or GPGPU data for different outputs (e.g., the first GPU for the first image, the second GPU for the second image). Each GPU may include its own memory or may share memory with other GPUs.
[0194] In addition to or replacing CPU 506 and / or GPU 508, logic unit 520 may be configured to execute at least some computer-readable instructions to control one or more components of computing device 500 to perform one or more methods and / or processes described herein. In embodiments, CPU 506, GPU 508, and / or logic unit 520 may execute any combination of methods, processes, and / or portions thereof discretely or jointly. One or more logic units 520 may be part of and / or integrated into one or more CPUs 506 and / or one or more GPUs 508, and / or one or more logic units 520 may be discrete components of CPU 506 and / or GPU 508 or otherwise external thereto. In embodiments, one or more logic units 520 may be processors of one or more CPUs 506 and / or one or more GPUs 508.
[0195] Examples of logic unit 520 include one or more processing cores and / or components thereof, such as data processing unit (DPU), tensor core (TC), tensor processing unit (TPU), pixel vision core (PVC), vision processing unit (VPU), graphics processing cluster (GPC), texture processing cluster (TPC), streaming multiprocessor (SM), tree traversal unit (TTU), artificial intelligence accelerator (AIA), deep learning accelerator (DLA), arithmetic logic unit (ALU)), application-specific integrated circuit (ASIC), floating-point unit (FPU), input / output (I / O) element, peripheral component interconnect (PCI) or peripheral component interconnect fast (PCIe) element, etc.
[0196] The communication interface 510 may include one or more receivers, transmitters, and / or transceivers that enable the computing device 500 to communicate with other computing devices via electronic communication networks, including wired and / or wireless communications. The communication interface 510 may include components and functions that enable communication via any of several different networks, such as wireless networks (e.g., Wi-Fi, Z-Wave, Bluetooth, Bluetooth LE, ZigBee, etc.), wired networks (e.g., communication via Ethernet or InfiniBand), low-power wide area networks (e.g., LoRaWAN, SigFox, etc.), and / or the Internet. In one or more embodiments, the logic unit 520 and / or the communication interface 510 may include one or more data processing units (DPUs) to directly transmit data received via a network and / or via interconnect system 502 to one or more GPUs 508 (e.g., memory within GPU 508).
[0197] I / O port 512 enables computing device 500 to be logically coupled to other devices, including I / O component 514, presentation component 518, and / or other components, some of which may be built into (e.g., integrated into) computing device 500. Illustrative I / O component 514 includes microphones, mice, keyboards, joysticks, game pads, game controllers, satellite dish antennas, browsers, printers, wireless devices, and so on. I / O component 514 can provide a Natural User Interface (NUI) for processing user-generated air gestures, voice, or other physiological input. In some instances, the input may be transmitted to appropriate network elements for further processing. The NUI can implement any combination of voice recognition, stylus recognition, facial recognition, biometric recognition, on-screen and adjacent-screen gesture recognition, air gestures, head and eye tracking, and touch recognition associated with the display of computing device 500 (as described in more detail in this disclosure). Computing device 500 may include depth cameras such as stereo camera systems, infrared camera systems, RGB camera systems, touchscreen technology, and combinations thereof for gesture detection and recognition. In addition, computing device 500 may include an accelerometer or gyroscope that enables motion detection (e.g., as part of an inertial measurement unit (IMU)). In some examples, the output of the accelerometer or gyroscope may be used by computing device 500 to render immersive augmented reality or virtual reality.
[0198] Power supply 516 may include a hard-wired power supply, a battery power supply, or a combination thereof. Power supply 516 may supply power to computing device 500 so that the components of computing device 500 can operate.
[0199] The presentation component 518 may include a display (such as a monitor, touch screen, television screen, head-up display (HUD), other display types, or combinations thereof), speakers, and / or other presentation components. The presentation component 518 may receive data from other components (such as GPU 508, CPU 506, DPU, etc.) and output that data (e.g., as images, videos, sounds, etc.).
[0200] Example Data Center
[0201] Figure 6 An example data center 600 is shown, which can be used in at least one embodiment of this disclosure. Data center 600 may include a data center infrastructure layer 610, a framework layer 620, a software layer 630, and an application layer 640.
[0202] like Figure 6As shown, the data center infrastructure layer 610 may include a resource coordinator 612, grouped computing resources 614, and node computing resources (“nodes CR”) 616(1)-616(N), where “N” represents any complete positive integer. In at least one embodiment, nodes CR 616(1)-616(N) may include, but are not limited to, any number of central processing units (CPUs) or other processors (including DPUs, accelerators, field-programmable gate arrays (FPGAs), graphics processing units or graphics processing units (GPUs), etc.), memory devices (e.g., dynamic read-only memory), storage devices (e.g., solid-state drives or disk drives), network input / output (NW I / O) devices, network switches, virtual machines (VMs), power modules, and cooling modules, etc. In some embodiments, one or more nodes CR 616(1)-616(N) may correspond to servers having one or more of the aforementioned computing resources. In addition, in some embodiments, nodes CR616(1)-616(N) may include one or more virtual components, such as vGPU, vCPU, etc., and / or one or more of nodes CR616(1)-616(N) may correspond to virtual machines (VMs).
[0203] In at least one embodiment, the grouped computing resources 614 may include individual groups (not shown) of nodes CR616 housed in one or more racks, or a plurality of racks (also not shown) housed in data centers in various geographic locations. Individual groups of nodes CR616 within the grouped computing resources 614 may include computing, networking, memory, or storage resources that can be configured or allocated to support groups of one or more workloads. In at least one embodiment, several nodes CR616, including CPUs, GPUs, DPUs, and / or other processors, may be grouped within one or more racks to provide computing resources to support one or more workloads. One or more racks may also include any number of power modules, cooling modules, and / or network switches in any combination.
[0204] Resource coordinator 612 can be configured or otherwise controlled to control one or more nodes CR616(1)-616(N) and / or groups of computing resources 614. In at least one embodiment, resource coordinator 612 may include a Software Design Infrastructure (SDI) management entity for data center 600. Resource coordinator 612 may include hardware, software, or some combination thereof.
[0205] In at least one embodiment, such as Figure 6As shown, framework layer 620 may include a job scheduler 633, a configuration manager 634, a resource manager 636, and a distributed file system 638. Framework layer 620 may include a framework of software 632 supporting software layer 630 and / or one or more applications 642 of application layer 640. Software 632 or application 642 may respectively include web-based service software or applications, such as service software or applications provided by Amazon Web Services, Google Cloud, and Microsoft Azure. Framework layer 620 may be, but is not limited to, a free and open-source software web application framework, such as Apache Spark, which can utilize distributed file system 638 for large-scale data processing (e.g., "big data"). TM (Hereinafter referred to as "Spark"). In at least one embodiment, the job scheduler 633 may include a Spark driver for facilitating the scheduling of workloads supported by various layers of data center 600. In at least one embodiment, the configuration manager 634 may be able to configure different layers, such as software layer 630 and framework layer 620 including Spark and a distributed file system 638 for supporting large-scale data processing. The resource manager 636 is able to manage cluster or grouped computing resources mapped to or allocated for supporting distributed file system 638 and job scheduler 633. In at least one embodiment, cluster or grouped computing resources may include grouped computing resources 614 at data center infrastructure layer 610. The resource manager 636 may coordinate with resource coordinator 612 to manage these mapped or allocated computing resources.
[0206] In at least one embodiment, the software 632 included in the software layer 630 may include software used by at least a portion of the nodes CR616(1)-616(N), the grouped computing resources 614, and / or the distributed file system 638 of the framework layer 620. One or more types of software may include, but are not limited to, Internet web page search software, email virus browsing software, database software, and streaming video content software.
[0207] In at least one embodiment, the application layer 640 may include one or more applications 642 that can be used by at least a portion of the nodes CR616(1)-616(N), the grouped computing resources 614, and / or the distributed file system 638 of the framework layer 620. The one or more types of applications may include, but are not limited to, any number of genomics applications, cognitive computing and machine learning applications, including training or inference software, machine learning framework software (e.g., PyTorch, TensorFlow, Caffe, etc.), and / or other machine learning applications used in conjunction with one or more embodiments.
[0208] In at least one embodiment, any of the configuration manager 634, resource manager 636, and resource coordinator 612 can perform any number and type of self-modification actions based on any amount and type of data acquired in any technically feasible manner. Self-modification actions can mitigate potentially poor configuration decisions by data center operators of data center 600 and can prevent underutilization and / or skewed portions of the data center.
[0209] Data center 600 may include tools, services, software, or other resources for training one or more machine learning models or using one or more machine learning models to predict or infer information according to one or more embodiments described herein. For example, a machine learning model can be trained by calculating weight parameters based on a neural network architecture using the software and computing resources described herein with respect to data center 600. In at least one embodiment, by using weight parameters calculated through one or more training techniques, the resources described herein with respect to data center 600 can be used to infer or predict information using trained machine learning models corresponding to one or more neural networks, such as, but not limited to, those described herein.
[0210] In at least one embodiment, the data center 600 may use a CPU, application-specific integrated circuit (ASIC), GPU, FPGA, and / or other hardware (or corresponding virtual computing resources) to perform training and / or inference. Furthermore, one or more software and / or hardware resources described in this disclosure may be configured as a service to allow a user to train or perform information inference, such as image recognition, speech recognition, or other artificial intelligence services.
[0211] Example network environment
[0212] A network environment suitable for implementing embodiments of this disclosure may include one or more client devices, servers, network-attached storage (NAS), other backend devices, and / or other device types. Client devices, servers, and / or other device types (e.g., each device) may... Figure 5 This is implemented on one or more instances of computing device 500—for example, each device may include similar components, features, and / or functions of computing device 500. Furthermore, in the case of implementing backend devices (e.g., servers, NAS, etc.), the backend devices may be included as part of data center 600, examples of which are described herein. Figure 6 To describe in more detail.
[0213] Components of a network environment can communicate with each other via a network, which can be wired, wireless, or both. A network can include multiple networks, or networks within multiple networks. For example, a network can include one or more wide area networks (WANs), one or more local area networks (LANs), one or more public networks (such as the Internet and / or the Public Switched Telephone Network (PSTN)), and / or one or more private networks. In cases where the network includes a wireless telecommunications network, components such as base stations, communication towers, or even access points (and other components) can provide wireless connectivity.
[0214] A compatible network environment may include one or more peer-to-peer network environments (in which case the server may not be included in the network environment) and one or more client-server network environments (in which case one or more servers may be included in the network environment). In a peer-to-peer network environment, the server functionality described herein can be implemented on any number of client devices.
[0215] In at least one embodiment, the network environment may include one or more cloud-based network environments, distributed computing environments, combinations thereof, etc. A cloud-based network environment may include a framework layer, a job scheduler, a resource manager, and a distributed file system implemented on one or more servers, which may include one or more core network servers and / or edge servers. The framework layer may include a framework for supporting software at the software layer and / or one or more applications at the application layer. The software or applications may respectively include network-based service software or applications. In embodiments, one or more client devices may use the network-based service software or applications (e.g., by accessing the service software and / or applications via one or more application programming interfaces (APIs)). The framework layer may be, but is not limited to, a type of free and open-source software network application framework, such as one that can use a distributed file system for large-scale data processing (e.g., "big data").
[0216] A cloud-based network environment can provide cloud computing and / or cloud storage for any combination of the computing and / or data storage functions (or one or more portions thereof) described herein. Any of these various functions can be distributed across multiple locations from a central or core server (e.g., distributed across one or more data centers at the state, region, country, global, etc.). If the connection to a user (e.g., a client device) is relatively close to an edge server, the core server can assign at least a portion of the functionality to the edge server. A cloud-based network environment can be private (e.g., limited to a single organization), public (e.g., available to many organizations), and / or a combination thereof (e.g., a hybrid cloud environment).
[0217] Client devices may include those described in this article. Figure 5 The example computing device 500 described includes at least some components, features, and functions. By way of example and not limitation, the client device can be a personal computer (PC), laptop computer, mobile device, smartphone, tablet computer, smartwatch, wearable computer, personal digital assistant (PDA), MP3 player, virtual reality headset, global positioning system (GPS) or device, video player, camera, surveillance equipment or system, vehicle, ship, aircraft, virtual machine, drone, robot, handheld communication device, hospital equipment, gaming device or system, entertainment system, in-vehicle computer system, embedded system controller, remote control, electrical appliance, consumer electronics device, workstation, edge device, any combination of these described devices, or any other suitable device.
[0218] This disclosure can be described in the general context of machine-usable instructions or computer code, including computer-executable instructions such as program modules, which are executed by a computer or other machine such as a personal digital assistant or other handheld device. Typically, a program module, including routines, programs, objects, components, data structures, etc., refers to code that performs a specific task or implements a specific abstract data type. This disclosure can be practiced in a wide variety of system configurations, including handheld devices, consumer electronics, general-purpose computers, more specialized computing devices, etc. This disclosure can also be practiced in distributed computing environments where tasks are performed by remote processing devices linked via a communication network.
[0219] As used herein, the phrase “and / or” relating to two or more elements should be interpreted as referring to only one element or a combination of elements. For example, “element A, element B, and / or element C” could include only element A, only element B, only element C, element A and element B, element A and element C, element B and element C, or element A, B, and C. Furthermore, “at least one of element A or element B” could include at least one of element A, at least one of element B, or at least one of element A and at least one of element B. Further, “at least one of element A and element B” could include at least one of element A, at least one of element B, or at least one of element A and at least one of element B. Moreover, the use of the term “based on” should not be interpreted as “based on only” or “based on only”. Rather, a first element “based on” a second element includes instances where the first element is based on the second element but may also be based on one or more additional elements.
[0220] This document describes in detail the subject matter of this disclosure to satisfy legal requirements. However, the description itself is not intended to limit the scope of this disclosure. Rather, the inventors have envisioned that the claimed subject matter may also be embodied in other ways to include steps different from or similar combinations of steps described herein in conjunction with other current or future techniques. Furthermore, although the terms “step” and / or “block” may be used herein to imply different elements of the method employed, these terms should not be construed as implying any particular order among or between the various steps disclosed herein, unless the order of the steps is explicitly described.
[0221] For example, the present invention is illustrated according to various aspects described below. For convenience, various examples of aspects of the present invention are described as numbered examples (1, 2, 3, etc.). These are provided as examples and do not limit the present invention. Unless the context otherwise indicates, aspects of the various implementations described herein may be omitted, substituted with aspects of other implementations, or combined with aspects of other implementations. For example, one or more aspects of Example 1 below may be omitted, replaced with one or more aspects of another example (e.g., Example 2), or combined with aspects of another example. The following is a non-limiting overview of some example implementations presented herein.
[0222] Example 1. A method comprising:
[0223] Obtain a data record, the data record comprising one or more frames, the one or more frames comprising corresponding frame data;
[0224] The frame data of one or more frames is compared with an annotated dataset, the annotated dataset including known features and annotations corresponding to the known features;
[0225] One or more features in the one or more frames are identified, at least based on the comparison between the frame data and the annotated dataset;
[0226] Determine a subset of the one or more frames, the subset comprising the one or more features associated with one or more operational fields; and
[0227] The subset of frames is provided as training data to the detection model.
[0228] The method described in Example 1 also includes:
[0229] After determining the subset of frames, filtering includes the subset of frames containing the operational domain of errors. These errors include one or more of intrinsic errors, extrinsic errors, or random errors.
[0230] According to the method described in Example 1, the one or more operational domains include scenes present in the one or more frames.
[0231] According to the method described in Example 1, the data record is obtained using one or more sensors.
[0232] According to the method of Example 1, the annotation includes at least one of a bounding shape or a polyline corresponding to the known feature corresponding to the annotation.
[0233] According to the method described in Example 1, the one or more operation domains are described using corresponding operation domain definitions.
[0234] According to the method described in Example 1, the annotated dataset corresponds to a road map comprising multiple road segments, each of the multiple road segments comprising a ground truth label corresponding to one or more features associated with the respective road segment.
[0235] The method described in Example 1 also includes:
[0236] After identifying one or more features in one or more frames, additional operations are performed to identify additional features not included in the frame data.
[0237] Example 2. A system comprising:
[0238] One or more processors, said one or more processors being configured to cause the execution of operations, said operations including:
[0239] Data records comprising one or more frames are obtained using one or more sensors, wherein the one or more frames include corresponding frame data;
[0240] The frame data of one or more frames is compared with an annotated dataset, which includes known features and annotations corresponding to the known features;
[0241] One or more features in the one or more frames are identified, at least based on the comparison between the frame data and the annotated dataset;
[0242] Determine a subset of the one or more frames that includes the one or more features associated with one or more operational domains;
[0243] The subset of frames is filtered to remove one or more frames associated with an incorrect projection of the associated operational domain; and
[0244] The subset of frames is used as training data to update one or more parameters of the detection model.
[0245] According to the system described in Example 2, the one or more operational domains include scenes present in the one or more frames.
[0246] According to the system described in Example 2, the annotation includes at least one of a bounding shape or a polyline corresponding to the known feature corresponding to the annotation.
[0247] According to the system described in Example 2, the one or more operation domains are described using corresponding operation domain definitions.
[0248] According to the system described in Example 2, the annotated dataset corresponds to a road map comprising multiple road segments, each of the multiple road segments including a ground truth label corresponding to one or more features associated with the respective road segment.
[0249] According to the system described in Example 2, the subset of frames is filtered at least based on truth data.
[0250] According to the system described in Example 2, the operation further includes:
[0251] After identifying one or more features in one or more frames, additional operations are performed to identify additional features not included in the frame data.
[0252] According to the system described in Example 2, the system is included in at least one of the following:
[0253] Control systems for autonomous or semi-autonomous machines;
[0254] Sensing systems for autonomous or semi-autonomous machines;
[0255] A system used to perform simulation operations;
[0256] Systems used to perform digital twin operations;
[0257] A system for performing optical transmission simulation;
[0258] A system for collaborative content creation of 3D assets;
[0259] A system used to perform deep learning operations;
[0260] A system for presenting at least one of augmented reality content, virtual reality content, or mixed reality content;
[0261] A system used to host one or more real-time streaming applications;
[0262] Systems implemented using edge devices;
[0263] Systems implemented using robots;
[0264] A system for performing conversational AI operations;
[0265] A system for performing one or more generative AI operations;
[0266] A system that implements one or more large language model LLMs;
[0267] A system for generating synthetic data;
[0268] A system containing one or more virtual machines (VMs);
[0269] A system that is at least partially implemented in a data center; or
[0270] A system that utilizes cloud computing resources at least in part.
[0271] Example 3. One or more processors, including:
[0272] Processing circuitry, the processing circuitry being used to perform operations, the operations including:
[0273] Obtain a data record comprising one or more frames, wherein the one or more frames include corresponding frame data;
[0274] The frame data of one or more frames is compared with an annotated dataset, which includes known features and annotations corresponding to the known features;
[0275] Based at least on the comparison between the frame data and the annotated dataset, identify one or more features in the one or more frames; and
[0276] Determine a subset of the one or more frames, the subset comprising the one or more features associated with one or more operational fields.
Claims
1. A method comprising: obtaining a data record comprising one or more frames, the one or more frames comprising corresponding frame data; comparing the frame data of the one or more frames to an annotated data set, the annotated data set comprising known features and annotations corresponding to the known features; identifying one or more features in the one or more frames based at least on the comparison between the frame data and the annotated data set; determining a subset of the one or more frames, the subset comprising the one or more features associated with one or more operational domains; and providing the subset of frames as training data to a detection model.
2. The method of claim 1, further comprising: after determining the subset of frames, filtering the subset of frames to include operational domains that contain errors.
3. The method of claim 2, wherein the errors comprise one or more of intrinsic errors, extrinsic errors, or random errors.
4. The method of claim 1, wherein the one or more operational domains comprise scenes present in the one or more frames.
5. The method of claim 1, wherein the data record is obtained using one or more sensors.
6. The method of claim 1, wherein the annotations comprise at least one of a bounding shape or a polyline corresponding to the known features corresponding to the annotations.
7. The method of claim 1, wherein the one or more operational domains are described using corresponding operational domain definitions.
8. The method of claim 1, wherein the annotated data set corresponds to a road map comprising a plurality of road segments.
9. The method of claim 8, wherein individual road segments of the plurality of road segments comprise ground truth labels corresponding to the one or more features associated with the individual road segments.
10. The method of claim 1, further comprising: after identifying one or more features in the one or more frames, performing additional operations to identify additional features not included in the frame data.
11. A system comprising: one or more processors to cause performance of operations comprising: obtaining a data record comprising one or more frames using one or more sensors, the one or more frames comprising corresponding frame data; comparing the frame data of the one or more frames to an annotated data set, the annotated data set comprising known features and annotations corresponding to the known features; identifying one or more features in the one or more frames based at least on the comparison between the frame data and the annotated data set; determining a subset of the one or more frames comprising the one or more features associated with one or more operational domains; filtering the subset of frames to remove one or more frames associated with incorrect projections of associated operational domains; and updating one or more parameters of a detection model using the subset of frames as training data.
12. The system of claim 11, wherein the one or more operational domains include a scene present in the one or more frames.
13. The system of claim 11, wherein the annotations include at least one of a bounding shape or a polyline corresponding to the known features corresponding to the annotations.
14. The system of claim 11, wherein the one or more operational domains are described using corresponding operational domain definitions.
15. The system of claim 11, wherein the annotated dataset corresponds to a road map including a plurality of road segments.
16. The system of claim 15, wherein each road segment of the plurality of road segments includes a ground truth label corresponding to the one or more features associated with the each road segment.
17. The system of claim 11, wherein the subset of frames is filtered based at least on ground truth data.
18. The system of claim 11, the operations further comprising: after identifying one or more features in the one or more frames, performing additional operations to identify additional features not included in the frame data.
19. The system of claim 11, wherein the system is included in at least one of: a control system for an autonomous or semi-autonomous machine; a perception system for an autonomous or semi-autonomous machine; a system for performing simulation operations; a system for performing digital twin operations; a system for performing optical transport simulations; a system for performing collaborative content creation of 3D assets; a system for performing deep learning operations; a system for presenting at least one of augmented reality content, virtual reality content, or mixed reality content; a system for hosting one or more real-time streaming applications; a system implemented using edge devices; a system implemented using robots; a system for performing conversational AI operations; a system for performing one or more generative AI operations; a system implementing one or more large language models (LLMs); a system for generating synthetic data; a system including one or more virtual machines (VMs); a system implemented at least partially in a data center; or a system implemented at least partially using cloud computing resources.
20. One or more processors comprising: processing circuitry to perform operations comprising: obtaining a data record including one or more frames, the one or more frames including corresponding frame data; comparing the frame data of the one or more frames to an annotated dataset, the annotated dataset including known features and annotations corresponding to the known features; based at least on the comparison between the frame data and the annotated dataset, identifying one or more features in the one or more frames; and determining a subset of the one or more frames, the subset including the one or more features associated with one or more operational domains.
Citation Information
Patent Citations
Method for programmable timeouts of tree traversal mechanisms in hardware
US10885698B2