Training an object detection model using transfer learning
Patent Information
- Application Number
- CN202210854819.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2021-09-16
- Filing Date
- 2022-07-18
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2042-07-18
Smart Images

Figure CN115830486B_ABST
Abstract
Description
Technical Field
[0001] At least one embodiment relates to processing resources for performing and facilitating operations using an object detection model trained with transfer learning. For example, at least one embodiment relates to a processor or computing system for providing and enabling one or more computing systems to use a transfer learning-trained object detection model to detect objects of a target category depicted in one or more images, according to the various new techniques described herein. Background Technology
[0002] Machine learning is frequently applied to image processing, such as recognizing objects depicted in images. Object recognition can be used in medical imaging, scientific research, autonomous driving systems, robotic automation, security applications, law enforcement practices, and many other environments. Machine learning involves training a computational system using training images and other training data to identify patterns in images that may contribute to object detection. Training can be supervised or unsupervised. Machine learning models can use various computational algorithms, such as decision tree algorithms (or other rule-based algorithms), artificial neural networks, etc. During the inference phase, new images are fed into the trained machine learning model, and the patterns and features identified during training can be used to identify various target objects of interest (e.g., vehicles in a road image). Attached Figure Description
[0003] Various embodiments according to this disclosure will be described with reference to the accompanying drawings, in which:
[0004] Figure 1 It is a block diagram of an example system architecture according to at least one embodiment;
[0005] Figure 2 It is a block diagram of an example training data generator and an example training engine according to at least one embodiment;
[0006] Figure 3 It is a block diagram of an example object detection engine according to at least one embodiment;
[0007] Figure 4A An example trained object detection model according to at least one embodiment is described;
[0008] Figure 4B An example trained object detection model updated to remove the mask head according to at least one embodiment is described;
[0009] Figure 5A A flowchart is shown illustrating an example method for training a machine learning model to detect objects of a target category according to at least one embodiment;
[0010] Figure 5BA flowchart illustrating an example method of using a machine learning model trained to detect objects of a target category, according to at least one embodiment;
[0011] Figure 6 A flowchart illustrating an example method for training a machine learning model and updating the trained machine learning model to remove the mask head, according to at least one embodiment, is shown.
[0012] Figure 7A The inference and / or training logic according to at least one embodiment is illustrated;
[0013] Figure 7B The inference and / or training logic according to at least one embodiment is illustrated;
[0014] Figure 8 An example data center system according to at least one embodiment is shown;
[0015] Figure 9 A computer system according to at least one embodiment is shown;
[0016] Figure 10 A computer system according to at least one embodiment is shown;
[0017] Figure 11 At least a portion of a graphics processor according to one or more embodiments is shown;
[0018] Figure 12 At least a portion of a graphics processor according to one or more embodiments is shown;
[0019] Figure 13 This is an example data flow diagram of an advanced computing pipeline according to at least one embodiment;
[0020] Figure 14 This is a system diagram of an example system for training, adapting, instantiating, and deploying machine learning models in an advanced computing pipeline, according to at least one embodiment; and
[0021] Figure 15A and Figure 15B A data flow diagram of a process for training a machine learning model according to at least one embodiment is shown, as well as a client-server architecture for enhancing annotation tools using a pre-trained annotation model;
[0022] Figure 16A An example of an autonomous vehicle according to at least one embodiment is shown;
[0023] Figure 16B A method for illustrating at least one embodiment is shown. Figure 16A Examples of camera positions and field of view for autonomous vehicles;
[0024] Figure 16C The illustration shows an embodiment according to at least one of the embodiments. Figure 16A Example system architecture for autonomous vehicles; and
[0025] Figure 16D An example of a method for connecting a cloud-based server to a cloud-based server is shown according to at least one embodiment. Figure 16A A system for communication between autonomous vehicles. Detailed Implementation
[0026] Accurately detecting and classifying objects contained in images depicting various environments is a challenging task. Advances have been made in machine learning models trained to detect objects contained in a given input image. However, the accuracy of object detection and classification provided by machine learning models depends on the data used to train the model. In one example, an intelligent surveillance system can use a machine learning model to detect objects in a target category (e.g., the "human" or "personal" category) in images captured by cameras (e.g., surveillance cameras, autonomous vehicles, etc.). In addition to detecting and classifying objects depicted in a given input image, the machine learning model can also be trained to determine one or more features associated with the detected objects. Following the previous example, the machine learning model used by the intelligent surveillance system to detect objects of a target category can also be trained to predict the location of detected objects in a given input image (e.g., relative to other objects depicted in the given input image).
[0027] To train a model to detect objects of a target category with high accuracy (e.g., 95% or higher), training data can be generated based on a large number of images (e.g., thousands, or in some cases millions) (referred to herein as training images). In some systems, the data used to train the model (referred to herein as training data) may include indications of regions in each training image that include objects (e.g., bounding boxes), indications of whether the objects within those regions correspond to the target category, and additional data indicating location (e.g., pose, layout, or orientation) or shape (e.g., mask data). Acquiring a large number of images suitable for training a model to detect and classify objects of the target category can take a significant amount of time (e.g., months or, in some cases, years). Furthermore, accurately determining the labeled data for each image (e.g., regions in each image that include objects, the categories associated with those objects, and the mask data associated with those objects) can take even longer. In some systems, the labeled data for the training images is determined and provided solely by highly trusted entities. For example, in some systems, humans are relied upon to provide indications of regions, categories, and additional data on the objects depicted in each training image. However, in such a system, acquiring labeled data for each training image can be too expensive and time-consuming, as the highly trusted entity providing the labeled data must review thousands or even millions of images and determine and assign labeled data to each one.
[0028] In some cases, labeled data can be generated based on a small number of training images (e.g., dozens or hundreds) and / or based on decisions made by non-highly trusted entities. In such systems, a machine learning model can initially be trained to detect and classify objects contained in a given input image with low accuracy (e.g., less than 95%). During deployment, the model can be retrained based on feedback provided by data determined by one or more of the model's outputs (e.g., regions in a given input image containing detected objects and / or determined categories for detected objects in a given input image). Ultimately, the model can be retrained when deployed to detect and classify objects with high accuracy. However, retraining the model to detect and classify objects with high accuracy can require more time and computational resources, and in some cases, this is because training the model consumes time and computational resources, and retraining the model during the inference phase consumes additional time and resources.
[0029] Embodiments of this disclosure address the aforementioned and other shortcomings by providing transfer learning techniques to train an object detection model to detect objects associated with a target category in a given input image. A first machine learning model (also called a teacher model) can be trained (e.g., via a training data generator and / or training engine) to detect one or more objects depicted in a given input image. In some embodiments, objects depicted in a given input image may correspond to at least one of a plurality of (e.g., dozens, hundreds, etc.) different categories. The teacher model can be trained using first training data, which may include training inputs and a target output, the training inputs comprising one or more images, and the target output comprising labeled data such as data associated with each object depicted in each of the image sets. In some embodiments, the data associated with each object may include indications of regions comprising the object in one or more images, indications of the category (i.e., multiple different categories) associated with the object, and / or mask data associated with the object. Mask data refers to data (e.g., a two-dimensional (2D) bit array) indicating whether one or more pixels (or groups of pixels) of an image correspond to an object. In some embodiments, the images and data associated with objects depicted in the images may be obtained from a publicly available repository or database, comprising a large number of different images and object data that can be used to train a machine learning model for object detection. A first training data may be used to train a teacher model to detect one or more objects depicted in a given set of input images, and for each detected object, predict at least mask data associated with the corresponding detected object. In some additional embodiments, the teacher model may be trained to further predict regions in images of a given set of input images, including the depicted objects (e.g., bounding boxes) and / or the categories corresponding to the detected objects (i.e., among multiple distinct categories).
[0030] Once the teacher model has been trained using the first training data, the trained teacher model can be used to generate second training data to train a second machine learning model (referred to as the student model) to detect objects of a target category depicted in a given input image. An image set can be provided as input to the teacher model. In some embodiments, each image in the image set can be selected (e.g., from a repository or database of a particular domain or organization) to generate the second training data for training the student model. Object data associated with objects detected in each of the image sets provided as input to the teacher model can be determined based on one or more outputs acquired by the teacher model. In some embodiments, the object data may include mask data associated with each detected object. In some other embodiments, for each detected object, the object data may also include data indicating an image region comprising the detected object and / or an indication of the category associated with the detected object.
[0031] Second training data, including one or more outputs of a teacher model, can be used to train a student model associated with a target object category. Specifically, in some embodiments, the second training data may include training input and target output, the training input including a set of images, and the target output including mask data associated with each object detected in the set of images. In some embodiments, a training data generator may use the acquired output of the teacher model to acquire the mask data associated with each object. The target output of the second training data may also include an indication of whether the category associated with each object detected in the set of images corresponds to a target category. For example, as described above, one or more outputs of the teacher model may include an indication of the category associated with objects detected in a given input image contained in the second training data. In some embodiments, the training data generator may determine whether the category associated with each detected object corresponds to a target category based on the acquired output of the teacher model. The target output of the second training data may also include ground truth data associated with each object detected in the set of images. The ground truth data may indicate regions (e.g., bounding boxes) of the image that include the corresponding detected object. In some embodiments, the training data generator may obtain ground truth data from a database that includes indications of one or more bounding boxes associated with images in a set of images (e.g., to replace bounding box data provided by a teacher model for higher accuracy). In some embodiments, the database may be a domain-specific or organization-specific database comprising the set of images. Each of the bounding boxes associated with an image may be provided by an accredited bounding box authority or a user of a computing platform.
[0032] The second training data can be used to train the student model to predict bounding boxes and mask data associated with objects detected in a given input image. The student model can also be trained to predict whether the category associated with an object detected in a given input image corresponds to the target category. As mentioned above, the teacher model can be trained to predict multiple categories of objects detected in a given input image. Because the student model is trained to predict objects of a specific category (i.e., the target category) rather than multiple object categories, it can provide more accurate predictions than the teacher model.
[0033] In some cases, the trained student model can be a multi-head machine learning model. For example, a trained student model may include a first head for predicting bounding boxes associated with objects detected in a given image, a second head for predicting the category associated with the detected objects, and a third head for predicting mask data associated with objects detected in a given image. In some cases, the object detection engine (e.g., of a computing device, cloud computing platform, etc.) can identify the mask data (called mask heads) in the student model corresponding to the predictions associated with the detected objects and can update the student model to remove the identified heads. After removing the mask heads from the student model, the updated student model can be used to predict the bounding boxes and categories associated with objects detected in a given input image. By including the mask heads in the student model at the beginning, the updated student model's object detection and classification predictions can be more accurate than object detection models trained on training data that do not include mask data associated with objects depicted in the provided training images. Furthermore, removing the mask heads from the student model can significantly improve (e.g., 10%–20%) the inference speed associated with the student model and significantly reduce the model size associated with the student model. Therefore, in some cases, updated student models can be transmitted over a network to edge devices and / or one or more endpoint devices (e.g., smart surveillance cameras, autonomous vehicles) for object detection.
[0034] Various aspects and embodiments of this disclosure provide a technique for training object detection models using transfer learning. Training data, using a large number of images depicting various categories of objects (i.e., from publicly available repositories or databases), can be used to train a teacher model to make predictions with sufficient accuracy. Image data can be obtained from domain-specific or organization-specific repositories or databases and used as input to the teacher model to obtain predictions, which can be used to train a student model to improve prediction accuracy for specific, targeted, or particular image categories. Thus, the student model can be trained to make predictions with high (or higher) accuracy (e.g., 95% or higher) without requiring labeled data for training from experts or other accredited authorities. Furthermore, embodiments of this disclosure provide the ability to provide object detection models for use on edge devices and / or endpoint devices (e.g., smart surveillance cameras, autonomous vehicles, etc.), wherein the object detection model is trained to detect objects of a target category with high accuracy and also meets size constraints and inference speed conditions associated with edge devices and / or endpoint devices.
[0035] System Architecture
[0036] Figure 1 This is a block diagram of an example system architecture 100 according to at least one embodiment. System architecture 100 (also referred to herein as the “system”) includes computing device 102, data stores 112A-112N (collectively referred to as data store 112), and server machines 130, 140, and / or 150. In various implementations, network 110 may include a public network (e.g., the Internet), a private network (e.g., a local area network (LAN) or a wide area network (WAN)), a wired network (e.g., Ethernet), a wireless network (e.g., an 802.11 network or a Wi-Fi network), a cellular network (e.g., a Long Term Evolution (LTE) network), routers, hubs, switches, server computers, and / or combinations thereof.
[0037] Computing device 102 may be a desktop computer, laptop computer, smartphone, tablet computer, server, or any suitable computing device capable of performing the techniques described herein. In some embodiments, computing device 102 may be a computing device of a cloud computing platform. For example, computing device 102 may be a server machine of a cloud computing platform or a component of a server machine of a cloud computing platform. In such embodiments, computing device 102 may be coupled to one or more edge devices (not shown) via network 110. An edge device refers to a computing device capable of communicating between computing devices at the boundary of two networks. For example, an edge device may be connected to computing device 102, data storage 112A-112N, server machine 130, server machine 140, and / or server machine 150 via network 110, and may be connected to one or more endpoint devices (not shown) via another network. In such an example, the edge device may enable communication between computing device 102, data storage 112A-112N, server machine 130, server machine 140, and / or server machine 150 and one or more client devices. In other or similar embodiments, computing device 102 may be an edge device or a component of an edge device. For example, computing device 102 can facilitate communication between data storage 112A-112N, server machine 130, server machine 140 and / or server machine 150 connected to computing device 102 via a network and one or more client devices connected to computing device 102 via another network.
[0038] In other or similar embodiments, computing device 102 may be an endpoint device or a component of an endpoint device. For example, computing device 102 may be a device or a component of a device, such as, but not limited to: televisions, smartphones, cellular phones, personal digital assistants (PDAs), portable media players, netbooks, laptops, e-book readers, tablets, desktop computers, set-top boxes, game consoles, autonomous vehicles, surveillance equipment, etc. In such embodiments, computing device 102 may be connected via network 110 to data storage 112A-112N, server machine 130, server machine 140, and / or server machine 150. In other or similar embodiments, computing device 102 may be connected via network to edge devices (not shown) of system 100, and edge devices of system 100 may be connected via network 110 to data storage 112A-N, server machine 130, server machine 140, and / or server machine 150.
[0039] Computing device 102 may include memory 104. Memory 104 may include one or more volatile and / or non-volatile storage devices configured to store data. In some embodiments, computing device 102 may include object detection engine 151. Object detection engine 151 may be configured to detect one or more objects depicted in an image (e.g., image 106), and in some embodiments, acquire data associated with one or more detected objects (e.g., object data 108). For example, object detection engine 151 may be configured to provide image 106 as input to a trained object detection model (e.g., model 160), and determine object data 108 associated with image 106 based on one or more outputs of the trained object detection model. It should be noted that although embodiments of this disclosure are discussed in accordance with object detection models, embodiments can also be broadly applied to any type of machine learning model. Further details regarding object detection engine 151 and object detection models are provided herein.
[0040] As described above, in some embodiments, computing device 102 may be an endpoint device or a component thereof. In such embodiments, computing device 102 may include audiovisual components capable of generating audio and / or visual data. In some embodiments, the audiovisual components may include image capture devices (e.g., a camera) to capture and generate image 106 and generate image and / or video data associated with the generated image 106. In other or similar embodiments, computing device 102 may be an edge device or a component thereof, as described above. In such embodiments, computing device 102 may receive image 106 from an endpoint device including the audiovisual components (i.e., via network 110 or another network). Also as described above, in some embodiments, computing device 102 may be a server machine (e.g., for a cloud computing platform) or a component thereof. In such embodiments, computing device 102 may receive image 106 from an endpoint device including the audiovisual components and / or an edge device connected to the endpoint device (i.e., via network 110 or another network).
[0041] In some implementations, data storage 112A-112N is a persistent storage device capable of storing content items (e.g., images) and data associated with the stored content items (e.g., object data), as well as data structures for tagging, organizing, and indexing the content items and / or object data. Data storage 112 may be hosted by one or more storage devices (e.g., main memory, magnetic or optical storage-based disks, tape or hard disk drives, NAS, SAN, etc.). In some implementations, data storage 112 may be a network-connected file server, while in other embodiments, data storage 112 may be some other type of persistent storage, such as an object-oriented database, a relational database, etc., which may be hosted by computing device 102 or by one or more different machines coupled to computing device 102 via network 110.
[0042] like Figure 1As shown, in some embodiments, system 100 may include multiple data stores 112. In some embodiments, a first data store (e.g., data store 112A) may be configured to store data accessible only by computing device 102, server machine 130, server machine 140, and / or server machine 150. For example, data store 112A may be or include a domain-specific or organization-specific repository or database. In some embodiments, computing device 102, server machine 130, server machine 140, and / or server machine 150 may only be able to access data store 112A via network 110, which may be a private network. In other or similar embodiments, the data stored at data store 112A may be encrypted and accessible by computing device 102, server machine 130, server machine 140, and / or server machine 150 via an encryption mechanism (e.g., a private encryption key, etc.). In additional or alternative embodiments, a second data store (e.g., data store 112B) may be configured to store data accessible by any device accessible by data store 112B via any network. For example, data storage 112B may be or include a publicly accessible repository or database. In some embodiments, data storage 112B may be a publicly accessible data storage that can be accessed by any device via a public network. In additional or alternative embodiments, system 100 may include data storage 112 configured to store first data (e.g., via a private network 110, via encryption mechanisms, etc.) accessible only by computing device 102, server machine 130, server machine 140, and / or server machine 150, and second data accessible by devices connected to the data storage via another network (e.g., a public network). In yet another additional or alternative embodiment, system 100 may include only a single data storage 112 configured to store data (e.g., via a private network 110, via encryption mechanisms, etc.) accessible only by computing device 102, server machine 130, server machine 140, and / or server machine 150. In such embodiments, data storage 112 may store data retrieved from a publicly accessible data storage (e.g., by computing device 102, training data generator 131, training engine 141, etc.).
[0043] Server machine 130 may include a training data generator 131 capable of generating training data (e.g., a set of training inputs and a set of target outputs) to train ML models 160A-160N. Training data may be images stored in a data store that is or includes a repository or database of a specific domain or organization (e.g., data store 112A, etc.) or a dedicated portion of data store 112, and / or images stored in a data store that is or includes a publicly accessible repository or database (e.g., data store 112B, etc.) or a publicly accessible portion of data store 112. For example, training data generator 131 may generate training data for a teacher machine learning model (e.g., a teacher object detection model) based on images stored in data store 112B, a publicly accessible portion of data store 112, or retrieved from a publicly accessible data store (not shown). In another example, training data generator 131 may generate training data for a student machine learning model (e.g., a student object detection model) based on images stored in data store 112A, a dedicated portion of data store 112, or from a single dedicated data store 112, and based on one or more outputs of the teacher object detection model. Further details regarding the generation of training data for the teacher object detection model and the student object detection model will be provided in [the following text is incomplete and likely refers to a separate topic]. Figure 2 supply.
[0044] Server machine 140 may include training engine 141. Training engine 141 can use training data from training data generator 131 to train machine learning models 160A-160N. Machine learning models 160A-160N can refer to model artifacts created by training engine 141 using training data that includes training inputs and corresponding target outputs (correct responses to the corresponding training inputs). Training engine 141 can find patterns in the training data that map training inputs to target outputs (responses to be predicted) and provide machine learning models 160A-160N that capture these patterns. Machine learning models 160A-160N can consist of, for example, single-level linear or non-linear operations (e.g., support vector machines (SVM) or can be deep networks, i.e., machine learning models composed of multiple non-linear operations). An example of a deep network is a neural network with one or more hidden layers, and such a machine learning model can be trained by adjusting the weights of the neural network, for example, according to a backpropagation learning algorithm. For convenience, the remainder of this disclosure will refer to this implementation as a neural network, although some implementations may use an SVM or other type of learning machine instead of a neural network, or as a supplement to a neural network. In some embodiments, training data may be obtained from a training data generator 131 hosted on server machine 130. For example, training engine 141 may obtain first training data for training a teacher object detection model and second training data for training a student object detection model from training set generator 131. Further details regarding the training of the object detection model (e.g., model 160A-N) are provided below. Figure 2 supply.
[0045] Server 150 may include object detection engine 151, which provides one or more images as input to trained machine learning models 160A-160N to obtain one or more outputs. In some embodiments, one or more images are stored in data storage 112 or a proprietary portion of data storage 112, as described above. For example, trained machine learning model 160A may be a trained teacher object detection model. In such an example, object detection engine 151 may provide one or more images as input to trained machine learning model 160A to obtain one or more first outputs. According to embodiments provided herein, training data generator 131 may use one or more first outputs of machine learning model 160A to generate training data to train student object detection models. In another example, trained machine learning model 160B may be a trained student object detection model. In such an example, object detection engine 151 may provide one or more images 106 acquired by computing device 102 as input to trained machine learning model 160B to obtain one or more second outputs. Object detection engine 151 can use one or more second outputs to detect objects depicted in one or more images 106 and determine object data 108 associated with one or more detected objects. Further details regarding object detection engine 151 are available for further information. Figure 3 supply.
[0046] In some implementations, computing device 102, data storage 112, and / or server machines 130-150 may be one or more computing devices (e.g., rack servers, router computers, server computers, personal computers, mainframe computers, laptop computers, tablet computers, desktop computers, etc.), data storage (e.g., hard disks, memory, databases), networks, software components, and / or hardware components that can be used for image-based (e.g., image 106) object detection. It should be noted that in some other implementations, the functionality of computing device 102, server machines 130, 140, and / or 150 may be provided by a fewer number of machines. For example, in some implementations, server machines 130 and 140 may be integrated into a single machine, while in other implementations, server machines 130, 140, and 150 may be integrated into multiple machines. Furthermore, in some implementations, one or more of server machines 130, 140, and 150 may be integrated into computing device 102. Generally, the functions described in the implementation as being performed by computing device 102 and / or server machines 130, 140, 150 may also be performed on one or more edge devices (not shown) and / or client devices (not shown), if appropriate. Furthermore, functions belonging to a particular component may be performed by different or more components operating together. Computing device 102 and / or server machines 130, 140, 150 may also be accessed as services provided to other systems or devices through appropriate application programming interfaces.
[0047] Figure 2 This is a block diagram of an example training data generator 131 and an example training engine 141 according to at least one embodiment. The training data generator 131 may include a teacher model training data generator 210 and a student model training data generator 220. The training engine 141 may include a teacher model training module 230 and a student model training module 232. As previously described, the training data generator 131 may reside on a server machine, for example... Figure 1 The server machine 130 is either part of or separate from the computing device 102. The training engine 141 may reside on the server machine 130 or be part of or separate from the computing device 102 on another server machine, such as server machine 140.
[0048] In some embodiments, the teacher model training data generator 210 can be configured to generate training data for training a teacher object detection model (e.g., model 160A), and the student model training data generator 220 can be configured to generate training data for training a student object detection model (e.g., model 160B). Figure 2As shown, the training data generator 131 can be connected to the data storage 250. The data storage 250 can be configured to store data used by the teacher model training data generator 210 to generate training data for training the teacher object detection model. For example, the data storage 250 can be configured to store one or more training images 252, and for each training image 252, training image region of interest (ROI) data 254, training image mask data 256, and / or training image object data 258 associated with the training image 252. In some embodiments, each training image 252 can depict an object associated with a specific category in a set of multiple different object categories. In an illustrative example, the training image 252 can depict a first object corresponding to the human category and a second object corresponding to the animal category (e.g., the dog category).
[0049] The training image ROI data 254 can indicate each region of the corresponding training image 252 that depicts a corresponding object. In some embodiments, the training image ROI data 254 can correspond to a bounding box or another boundary shape (e.g., sphere, ellipse, cylinder, etc.) indicating the region of the training image 252 that depicts a corresponding object. According to the preceding example, the training image ROI data 254 associated with the example training image 252 can include a first bounding box indicating a first region of the training image 252 that depicts a first object and a second bounding box indicating a second region of the training image 252 that depicts a second object.
[0050] Training image mask data 256 may be data (e.g., a two-dimensional (2D) bit array) indicating whether one or more pixels (or groups of pixels) in a corresponding training image 252 correspond to a specific object. Based on the preceding example, training image mask data 256 associated with example training image 252 may include an indication of a first group of pixels corresponding to a first object and an indication of a second group of pixels corresponding to a second object.
[0051] Training image object data 258 may refer to data indicating one or more features associated with each object depicted in a respective training image 252. In some embodiments, training image object data 258 may include data indicating a category (i.e., one of multiple categories) associated with the depicted object. For example, training image object data 258 associated with example training image 252 may include data indicating that a first object depicted in training image 252 is associated with a first category (e.g., the human category) and a second object depicted in training image 252 is associated with a second category (e.g., the animal category). In additional or alternative embodiments, training image object data 258 may include data indicating other features associated with each depicted object, such as the object's position (e.g., orientation, etc.) or shape.
[0052] In some embodiments, training images 252 may be included in an image set that can be used to train an object detection model. For example, training images 252 may be included in a publicly accessible image set that can be retrieved from a publicly accessible data store (e.g., data store 112B) or a publicly accessible portion of a data store (e.g., data store 112) and used to train the object detection model. In some embodiments, each of the images in the image set may be associated with image data, which may also be included in a publicly accessible data store or a publicly accessible portion of a data store. In some embodiments, each of the images in the image set and the image data associated with each of the images in the image set may be provided by one or more users of the object detection platform. For example, a user of the object detection platform may provide (i.e., via a corresponding client device associated with the user) images depicting one or more objects. The user may also provide (i.e., via a graphical user interface of the corresponding client device) indications of ROI data, mask data, and object data associated with each object depicted in the provided images. In another example, a first user of the object detection platform may provide (i.e., via a first client device associated with the first user) images depicting one or more objects, and a second user may provide (i.e., via a graphical user interface of a second client device associated with the second user) indications of ROI data, mask data, and object data associated with each object depicted in the image provided by the first user. In some embodiments, according to the embodiments provided herein, the data retrieval module 212 of the teacher model training data generator 210 may retrieve training images 252 from publicly accessible data storage or publicly accessible portions of data storage for training the teacher object detection model.
[0053] In some embodiments, data storage 250 may correspond to data storage for... Figure 1The described publicly accessible data storage is, for example, data storage 112B. In other or similar embodiments, data storage 250 may correspond to a publicly accessible portion of data storage 112A. Data retrieval module 212 may retrieve training images 252, training image ROI data 254, training image mask data 256, and / or training image object data 258 from data storage 250 (i.e., from data storage 112A or data storage 112B). In other or similar embodiments, data storage 250 may correspond to data storage 112, which may only be accessed via a private network and / or via encryption mechanisms (e.g., data storage 112A). In such an embodiment, the data retrieval module 212 can retrieve training images 252, training image ROI data 254, training image mask data 256, and / or training image object data 258 from publicly accessible data storage (not shown) and from a repository in such an embodiment. The data retrieval module 212 can also retrieve training images 252, training image ROI data 254, training image mask data 256, and / or training image object data 258 from data storage 250.
[0054] The training data generator module 214 can generate training data for training a teacher object detection model in response to the data retrieval module 212 retrieving training images 252, training image ROI data 254, training image mask data 256, and / or training image object data 258. In some embodiments, the training data may include a training input set and a target output set. As described above, the training input set may include one or more training images 252 retrieved by the data retrieval module 212. In some embodiments, the training data generator module 214 may apply one or more image transformations to one or more training images 252 retrieved by the data retrieval module 212. For example, the training images 252 retrieved by the data retrieval module 212 may be associated with a specific amount of image noise. The training data generator module 214 may apply one or more image transformations to the retrieved training images 252 to generate modified training images. The modified training images may include a different amount of image noise (e.g., less image noise) compared to the retrieved training images 252. According to the above embodiment, the training data generator module 214 may include modified training images in the training image set 252. The target output set may include training image ROI data 254, training image mask data 256, and / or training image object data 258. In response to generating the training input set and the target output set, the training data generator module 214 may generate a mapping between the training input set and the target output set to generate teacher model training data 272.
[0055] In some embodiments, the teacher model training data generator 210 may store the teacher model training data 272 at the data storage 270. The data storage 270 may be, or include, a domain-specific or organization-specific repository or database (e.g., data storage 112A) or a dedicated portion of the data storage 112 accessible by a computing device via a dedicated network and / or through encryption mechanisms. According to... Figure 1 In the described embodiments, training data generator 131 and / or training engine 141 can access data storage 270. In additional or alternative embodiments, teacher model training data generator 210 can provide the generated mappings to teacher model training module 230 of training engine 141 to train the teacher object detection model.
[0056] In response to acquiring training data 272 (i.e., from the teacher model training data generator 210 or from the data storage 270), the teacher model training engine 230 can use the training data 272 to train a teacher object detection model. In some embodiments, the teacher object detection model may be trained to detect one or more objects of multiple categories depicted in a given input image, and predict mask data and / or ROI data associated with each detected object. In some embodiments, the teacher object detection model may also be trained to predict object data (e.g., object category, other feature data associated with the object, etc.) for each object detected in a given input image. In other or similar embodiments, the teacher object detection model may be trained to detect one or more objects of a single category depicted in a given input image, and predict mask data, ROI data, object category data, and / or other object feature data associated with each detected object. In some embodiments, in response to training the teacher object detection model, the training engine 141 may store the trained teacher object detection model in the data storage 270 as a teacher model 274. According to the embodiments provided herein, the student model training data generator 220 can use the trained teacher model 274 to generate training data to train the student object detection model.
[0057] Data storage 270 may store one or more training images 276, which will be used to train a student object detection model to detect objects associated with a target category. For example, the target category may correspond to the human category. In such an example, one or more training images 276 at data storage 270 may depict one or more objects associated with the human category. In some embodiments, each object depicted in a corresponding training image 276 may be associated with different features (e.g., people in different environments, people in different locations, etc.). Therefore, one or more training images 276 can be used to train a student object detection model to detect objects associated with the human category and also associated with different features. In another example, one or more training images 276 in data storage 270 may depict one or more objects that are not associated with the human category but correspond to one or more features similar to those of objects associated with the human category (e.g., objects depicted in training image 276 are associated with similar locations or environments to those associated with the human category). Therefore, one or more training images 276 can be used to train a student object detection model to detect objects that correspond to one or more features associated with the human category but are not associated with the human category.
[0058] The teacher inference module 222 can retrieve one or more training images 276 from the data storage 270 and provide one or more training images 276 as input to the trained teacher model 274. The teacher output module 224 can acquire one or more outputs of the trained teacher model 274 and can determine object data associated with each input training image 276 from one or more acquired outputs. In some embodiments, the determined object data for a corresponding training image 274 may include output image ROI data 278, output image mask data 280, and / or output image feature data 282. The output image ROI data 278 associated with a corresponding training image 276 may be a region of the corresponding training image 276 depicting a specific object. The output image mask data 280 may indicate mask data associated with a specific object. The output image feature data 282 may indicate one or more features associated with a specific object, such as the object's category, object location, object shape, etc. In some embodiments, the teacher output module 224 may store the output image ROI data 278, the output image mask data 280, and / or the output image feature data 282 in the data storage 270.
[0059] The training data generator module 226 can generate training data for training a student object detection model to detect one or more objects associated with a target category for a given input image. In some embodiments, the training data may include a training input set and a target output set. As described above, the training input set may include one or more training images 276 provided as input to the trained teacher model 274. The target output set may include at least output image mask data 280 determined from one or more outputs of the trained teacher model 274. In some embodiments, the target output set may include output image mask data 280 and output image feature data 282.
[0060] In some embodiments, the training data generator module 226 may generate updated image feature data based on output image feature data 282 determined from one or more outputs of the trained teacher model 274. For example, a training image 276 provided as input to the training teacher model 274 may depict a first object associated with a first category (e.g., human category, etc.) and a second object associated with a second category (e.g., animal category, etc.). The teacher output module 224 may obtain from the trained teacher model 274 one or more outputs indicating the first and second objects detected in a given input image, and output image feature data 282 indicating that the first object is associated with a first category and the second object is associated with a second category. The training data generator module 226 may determine whether the first category and / or the second category corresponds to a target category, and may generate updated image object data based on this determination. For example, if the target category is the human category, the training data generator module 226 may determine that the first category corresponds to the target category and generate updated image feature data to indicate that the first object corresponds to the target category. The training data generator module 226 can also determine that the second category does not correspond to the target category and can generate updated image feature data to indicate that the second object does not correspond to the target category. In some embodiments, the training data generator module 226 may include updated object data in the target output set instead of the output image feature data 282.
[0061] In some embodiments, the target output set may additionally include image ground truth data 284. Image ground truth data 284 may indicate regions of a corresponding training image 276 that include objects detected by the trained teacher model 274. For example, ground truth data 284 may include indications of one or more bounding boxes associated with the training image 276, which was obtained from an accredited bounding box authority or a user of a computing device or object detection platform. In some embodiments, ground truth data 284 is obtained from an accredited bounding box authority or user before or after providing one or more training images 276 as input to the trained teacher model 274. In an illustrative example, image ground truth data 284 may correspond to output image ROI data 278 unless the bounding boxes of image ground truth data 284 can more accurately identify regions of image 276 depicting a particular object than the bounding boxes of output image ROI data 278. In additional or alternative embodiments, image ground truth data 284 may indicate the category of objects detected by the trained teacher model 274. In some embodiments, as described above, the target output set may include the categories of objects indicated by image ground truth data 284 instead of the categories of objects indicated by output image feature data 282. In some embodiments, according to the previously described embodiments, training data generation module 226 may generate updated image feature data based on the object categories indicated by image ground truth data 284.
[0062] In response to generating a training input set and a target output set for the corresponding training image 276, the training data generator module 226 may generate a mapping between the training input set and the target output set to generate student model training data 286. In some embodiments, the student model training data generator 220 may store the student model training data 286 in the data storage 270. In other or similar embodiments, the student model training data generator 220 may send the student model training data 286 to the training engine 141. In response to acquiring the training data 286 (i.e., from the student model training data generator 220 or from the data storage 270), the student model training module 232 may use the training data 286 to train a student object detection model. The student object detection model may be trained to detect one or more objects of a target category depicted in a given input image for a given input, and predict mask data, ROI data, and / or feature data associated with each detected object. In some embodiments, the training engine 141 may provide the trained student object detection model to the object detection engine, for example... Figure 1 The object detection engine 151 in the system.
[0063] Figure 3This is a block diagram of an example object detection engine 310 according to at least one embodiment. In some embodiments, the object detection engine 310 may be related to a reference... Figure 1 The object detection engine described corresponds to 151. For example... Figure 3 As shown, the object detection engine 310 may include an input image component 312, an object data component 314, a model head component 316, and / or a model update component 318. In some embodiments, the object detection engine 310 may be coupled to a memory 320. In some embodiments, the object detection engine 310 may reside at a computing device 102. In such embodiments, the memory 320 may be associated with a computing device 102. Figure 1 The memory 104 described corresponds to this. In other or similar embodiments, the object detection engine 310 may reside at server 150. In such embodiments, memory 320 may correspond to memory for data storage (e.g., data storage 112), memory 104, or memory in another memory device associated with system 100.
[0064] The input image component 312 can be configured to acquire an image (e.g., image 106) and provide the acquired image as input to a trained object detection model 322 stored in memory 320. In some embodiments, the trained object detection model 322 may correspond to a student object detection model, which is trained by training engine 141 using training data generated by training data generator 131, as per [reference to...]. Figure 1 and Figure 2 As described. In other or similar embodiments, the trained object detection model 322 may correspond to another trained object detection model that was not trained by the training engine 141 using training data generated by the training data generator 131.
[0065] Such as about Figure 1 As described, in some embodiments, computing device 102 may be a computing device of a cloud computing platform. In such embodiments, computing device 102 may be coupled to one or more edge devices (e.g., edge device 330), each edge device being coupled to one or more endpoint devices (e.g., endpoint devices 332A-332N (collectively referred to herein as point devices 332)). Figure 3As shown. In some embodiments, the audiovisual and / or sensor components of endpoint device 332 may generate image 106 as described above and transmit image 106 to edge device 330 (e.g., via a network). Edge device 330 may transmit the received image 106 to computing device 102 (e.g., via network 110). In such embodiments, computing device 102 may send image 106 to input image component 312 (e.g., via network 110 or the bus of computing device 102). In other or similar embodiments, as also described above, computing device 102 may be edge device 330 or may be a component of edge device 330. In such embodiments, edge device 330 may receive image 106 from endpoint device 332 (e.g., via a network) and may transmit image 106 to input image component 312 (e.g., via network 110 or the bus of computing device 102). In other or similar embodiments, as previously described, computing device 102 may be one or more terminal devices 332A-332N, or may be a component of one or more terminal devices 332A-332N. In such embodiments, one or more endpoint devices 332A-332N may generate image 106 and send image 106 to input image component 312 (e.g., via a network, network 110, or bus of computing device 102).
[0066] In some embodiments, in response to receiving image 106, input image component 312 provides image 106 as input to trained object detection model 322. In other or similar embodiments, as described above, input image component 312 may apply one or more image transformations (e.g., to reduce the amount of noise included in image 106) to generate a modified image and provide the modified image as input to trained object detection model 322. Object data component 314 may acquire one or more outputs of trained object detection model 322 and may determine object data 108 based on the acquired one or more outputs. In some embodiments, object data 108 determined based on the acquired one or more outputs may correspond to one or more objects detected in a given input image 106 (or a modified input image). For example, object data 108 may include indications of regions of image 106 that include detected objects (e.g., bounding boxes). In some additional or alternative embodiments, object data 108 may also include mask data associated with the detected objects. In some additional or alternative embodiments, object data 108 may also include data indicating the category and / or one or more characteristics associated with the detected object.
[0067] In some embodiments, one or more outputs of the trained object detection model 322 may include indications of multiple regions of image 106 and indications of the confidence level of each region including a detected object. Object data component 314 can determine a specific region in image 106 that includes a detected object by determining that the confidence level associated with a specific region of image 106 meets a confidence level criterion (e.g., the confidence level exceeds a threshold, etc.). In response to determining that a specific region of image 106 meets the confidence level criterion, object data component 314 may include an indication in memory 320 of the specific region in image 106 having object data. In additional or alternative embodiments, one or more outputs of the trained object detection model 322 may include multiple mask data sets and indications of the confidence level associated with each mask data set for a detected object. In response to determining that a particular mask data set meets the confidence level criterion, object data component 314 may include in memory 320 an indication of specific mask data having object data 108. In yet another additional or alternative embodiment, one or more output features of the trained object detection model 322 may include multiple categories and / or features and an indication of the confidence level of each category and / or feature corresponding to the detected object. In response to determining that a particular category and / or feature meets the confidence level criteria, the object data component 314 may include an indication of the particular category and / or feature of the object data 108 in the memory 320.
[0068] In some embodiments, object data component 314 may determine, based on object data 108, whether an object detected in a given input image 106 corresponds to a target category. For example, in some embodiments, object data 108 may indicate a category associated with the detected object, as described above. Object data component 312 may compare the indicated category with a target category to determine whether the detected object corresponds to the target category. In another example, according to the previously described embodiments, object data 108 may include data indicating whether the detected object corresponds to a target category. In such an example, object data component 312 may determine whether the detected object corresponds to a target category based on the included data. In some embodiments, object data component 314 may update object data 108 to include an indication of whether the detected object corresponds to a target category. Object data component 314 may transmit object data 108 to computing device 102. In some embodiments, object data component 312 may additionally or alternatively transmit a notification to computing device 102 indicating whether the detected object corresponds to a target category.
[0069] As described above, in some embodiments, the trained object detection model 322 (e.g., Figure 2The trained object detection model 322 can be trained to predict ROI data, image mask data, and / or image feature data associated with a given input image. In such embodiments, the trained object detection model 322 can be a multi-head model, where each head of the multi-head model is used to predict a specific type of data associated with an object detected in a given input image. For example, the trained object detection model 322 may include a first head corresponding to predicting ROI data associated with the detected object, a second head corresponding to predicting mask data associated with the detected object, and / or a third head corresponding to predicting feature data (e.g., the category of the detected object) associated with the detected object. As described above, in some embodiments, a large number of training images (e.g., training image 276) may be used to train the object detection model. For example, in some systems, hundreds, thousands, or in some cases millions of images may be used to train the object detection model. In view of the above, in some embodiments, a multi-head object detection model 322 trained on a large number of images may consume a large amount of system resources (e.g., storage space, processing resources, etc.). In some embodiments, the object detection engine 310 may remove one or more heads (e.g., mask heads) of the trained multi-head object detection model 322 to reduce the amount of system resources consumed by the object detection model 322 before and / or during inference.
[0070] Figure 4A An example trained multi-head object detection model 322 according to at least one embodiment is described. Figure 4A As shown, model 322 may include at least a ROI head 412 and a mask head 414. It should be noted that, although as... Figure 4A The illustrated model 322 may resemble a neural network, but embodiments of this disclosure can be applied to any type of machine learning model. According to the previously described embodiments, the input image component 312 may provide image 106 as input to model 322. Input image 106 may be provided to the ROI head 412 and mask head 414 of model 322. As described above, model 322 may provide one or more outputs based on a given input image 106. In some embodiments, the provided outputs may include ROI head output 416 and mask head output 418. ROI head output 416 may be provided based on inference performed according to ROI head 412. Mask head output 418 may be provided based on inference performed according to mask head 414.
[0071] Return to reference Figure 3In some embodiments, the model head component 316 and model update component 318 of the object detection engine 310 can remove mask heads from the object detection model 322. For example, the model head component 316 can identify one or more heads (e.g., mask head 414) of the model 322 that correspond to providing a specific output (e.g., predicted mask data) associated with a given input image. In response to the model head component 316 identifying the corresponding head of the model 322, the model update component 318 can update the model 322 to remove one or more identified heads. Figure 4B An updated trained object detection model 324 according to at least one embodiment is depicted, which is updated to remove mask heads 414. Figure 4B As shown, the model update component 318 can remove the mask head 414 of model 322 to generate an updated model 324. Therefore, in some embodiments, model 324 may provide an ROI head output 416 and may not provide a mask head output 418. It should be noted that, although Figure 3 and Figures 4A-4B This includes embodiments involving the removal of mask head 414 from model 322, and embodiments of this disclosure can be applied to the removal of any head from model 322.
[0072] Return to reference Figure 3 In some embodiments, the object detection engine 310 may transmit the updated object detection model 324 to the computing device 102 (e.g., via network 110 or the bus of the computing device 102). As described above, in some embodiments, the computing device 102 may be a cloud computing platform or a component of a cloud computing platform. In such embodiments, the computing device 102 may transmit the updated object detection model 324 to the edge device 330 (e.g., via network 110). In some embodiments, the edge device 330 may use the updated object detection model 324 to perform object detection based on the image 106 generated by the endpoint devices 332A-332N. In other or similar embodiments, the edge device 330 may transmit the updated object detection model 324 to the endpoint devices 332A-332N. Also as described above, the computing device 102 may be the edge device 330 or a component of the edge device 330. In such embodiments, the edge device 330 may use the updated object detection model 324 to perform object detection and / or may transmit the updated object detection model 324 to the endpoint device 332A. As described above, computing device 102 may be one or more endpoint devices 332A-332N or a component of one or more endpoint devices 332A-332N. In such an embodiment, according to the above embodiment, one or more endpoint devices 332A-332N may use an updated object detection model 324 to perform object detection.
[0073] Figures 5A-5B and Figure 6 The following are flowcharts of example methods 500, 550, and 600 related to training an object detection model, according to at least some embodiments. In at least one embodiment, methods 500, 550, and / or 600 may be executed by computing device 102, server machine 130, server machine 140, server machine 150, one or more edge devices, one or more endpoint devices, or other computing devices, or combinations of one or more computing devices. Methods 500, 550, and / or 600 may be executed by one or more processing units (e.g., CPU and / or GPU), which may include one or more memory devices (or communicate with one or more memory devices). In at least one embodiment, methods 500, 550, and / or 600 may be executed by multiple processing threads (e.g., CPU threads and / or GPU threads), each thread performing the operation of one or more individual functions, routines, subroutines, or methods. In at least one embodiment, the processing threads implementing methods 500, 550, and / or 600 may be synchronized (e.g., using semaphores, critical sections, and / or other thread synchronization mechanisms). Alternatively, the processing threads implementing methods 500, 550, and / or 600 can be executed asynchronously relative to each other. The various operations of methods 500, 550, and / or 600 can be performed in conjunction with... Figures 5A-5B and Figure 6 The different sequences shown are executed. Some operations of these methods can be executed simultaneously with other operations. In at least one embodiment, Figures 5A-5B and Figure 6 One or more of the operations shown may not always be performed.
[0074] Figure 5A A flowchart illustrating an example method 500 for training a machine learning model to detect objects of a target category according to at least one embodiment is shown. In some embodiments, one or more operations of method 500 may be performed by one or more components or modules of training data generator 131, as described herein. At block 510, a processing unit performing method 500 may identify a first set of images containing multiple objects of multiple categories. In some embodiments, the processing unit may obtain the first set of images from data storage 270, as previously described.
[0075] At block 512, the processing unit performing method 500 may provide a first set of images as input to a first machine learning model. The first machine learning model may be a trained teacher object detection model trained to detect one or more objects of multiple categories depicted in a given input image. The trained teacher object detection model may also be trained to predict mask data for each of the one or more detected objects, and, in some embodiments, ROI data associated with the respective detected object. In additional or alternative embodiments, the trained teacher object detection model may be trained to predict a specific category of multiple categories associated with each detected object. At block 514, the processing unit performing method 500 may determine object data associated with the first set of images from one or more outputs of the first machine learning model. The object data for each corresponding image in the first set of images may include mask data associated with each object detected in the corresponding image.
[0076] At block 516, the processing unit performing method 500 may train a second machine learning model to detect objects of a target category in the second image set using a first image set and a portion of object data determined from one or more outputs of the first machine learning model. The second machine learning model may be a student object detection model. The processing unit may train the student object detection model using training inputs and a target output. The training inputs may include the first image set (i.e., provided as input to the trained teacher object detection model). The target output may include mask data associated with each detected object in the first image set, which is included in one or more outputs of the teacher object detection model. The target output may also include an indication of whether the category associated with each object detected in the first image set corresponds to a target category. In some embodiments, the processing unit may determine whether a specific category associated with each object detected in the first image set corresponds to a target category. The target output may include an indication of whether that specific category corresponds to the target category. In another embodiment, the target output may include an indication of the specific category associated with each object detected in the first image set.
[0077] In some embodiments, the target output may further include ground truth data associated with each object detected in the first image set. As described above, the ground truth data associated with each detected object can indicate a region in the image that includes the corresponding detected object. According to the previously described embodiments, the processing unit may use a database (e.g., in data storage 270) that includes indications of one or more ROIs (e.g., bounding boxes) associated with the image to identify the ground truth data. Each ROI included in the database may be provided by an accredited ROI authority or a user of the platform. In some embodiments, the processing unit may also use the database to identify the category associated with each object detected in the first image set.
[0078] In some embodiments, the trained second machine learning model may be a multi-head machine learning model, as described above. In some embodiments, the processing unit may identify one or more heads in the second machine learning model corresponding to mask data for predicting a given input image, and may update the second machine learning model to remove one or more identified heads. In some embodiments, the processing unit may provide a second set of images as input to the second machine learning model and obtain one or more outputs of the second machine learning model. The processing unit may determine additional object data associated with each of the one or more obtained outputs. In some embodiments, the additional object data may include indications of regions (e.g., bounding boxes) in the corresponding image that include objects detected in the corresponding image, and the category associated with the detected objects.
[0079] Figure 5B A flowchart of an example method 550 using a machine learning model trained to detect objects of a target category, according to at least one embodiment, is shown. In some embodiments, one or more operations of method 550 may be performed by one or more components or modules of object detection engine 151, as described herein. At block 552, the processing unit performing method 550 may provide a current set of images as input to a first machine learning model. In some embodiments, the current set of images may be generated by an audiovisual component (e.g., a camera) located at or coupled to an endpoint device, edge device, or server, as described above.
[0080] A first machine learning model can be trained to detect objects of a target category in a given set of images. In some embodiments, the first machine learning model may correspond to a student object detection model, as described above. In some embodiments, the first machine learning model may be trained according to the previously described embodiments. For example, the first machine learning model may be trained using training inputs comprising a set of training images and a target output for the training inputs. For each corresponding training image in the set of training images, the target output may include ground truth data associated with each object depicted in the corresponding training image. The ground truth data may indicate the region of the corresponding training image that includes the corresponding object. In some embodiments, the ground truth data may be obtained using a database including indications of one or more bounding boxes associated with the set of training images. Each of the one or more bounding boxes may be provided by an accredited bounding box authority and / or a user of the platform.
[0081] The target output may also include mask data associated with each object depicted in the respective training images. The mask data may be obtained based on one or more outputs of a second machine learning model. In some embodiments, the second machine learning model may correspond to the teacher model described herein. For example, the set of training images may be provided as input to the second machine learning model. The second machine learning model may be trained to detect one or more objects of at least one of a plurality of categories depicted in a given input image, and predict at least mask data associated with each of the one or more detected objects. As described above, object data may be determined from one or more outputs of the second machine learning model. The object data for each respective training image may include mask data associated with each object detected in the respective image. According to the previously described embodiments, the target output may also include an indication of whether the category associated with each object depicted in the respective training image corresponds to a target category.
[0082] At block 554, the processing unit performing method 550 may acquire one or more outputs of the first machine learning model. At block 556, the processing unit performing method 550 may determine object data associated with each of the current image sets based on the acquired one or more outputs. In some embodiments, the determined object data for each corresponding image in the current image set may include an indication of a region in the corresponding image that includes an object detected in the corresponding image and an indication of whether the detected object corresponds to a target category. In some embodiments, the object data may also include mask data associated with an object detected in the corresponding image. In some embodiments, object data associated with each of the image sets may be determined by extracting one or more sets of object data from one or more outputs of the first machine learning model. Each of the one or more sets of object data may be associated with a confidence level corresponding to an object detected in the corresponding image. The processing unit may determine whether the confidence level associated with the corresponding set of object data meets a confidence level criterion (e.g., exceeds a confidence level threshold). In response to determining that the confidence level associated with the corresponding set of object data meets the confidence level criterion, the processing unit may determine the set of object data corresponding to the detected object.
[0083] Figure 6 A flowchart illustrating an example method for training a machine learning model and updating the trained machine learning model to remove a mask head, according to at least one embodiment, is shown. In some embodiments, one or more operations of method 600 may be performed by one or more components or modules of training data generator 131 and / or training engine 141, as described herein. At block 610, the processing unit performing method 600 may be identified as generating training data for the machine learning model.
[0084] At block 612, the processing unit performing method 600 may generate training input including an image of the depicted object. At block 614, the processing unit performing method 600 may generate a target output for the training input. The target output may include a bounding box associated with the depicted object, mask data associated with the depicted object, and an indication of a category associated with the depicted object. In some embodiments, the processing unit may generate the target output by providing an image of the depicted object as input to an additional machine learning model trained to detect one or more objects depicted in a given input image and predict at least mask data associated with each of the one or more detected objects. In some embodiments, the additional machine learning model is further trained to predict a category associated with the corresponding detected object. The processing unit may use the additional machine learning model to obtain mask data (and, in some embodiments, an indication of a category) associated with the objects depicted in the training input image.
[0085] In additional or alternative embodiments, the processing unit may generate the target output by acquiring ground truth data associated with the image. As described above, the ground truth data may include bounding boxes associated with the depicted objects, any of which can be obtained from a database storing indications of bounding boxes associated with objects depicted in the image set. The indications of the bounding boxes may be provided by an accredited bounding box authority or a user of the platform.
[0086] At box 616, the processing unit performing method 600 may provide training data to train a machine learning model on (i) a set of training inputs including generated training inputs and (ii) a set of target outputs including generated target outputs. At box 618, the processing unit performing method 600 may identify one or more heads of the trained machine learning model corresponding to the mask data predicted for a given input image. At box 620, the processing unit performing method 600 may update the trained machine learning model to remove one or more identified heads.
[0087] In some embodiments, the processing unit or other processing unit performing method 600 may provide a set of images as input to an updated trained machine learning model and obtain one or more outputs of the updated trained machine learning model. The processing unit may determine object data associated with each of the images from the one or more outputs. The object data may include regions in the respective images that include objects detected in the respective images and indications of the categories associated with the detected objects. In some embodiments, the processing unit or other processing unit performing method 600 may transmit the updated trained machine learning model via a network to at least one of an edge device or an endpoint device.
[0088] Reasoning and training logic
[0089] Figure 7A Inference and / or training logic 715 is shown for performing inference and / or training operations associated with one or more embodiments. The following is in conjunction with... Figure 7A and / or Figure 7B Provide details about reasoning and / or training logic 715.
[0090] In at least one embodiment, the inference and / or training logic 715 may include, but is not limited to, code and / or data memory 701 for storing forward and / or output weights and / or input / output data, and / or other parameters configuring neurons or layers of a neural network trained for and / or used for inference in one or more embodiments. In at least one embodiment, the training logic 715 may include or be coupled to code and / or data memory 701 for storing graph code or other software to control timing and / or sequence, wherein weight and / or other parameter information is loaded to configure logic, including integer and / or floating-point units (collectively, arithmetic logic units (ALUs)). In at least one embodiment, code (such as graph code) loads weight or other parameter information into the processor ALU based on the architecture of the neural network to which the code corresponds. In at least one embodiment, code and / or data memory 701 stores weight parameters and / or input / output data of each layer of a neural network trained or used in one or more embodiments during forward propagation of input / output data and / or weight parameters during training and / or inference using one or more embodiments. In at least one embodiment, any portion of the code and / or data memory 701 may be included in other on-chip or off-chip data memory, including the processor's L1, L2, or L3 cache or system memory.
[0091] In at least one embodiment, any portion of the code and / or data memory 701 may be internal or external to one or more processors or other hardware logic devices or circuits. In at least one embodiment, the code and / or data memory 701 may be a cache memory, dynamic random-addressable memory (“DRAM”), static random-addressable memory (“SRAM”), non-volatile memory (e.g., flash memory), or other memory. In at least one embodiment, the choice of whether the code and / or data memory 701 is internal or external to the processor, for example, or composed of DRAM, SRAM, flash memory, or some other memory type, may depend on the available on-chip or off-chip storage space, the latency requirements of the training and / or inference functions being performed, the batch size of the data used in the inference and / or training of the neural network, or some combination of these factors.
[0092] In at least one embodiment, the inference and / or training logic 715 may include, but is not limited to, code and / or data memory 705 for storing backpropagation and / or output weights and / or input / output data neural networks corresponding to neurons or layers of a neural network trained and / or used for inference in one or more embodiments. In at least one embodiment, during training and / or inference using one or more embodiments, the code and / or data memory 705 stores weight parameters and / or input / output data for each layer of a neural network trained or used in one or more embodiments during backpropagation of input / output data and / or weight parameters. In at least one embodiment, the training logic 715 may include or be coupled to the code and / or data memory 705 for storing graph code or other software to control timing and / or sequence, wherein weight and / or other parameter information is loaded to configure logic including integer and / or floating-point units (collectively, an arithmetic logic unit (ALU)). In at least one embodiment, code (such as graph code) loads weight or other parameter information into the processor ALU based on the architecture of the neural network to which the code corresponds. In at least one embodiment, any portion of the code and / or data memory 705 may be included together with other on-chip or off-chip data storage, including the processor's L1, L2, or L3 cache or system memory. In at least one embodiment, any portion of the code and / or data memory 705 may be internal or external to one or more processors or other hardware logic devices or circuits. In at least one embodiment, the code and / or data memory 705 may be cache memory, DRAM, SRAM, non-volatile memory (e.g., flash memory), or other memory. In at least one embodiment, the choice between the code and / or data storage 705 being internal or external to the processor, for example, whether it consists of DRAM, SRAM, flash memory, or some other type of storage, depends on whether the available storage is on-chip or off-chip, the latency requirements of the training and / or inference functions being performed, the data batch size used in the inference and / or training of the neural network, or some combination of these factors.
[0093] In at least one embodiment, code and / or data memory 701 and code and / or data memory 705 may be separate memory structures. In at least one embodiment, code and / or data memory 701 and code and / or data memory 705 may be the same memory structure. In at least one embodiment, code and / or data memory 701 and code and / or data memory 705 may have partially the same memory structure and partially separate memory structures. In at least one embodiment, any portion of code and / or data memory 701 and code and / or data memory 705 may be included together with other on-chip or off-chip data memory, including the processor's L1, L2, or L3 cache or system memory.
[0094] In at least one embodiment, the inference and / or training logic 715 may include, but is not limited to, one or more arithmetic logic units (“ALUs”) 710 (including integer and / or floating-point units) for performing logical and / or mathematical operations based at least in part on or instructed by training and / or inference code (e.g., graph code), the results of which may produce activations (e.g., output values from layers or neurons within a neural network) stored in activation memory 720, which are functions of input / output and / or weight parameter data stored in code and / or data memory 701 and / or code and / or data memory 705. In at least one embodiment, activation is activated in response to execution instructions or other code, linear algebraic and / or matrix-based mathematical generation performed by ALU 710, and the activation is stored in activation storage 720, wherein weight values stored in code and / or data memory 705 and / or code and / or data storage 701 are used as operands with other values, such as bias values, gradient information, momentum values, or other parameters or hyperparameters, and any or all of these can be stored in code and / or data memory 705 or code and / or data memory 701 or other on-chip or off-chip memory.
[0095] In at least one embodiment, one or more processors or other hardware logic devices or circuits include one or more ALUs 710, while in another embodiment, one or more ALUs 710 may be located outside the processor or other hardware logic device or the circuitry that uses them (e.g., a coprocessor). In at least one embodiment, one or more ALUs 710 may be included within an execution unit of a processor, or otherwise included in a group of ALUs accessible by the execution unit of the processor, which may be within the same processor or distributed among different processors of different types (e.g., a central processing unit, a graphics processing unit, a fixed-function unit, etc.). In at least one embodiment, code and / or data memory 701, code and / or data memory 705, and activation memory 720 may be the same processor or other hardware logic device or circuitry, while in another embodiment, they may be in different processors or other hardware logic devices or circuitry, or some combination of the same and different processors or other hardware logic devices or circuitry. In at least one embodiment, any portion of activation memory 720 may be included together with other on-chip or off-chip data storage, including the processor's L1, L2, or L3 cache or system memory. Furthermore, inference and / or training code may be stored together with other code accessible to the processor or other hardware logic or circuitry, and may be retrieved and / or processed using the processor’s fetch, decode, schedule, execute, exit, and / or other logic circuitry.
[0096] In at least one embodiment, the active memory 720 may be a cache memory, DRAM, SRAM, non-volatile memory (e.g., flash memory), or other memory. In at least one embodiment, the active memory 720 may be wholly or partially located inside or outside one or more processors or other logic circuits. In at least one embodiment, the choice of whether the active memory 720 is internal to or external to the processor may depend on the availability of on-chip or off-chip memory, the latency requirements for training and / or inference functions, the batch size of data used in inference and / or training the neural network, or some combination of these factors. For example, it may include DRAM, SRAM, flash memory, or other memory types. In at least one embodiment, Figure 7A The inference and / or training logic 715 shown can be used in conjunction with an application-specific integrated circuit (“ASIC”), such as those from Google. Processing unit, from Graphcore TM Inference processing unit (IPU) or from Intel Corp. (e.g., "Lake Crest") processor. In at least one embodiment, Figure 7AThe inference and / or training logic 715 shown can be used in conjunction with central processing unit (“CPU”) hardware, graphics processing unit (“GPU”) hardware, or other hardware such as field programmable gate array (“FPGA”)
[0097] Figure 7B Inference and / or training logic 715 according to at least one or more embodiments is illustrated. In at least one embodiment, the inference and / or training logic 715 may include, but is not limited to, hardware logic, wherein computational resources are dedicated or otherwise uniquely used in conjunction with weight values or other information corresponding to one or more layers of neurons within a neural network. In at least one embodiment, Figure 7B The inference and / or training logic 715 shown can be used in conjunction with an application-specific integrated circuit (ASIC), such as those from Google. Processing unit, from Graphcore TM Inference processing units (IPUs) or from Intel Corp. (e.g., "Lake Crest") processor. In at least one embodiment, Figure 7B The inference and / or training logic 715 shown can be used in conjunction with central processing unit (CPU) hardware, graphics processing unit (GPU) hardware, or other hardware (e.g., field-programmable gate array (FPGA)). In at least one embodiment, the inference and / or training logic 715 includes, but is not limited to, code and / or data memory 701 and code and / or data memory 705, which can be used to store code (e.g., graph code), weight values, and / or other information, including bias values, gradient information, momentum values, and / or other parameter or hyperparameter information. Figure 7B In at least one embodiment shown, each of code and / or data memory 701 and code and / or data memory 705 is associated with dedicated computing resources (e.g., computing hardware 702 and computing hardware 706), respectively. In at least one embodiment, each of computing hardware 702 and computing hardware 706 includes one or more ALUs that perform mathematical functions (e.g., linear algebraic functions) on information stored in code and / or data memory 701 and code and / or data memory 705, respectively, and the results of the function execution are stored in activation memory 720.
[0098] In at least one embodiment, each of the code and / or data memories 701 and 705 and the corresponding computing hardware 702 and 706 corresponds to a different layer of the neural network, such that activation obtained from one “store / computation pair 701 / 702” of the code and / or data memories 701 and computing hardware 702 provides input as input to the next “store / computation pair 705 / 706” of the code and / or data memories 705 and computing hardware 706, in order to reflect the conceptual organization of the neural network. In at least one embodiment, each store / computation pair 701 / 702 and 705 / 706 may correspond to more than one neural network layer. In at least one embodiment, additional store / computation pairs (not shown) may be included in the inference and / or training logic 715 after or in parallel with the store / computation pairs 701 / 702 and 705 / 706.
[0099] Data Center
[0100] Figure 8 An example data center 800 that can be used with at least one embodiment is shown. In at least one embodiment, the data center 800 includes a data center infrastructure layer 810, a framework layer 820, a software layer 830, and an application layer 840.
[0101] In at least one embodiment, such as Figure 8 As shown, the data center infrastructure layer 810 may include a resource coordinator 812, packet computing resources 814, and node computing resources (“nodes CR”) 816(1)-816(N), where “N” represents any positive integer. In at least one embodiment, nodes CR 816(1)-816(N) may include, but are not limited to, any number of central processing units (“CPUs”) or other processors (including accelerators, field-programmable gate arrays (FPGAs), graphics processors, etc.), memory devices (e.g., dynamic read-only memory), storage devices (e.g., solid-state drives or disk drives), network input / output (“NW I / O”) devices, network switches, virtual machines (“VMs”), power modules, and cooling modules, etc. In at least one embodiment, one or more nodes CR 816(1)-816(N) may be servers having one or more of the aforementioned computing resources.
[0102] In at least one embodiment, the grouped computing resource 814 may include individual groups (not shown) of node CRs housed within one or more racks, or a plurality of racks (also not shown) housed within data centers in various geographic locations. The individual groups of node CRs within the grouped computing resource 814 may include computing, networking, memory, or storage resources that can be configured or allocated to support groups of one or more workloads. In at least one embodiment, several node CRs, including CPUs or processors, may be grouped within one or more racks to provide computing resources to support one or more workloads. In at least one embodiment, the one or more racks may also include any number of power modules, cooling modules, and network switches, in any combination.
[0103] In at least one embodiment, resource coordinator 812 may configure or otherwise control one or more nodes CR816(1)-816(N) and / or grouped computing resources 814. In at least one embodiment, resource coordinator 812 may include a software design infrastructure (“SDI”) management entity for data center 800. In at least one embodiment, resource coordinator may include hardware, software, or some combination thereof.
[0104] In at least one embodiment, such as Figure 8 As shown, framework layer 820 includes a job scheduler 822, a configuration manager 824, a resource manager 826, and a distributed file system 828. In at least one embodiment, framework layer 820 may include a framework of software 832 supporting software layer 830 and / or one or more applications 842 supporting application layer 840. In at least one embodiment, software 832 or application 842 may respectively include web-based service software or applications, such as services or applications provided by Amazon Web Services, Google Cloud, and Microsoft Azure. In at least one embodiment, framework layer 820 may be, but is not limited to, a free and open-source software web application framework, such as Apache Spark, which can utilize distributed file system 828 for large-scale data processing (e.g., "big data"). TM(Hereinafter referred to as "Spark"). In at least one embodiment, the job scheduler 822 may include a Spark driver to facilitate the scheduling of workloads supported by various layers of the data center 800. In at least one embodiment, the configuration manager 824 may be able to configure different layers, such as the software layer 830 and the framework layer 820, which includes Spark and a distributed file system 828 for supporting large-scale data processing. In at least one embodiment, the resource manager 826 is able to manage cluster or group computing resources mapped to or allocated to support the distributed file system 828 and the job scheduler 822. In at least one embodiment, the cluster or group computing resources may include group computing resources 814 on the data center infrastructure layer 810. In at least one embodiment, the resource manager 826 may coordinate with the resource coordinator 812 to manage these mapped or allocated computing resources.
[0105] In at least one embodiment, the software 832 included in the software layer 830 may include software used by at least a portion of the nodes CR816(1)-816(N), the grouped computing resources 814, and / or the distributed file system 828 of the framework layer 820. One or more types of software may include, but are not limited to, Internet web page search software, email virus scanning software, database software, and streaming video content software.
[0106] In at least one embodiment, one or more applications 842 included in application layer 840 may include one or more types of applications used by at least a portion of nodes CR816(1)-816(N), grouped computing resources 814, and / or the distributed file system 828 of framework layer 820. One or more types of applications may include, but are not limited to, any number of genomics applications, cognitive computing and machine learning applications, including training or inference software, machine learning framework software (e.g., PyTorch, TensorFlow, Caffe, etc.), or other machine learning applications used in conjunction with one or more embodiments.
[0107] In at least one embodiment, any of the configuration manager 824, resource manager 826, and resource coordinator 812 can perform any number and type of self-modification actions based on any amount and type of data acquired in any technically feasible manner. In at least one embodiment, self-modification actions can mitigate potentially poor configuration decisions by data center operators of data center 800 and can prevent underutilization and / or poor performance of the data center.
[0108] In at least one embodiment, data center 800 may include tools, services, software, or other resources to train one or more machine learning models or to use one or more machine learning models to predict or infer information according to one or more embodiments described herein. For example, in at least one embodiment, a machine learning model can be trained by calculating weight parameters based on a neural network architecture using the software and computing resources described above with respect to data center 800. In at least one embodiment, information can be inferred or predicted using trained machine learning models corresponding to one or more neural networks using the resources described above with respect to data center 800 by using weight parameters calculated through one or more training techniques described herein.
[0109] In at least one embodiment, the data center may use a CPU, application-specific integrated circuit (ASIC), GPU, FPGA, or other hardware to utilize the aforementioned resources to perform training and / or inference. Furthermore, one or more of the aforementioned software and / or hardware resources may be configured as a service to allow a user to train or perform information inference, such as image recognition, speech recognition, or other artificial intelligence services.
[0110] Inference and / or training logic 715 is used to perform inference and / or training operations associated with one or more embodiments. The following is combined with... Figure 7A and / or Figure 7B Details are provided regarding the inference and / or training logic 715. In at least one embodiment, the inference and / or training logic 715 can be implemented in the system. Figure 8 Used in this context for inference or prediction operations based at least in part on weight parameters computed using neural network training operations, neural network functions and / or architectures, or neural network use cases described herein.
[0111] These components can be used to generate synthetic data that simulates failures during network training, which can help improve network performance while limiting the amount of synthetic data to avoid overfitting.
[0112] Computer System
[0113] Figure 9This is a block diagram illustrating an exemplary computer system according to at least one embodiment. The exemplary computer system may be a system of interconnected devices and components, a system-on-a-chip (SoC), or some combination thereof formed with a processor, which may include an execution unit to execute instructions. In at least one embodiment, according to this disclosure, such as the embodiments described herein, computer system 900 may include, but is not limited to, components such as processor 902, whose execution unit includes logic to execute algorithms for process data. In at least one embodiment, computer system 900 may include a processor, such as one available from Intel Corporation of Santa Clara, California. Processor family, Xeon™ XScale™ and / or StrongARM™ Core TM or Nervana TM A microprocessor may be used, although other systems (including PCs, engineering workstations, set-top boxes, etc.) with other microprocessors may also be used. In at least one embodiment, computer system 900 may execute a version of the Windows operating system available from Microsoft Corporation of Redmond, Washington, although other operating systems (such as UNIX and Linux), embedded software, and / or graphical user interfaces may also be used.
[0114] The embodiments can be used in other devices, such as handheld devices and embedded applications. Some examples of handheld devices include cellular phones, Internet Protocol (IP) devices, digital cameras, personal digital assistants (“PDAs”), and handheld PCs. In at least one embodiment, the embedded application may include a microcontroller, a digital signal processor (“DSP”), a system-on-a-chip (SoC), a network computer (“NetPC”), a set-top box, a network hub, a wide area network (“WAN”) switch, or any other system that can execute one or more instructions according to at least one embodiment.
[0115] In at least one embodiment, the computer system 900 may include, but is not limited to, a processor 902, which may include, but is not limited to, one or more execution units 908, to perform machine learning model training and / or inference according to the techniques described herein. In at least one embodiment, the computer system 900 is a single-processor desktop or server system, but in another embodiment, the computer system 900 may be a multiprocessor system. In at least one embodiment, the processor 902 may include, but is not limited to, a Complex Instruction Set Computer (“CISC”) microprocessor, a Reduced Instruction Set Computing (“RISC”) microprocessor, a Very Long Instruction Word (“VLIW”) microprocessor, a processor implementing instruction set combination, or any other processor device, such as a digital signal processor. In at least one embodiment, the processor 902 may be coupled to a processor bus 910, which can transmit data signals between the processor 902 and other components in the computer system 900.
[0116] In at least one embodiment, processor 902 may include, but is not limited to, a Level 1 (“L1”) internal cache memory (“cache”) 904. In at least one embodiment, processor 902 may have a single internal cache or multiple levels of internal cache. In at least one embodiment, the cache memory may reside external to processor 902. Depending on specific implementation and requirements, other embodiments may also include a combination of internal and external caches. In at least one embodiment, register file 906 may store different types of data in various registers, including but not limited to integer registers, floating-point registers, status registers, and instruction pointer registers.
[0117] In at least one embodiment, an execution unit 908, including but not limited to logic for performing integer and floating-point operations, is also located within the processor 902. In at least one embodiment, the processor 902 may further include a microcode (“ucode”) read-only memory (“ROM”) for storing microcode of certain macro instructions. In at least one embodiment, the execution unit 908 may include logic for processing a packaged instruction set 909. In at least one embodiment, by including the packaged instruction set 909 in the instruction set of a general-purpose processor, along with the associated circuitry for executing the instructions, the packaged data in the general-purpose processor 902 can be used to perform operations used by numerous multimedia applications. In one or more embodiments, the execution of numerous multimedia applications can be accelerated and performed more efficiently by using the full width of the processor's data bus to perform operations on the packaged data, which may eliminate the need to transfer smaller data units on the processor's data bus to perform one or more operations on one data element at a time.
[0118] In at least one embodiment, execution unit 908 may also be used in a microcontroller, embedded processor, graphics device, DSP, and other types of logic circuitry. In at least one embodiment, computer system 900 may include, but is not limited to, memory 920. In at least one embodiment, memory 920 may be implemented as a dynamic random access memory (“DRAM”) device, a static random access memory (“SRAM”) device, a flash memory device, or other storage device. In at least one embodiment, memory 920 may store instructions 919 and / or data 921 represented by data signals that can be executed by processor 902.
[0119] In at least one embodiment, the system logic chip may be coupled to the processor bus 910 and the memory 920. In at least one embodiment, the system logic chip may include, but is not limited to, a memory controller hub (“MCH”) 916, and the processor 902 may communicate with the MCH 916 via the processor bus 910. In at least one embodiment, the MCH 916 may provide a high-bandwidth memory path 918 to the memory 920 for instruction and data storage, as well as for storage of graphics commands, data, and textures. In at least one embodiment, the MCH 916 may initiate data signals between the processor 902, the memory 920, and other components in the computer system 900, and bridge data signals between the processor bus 910, the memory 920, and the system I / O 922. In at least one embodiment, the system logic chip may provide a graphics port for coupling to a graphics controller. In at least one embodiment, the MCH 916 may be coupled to the memory 920 via the high-bandwidth memory path 918, and the graphics / video card 912 may be coupled to the MCH 916 via an Accelerated Graphics Port (“AGP”) interconnect 914.
[0120] In at least one embodiment, computer system 900 may use system I / O 922, a proprietary hub interface bus, for coupling MCH 916 to I / O controller hub (“ICH”) 930. In at least one embodiment, ICH 930 may provide direct connectivity to certain I / O devices via a local I / O bus. In at least one embodiment, the local I / O bus may include, but is not limited to, a high-speed I / O bus for connecting peripheral devices to memory 920, chipset, and processor 902. Examples may include, but are not limited to, audio controller 929, firmware hub (“Flash BIOS”) 928, wireless transceiver 926, data memory 924, a conventional I / O controller 923 including user input and keyboard interfaces, serial expansion port 927 (e.g., a Universal Serial Bus (USB) port), and network controller 934. Data memory 924 may include hard disk drives, floppy disk drives, CD-ROM devices, flash memory devices, or other mass storage devices.
[0121] In at least one embodiment, Figure 9 A system including interconnected hardware devices or "chips" is shown, while in other embodiments, Figure 9 An exemplary system-on-a-chip (“SoC”) may be illustrated. In at least one embodiment, the device may be interconnected with a proprietary interconnect, a standardized interconnect (e.g., PCIe), or some combination thereof. In at least one embodiment, one or more components of the computer system 900 are interconnected using a compute fast link (CXL) interconnect.
[0122] The inference and / or training logic 715 is used to perform inference and / or training operations associated with one or more embodiments. This document combines... Figure 7A and / or Figure 7B Details regarding the inference and / or training logic 715 are provided. In at least one embodiment, the inference and / or training logic 715 may be... Figure 9 Used in systems for reasoning or predicting operations based at least in part on weight parameters calculated using neural network training operations, neural network functions and / or architectures or neural network use cases described herein.
[0123] Such components can be used to generate synthetic data that simulates failures during network training, which can help improve network performance while limiting the amount of synthetic data to avoid overfitting.
[0124] Figure 10This is a block diagram illustrating an electronic device 1000 for utilizing a processor 1010 according to at least one embodiment. In at least one embodiment, the electronic device 1000 may be, for example, but not limited to, a laptop computer, tower server, rack server, blade server, desktop computer, tablet computer, mobile device, telephone, embedded computer, or any other suitable electronic device.
[0125] In at least one embodiment, system 1000 may include, but is not limited to, processor 1010 communicatively coupled to any suitable number or type of components, peripherals, modules, or devices. In at least one embodiment, processor 1010 is coupled using a bus or interface, such as I... 2 C-bus, System Management Bus (“SMBus”), Low Pin Count (LPC) bus, Serial Peripheral Interface (“SPI”), High Definition Audio (“HDA”) bus, Serial Advanced Technology Accessory (“SATA”) bus, Universal Serial Bus (“USB”) (versions 1, 2, and 3), or Universal Asynchronous Receiver / Transmitter (“UART”) bus. In at least one embodiment, Figure 10 The system shown includes interconnected hardware devices or "chips," while in other embodiments, Figure 10 An exemplary system-on-a-chip (“SoC”) may be illustrated. In at least one embodiment, Figure 10 The device shown can be interconnected with proprietary interconnects, standardized interconnects (e.g., PCIe), or some combination thereof. In at least one embodiment, Figure 10 One or more components are interconnected using Computational Fast Link (CXL) interconnects.
[0126] In at least one embodiment, Figure 10 It may include a display 1024, a touch screen 1025, a touchpad 1030, a near field communication unit (“NFC”) 1045, a sensor hub 1040, a thermal sensor 1046, a fast chipset (“EC”) 1035, a trusted platform module (“TPM”) 1038, a BIOS / firmware / flash (“BIOS, FW Flash”) 1022, a DSP 1060, a drive 1020 (e.g., a solid-state drive (“SSD”) or a hard disk drive (“HDD”)), a wireless local area network unit (“WLAN”) 1050, a Bluetooth unit 1052, a wireless wide area network unit (“WWAN”) 1056, a global positioning system (GPS) unit 1055, a camera (“USB 3.0 camera”) 1054 (e.g., a USB 3.0 camera), and / or a low-power double data rate (“LPDDR”) memory unit (“LPDDR3”) 1015 implemented in, for example, the LPDDR3 standard. These components can each be implemented in any suitable way.
[0127] In at least one embodiment, other components may be communicatively coupled to processor 1010 via the components discussed herein. In at least one embodiment, accelerometer 1041, ambient light sensor (“ALS”) 1042, compass 1043, and gyroscope 1044 may be communicatively coupled to sensor hub 1040. In at least one embodiment, thermal sensor 1039, fan 1037, keyboard 1036, and touchpad 1030 may be communicatively coupled to EC 1035. In at least one embodiment, speaker 1063, earphone 1064, and microphone (“mic”) 1065 may be communicatively coupled to audio unit (“audio codec and Class D amplifier”) 1062, which in turn may be communicatively coupled to DSP 1060. In at least one embodiment, audio unit 1062 may include, for example, but not limited to, audio encoder / decoder (“codec”) and Class D amplifier. In at least one embodiment, SIM card (“SIM”) 1057 may be communicatively coupled to WWAN unit 1056. In at least one embodiment, components such as WLAN unit 1050, Bluetooth unit 1052, and WWAN unit 1056 can be implemented as next-generation form factor (NGFF).
[0128] Inference and / or training logic 715 is used to perform inference and / or training operations associated with one or more embodiments. This document combines... Figure 7A and / or Figure 7B Details regarding the inference and / or training logic 715 are provided. In at least one embodiment, the inference and / or training logic 715 can be in the system. Figure 10 It is used in the context of reasoning or predicting operations based at least in part on weight parameters calculated using neural network training operations, neural network functions and / or architectures or neural network use cases described herein.
[0129] Such components can be used to generate synthetic data that simulates failures during network training, which can help improve network performance while limiting the amount of synthetic data to avoid overfitting.
[0130] Figure 11 This is a block diagram of a processing system according to at least one embodiment. In at least one embodiment, system 1100 includes one or more processors 1102 and one or more graphics processors 1108, and may be a single-processor desktop system, a multi-processor workstation system, or a server system having a large number of processors 1102 or processor cores 1107. In at least one embodiment, system 1100 is a processing platform incorporated within a system-on-a-chip (SoC) integrated circuit for use in mobile, handheld, or embedded devices.
[0131] In at least one embodiment, system 1100 may include or be integrated into a server-based gaming platform, including a game console, mobile game console, handheld game console, or online game console, which are game and media consoles. In at least one embodiment, system 1100 is a mobile phone, smartphone, tablet computing device, or mobile internet device. In at least one embodiment, processing system 1100 may also include components coupled to or integrated into a wearable device, such as a smartwatch, smart glasses, augmented reality, or virtual reality device. In at least one embodiment, processing system 1100 is a television or set-top box device having one or more processors 1102 and a graphical interface generated by one or more graphics processors 1108.
[0132] In at least one embodiment, each of the one or more processors 1102 includes one or more processor cores 1107 for processing instructions that, when executed, perform operations against the system and user software. In at least one embodiment, each of the one or more processor cores 1107 is configured to process a particular instruction set 1109. In at least one embodiment, the instruction set 1109 may facilitate Complex Instruction Set Computing (CISC), Reduced Instruction Set Computing (RISC), or computation via Very Long Instruction Word (VLIW). In at least one embodiment, each processor core 1107 may process a different instruction set 1109, and the instruction sequence may include instructions that facilitate the emulation of other instruction sets. In at least one embodiment, the processor core 1107 may also include other processing devices, such as a digital signal processor (DSP).
[0133] In at least one embodiment, processor 1102 includes cache memory 1104. In at least one embodiment, processor 1102 may have a single internal cache or multiple levels of internal caches. In at least one embodiment, the cache memory is shared among various components of processor 1102. In at least one embodiment, processor 1102 also uses an external cache (e.g., a Level 3 (L3) cache or a last-level cache (LLC)) (not shown), which can be shared among processor cores 1107 using known cache coherence techniques. In at least one embodiment, processor 1102 further includes a register file 1106, which may include different types of registers for storing different types of data (e.g., integer registers, floating-point registers, status registers, and instruction pointer registers). In at least one embodiment, register file 1106 may include general-purpose registers or other registers.
[0134] In at least one embodiment, one or more processors 1102 are coupled to one or more interface buses 1110 to transmit communication signals, such as address, data, or control signals, between the processors 1102 and other components in the system 1100. In at least one embodiment, the interface bus 1110 may be a processor bus, such as a version of the Direct Media Interface (DMI) bus. In at least one embodiment, the interface bus 1110 is not limited to the DMI bus and may include one or more peripheral component interconnect buses (e.g., PCI, PCI Express), memory buses, or other types of interface buses. In at least one embodiment, the processor 1102 includes an integrated memory controller 1116 and a platform controller hub 1130. In at least one embodiment, the memory controller 1116 facilitates communication between memory devices and other components of the processing system 1100, while the platform controller hub (PCH) 1130 provides connectivity to input / output (I / O) devices via a local I / O bus.
[0135] In at least one embodiment, memory device 1120 may be a dynamic random access memory (DRAM) device, a static random access memory (SRAM) device, a flash memory device, a phase-change memory device, or a device with suitable performance for use as processor memory. In at least one embodiment, memory device 1120 may be used as system memory of processing system 1100 to store data 1122 and instructions 1121 for use when one or more processors 1102 execute an application or process. In at least one embodiment, memory controller 1116 is also coupled to an optional external graphics processor 1112, which may communicate with one or more graphics processors 1108 of processor 1102 to perform graphics and media operations. In at least one embodiment, display device 1111 may be connected to processor 1102. In at least one embodiment, display device 1111 may include one or more internal display devices, such as in mobile electronic devices or laptop devices, or external display devices connected via a display interface (e.g., DisplayPort). In at least one embodiment, the display device 1111 may include a head-mounted display (HMD), such as a stereoscopic display device for virtual reality (VR) or augmented reality (AR) applications.
[0136] In at least one embodiment, the platform controller hub 1130 enables peripheral devices to connect to the storage device 1120 and the processor 1102 via a high-speed I / O bus. In at least one embodiment, the I / O peripheral devices include, but are not limited to, an audio controller 1146, a network controller 1134, a firmware interface 1128, a wireless transceiver 1126, a touch sensor 1125, and a data storage device 1124 (e.g., a hard disk drive, flash memory, etc.). In at least one embodiment, the data storage device 1124 may be connected via a memory interface (e.g., SATA) or via a peripheral bus, such as a peripheral component interconnect bus (e.g., PCI, PCIe). In at least one embodiment, the touch sensor 1125 may include a touchscreen sensor, a pressure sensor, or a fingerprint sensor. In at least one embodiment, the wireless transceiver 1126 may be a Wi-Fi transceiver, a Bluetooth transceiver, or a mobile network transceiver, such as a 3G, 4G, or LTE transceiver. In at least one embodiment, the firmware interface 1128 enables communication with the system firmware and may be, for example, a Unified Extensible Firmware Interface (UEFI). In at least one embodiment, network controller 1134 may enable network connectivity to a wired network. In at least one embodiment, a high-performance network controller (not shown) is coupled to interface bus 1110. In at least one embodiment, audio controller 1146 is a multi-channel high-definition audio controller. In at least one embodiment, processing system 1100 includes an optional legacy I / O controller 1140 for coupling legacy (e.g., Personal System 2 (PS / 2)) devices to system 1100. In at least one embodiment, platform controller hub 1130 may also be connected to one or more Universal Serial Bus (USB) controllers 1142 that connect input devices, such as a keyboard and mouse combination 1143, a camera 1144, or other USB input devices.
[0137] In at least one embodiment, instances of the memory controller 1116 and platform controller hub 1130 may be integrated into a discrete external graphics processor, such as external graphics processor 1112. In at least one embodiment, the platform controller hub 1130 and / or the memory controller 1116 may be external to one or more processors 1102. For example, in at least one embodiment, system 1100 may include external memory controller 1116 and platform controller hub 1130, which may be configured as a memory controller hub and peripheral controller hub in a system chipset communicating with processor 1102.
[0138] Inference and / or training logic 715 is used to perform inference and / or training operations associated with one or more embodiments. This document combines... Figure 7A and / or Figure 7BDetails are provided regarding the inference and / or training logic 715. In at least one embodiment, some or all of the inference and / or training logic 715 may be incorporated into the graphics processor 1500. For example, in at least one embodiment, the training and / or inference techniques described herein may use one or more ALUs embodied in the graphics processor. Furthermore, in at least one embodiment, the inference and / or training operations described herein may use, in addition to Figure 7A or Figure 7B The logic is performed using logic other than that shown. In at least one embodiment, the weight parameters may be stored in on-chip or off-chip memory and / or registers (shown or not shown), which configure the ALU of the graphics processor to execute one or more machine learning algorithms, neural network architectures, use cases, or training techniques described herein.
[0139] Such components can be used to generate synthetic data that simulates failures during network training, which can help improve network performance while limiting the amount of synthetic data to avoid overfitting.
[0140] Figure 12 This is a block diagram of a processor 1200 according to at least one embodiment, the processor having one or more processor cores 1202A-1202N, an integrated memory controller 1214, and an integrated graphics processor 1208. In at least one embodiment, the processor 1200 may include additional cores, up to and including additional cores 1202N indicated by dashed boxes. In at least one embodiment, each processor core 1202A-1202N includes one or more internal cache units 1204A-1204N. In at least one embodiment, each processor core may also access one or more shared cache units 1206.
[0141] In at least one embodiment, internal cache units 1204A-1204N and shared cache unit 1206 represent a cache memory hierarchy within processor 1200. In at least one embodiment, cache memory units 1204A-1204N may include at least one level of instruction and data cache within each processor core and one or more levels of cache in a shared intermediate cache, such as Level 2 (L2), Level 3 (L3), Level 4 (L4), or other levels of cache, wherein the highest level of cache preceding external memory is classified as LLC. In at least one embodiment, cache coherence logic maintains coherence between the various cache units 1206 and 1204A-1204N.
[0142] In at least one embodiment, the processor 1200 may further include a set of one or more bus controller units 1216 and a system agent core 1210. In at least one embodiment, one or more bus controller units 1216 manage a set of peripheral buses, such as one or more PCI or PCIe buses. In at least one embodiment, the system agent core 1210 provides management functions for various processor components. In at least one embodiment, the system agent core 1210 includes one or more integrated memory controllers 1214 to manage access to various external memory devices (not shown).
[0143] In at least one embodiment, one or more processor cores 1202A-1202N include support for multi-threaded concurrent processing. In at least one embodiment, system agent core 1210 includes components for coordinating and operating cores 1202A-1202N during multi-threaded processing. In at least one embodiment, system agent core 1210 may additionally include a power control unit (PCU) including logic and components for regulating one or more power states of processor cores 1202A-1202N and graphics processor 1208.
[0144] In at least one embodiment, processor 1200 further includes a graphics processor 1208 for performing graph processing operations. In at least one embodiment, graphics processor 1208 is coupled to a shared cache unit 1206 and a system proxy core 1210 including one or more integrated memory controllers 1214. In at least one embodiment, system proxy core 1210 further includes a display controller 1211 for driving graphics processor outputs to one or more coupled displays. In at least one embodiment, display controller 1211 may also be a separate module coupled to graphics processor 1208 via at least one interconnect, or it may be integrated within graphics processor 1208.
[0145] In at least one embodiment, ring-based interconnect unit 1212 is used to couple internal components of processor 1200. In at least one embodiment, alternative interconnect units, such as point-to-point interconnects, switched interconnects, or other technologies, may be used. In at least one embodiment, graphics processor 1208 is coupled to ring interconnect 1212 via I / O link 1213.
[0146] In at least one embodiment, I / O link 1213 represents at least one of a variety of I / O interconnects, including packaged I / O interconnects that facilitate communication between various processor components and high-performance embedded memory module 1218 (e.g., eDRAM module). In at least one embodiment, each of processor cores 1202A-1202N and graphics processor 1208 uses embedded memory module 1218 as a shared last-level cache.
[0147] In at least one embodiment, processor cores 1202A-1202N are homogeneous cores executing a common instruction set architecture. In at least one embodiment, processor cores 1202A-1202N are heterogeneous in terms of instruction set architecture (ISA), with one or more processor cores 1202A-1202N executing a common instruction set, while one or more other processor cores 1202A-1202N execute a subset of the common instruction set or a different instruction set. In at least one embodiment, processor cores 1202A-1202N are heterogeneous in terms of microarchitecture, with one or more cores having relatively high power consumption coupled to one or more power cores having lower power consumption. In at least one embodiment, processor 1200 may be implemented on one or more chips or implemented as a SoC integrated circuit.
[0148] Inference and / or training logic 715 is used to perform inference and / or training operations associated with one or more embodiments. This document combines... Figure 7A and / or Figure 7B Details regarding the inference and / or training logic 715 are provided. In at least one embodiment, some or all of the inference and / or training logic 715 may be incorporated into the processor 1200. For example, in at least one embodiment, the training and / or inference techniques described herein may use one or more ALUs embodied in the graphics processor 1512, graphics cores 1202A-1202N, or... Figure 12 Among other components. Furthermore, in at least one embodiment, the inference and / or training operations described herein can use, except... Figure 7A or Figure 7B The logic is performed using logic other than that shown. In at least one embodiment, the weight parameters may be stored in on-chip or off-chip memory and / or registers (shown or not shown), which configure the ALU of the graphics processor 1200 to execute one or more machine learning algorithms, neural network architectures, use cases, or training techniques described herein.
[0149] Such components can be used to generate synthetic data that simulates failures during network training, which can help improve network performance while limiting the amount of synthetic data to avoid overfitting.
[0150] Virtualization computing platform
[0151] Figure 13 This is an example data flow diagram of process 1300 for generating and deploying an image processing and inference pipeline according to at least one embodiment. In at least one embodiment, process 1300 can be deployed for use with imaging devices, processing devices, and / or other device types at one or more facilities 1302. Process 1300 can be executed within training system 1304 and / or deployment system 1306. In at least one embodiment, training system 1304 can be used to train, deploy, and implement machine learning models (e.g., neural networks, object detection algorithms, computer vision algorithms, etc.) for use in deployment system 1306. In at least one embodiment, deployment system 1306 can be configured to offload processing and computing resources between distributed computing environments to reduce infrastructure requirements at facility 1302. In at least one embodiment, one or more applications in the pipeline can use or invoke services of deployment system 1306 (e.g., inference, visualization, computation, AI, etc.) during application execution.
[0152] In at least one embodiment, some applications in the advanced processing and inference pipeline may use machine learning models or other AI to perform one or more processing steps. In at least one embodiment, the machine learning model may be trained at facility 1302 using data 1308 (such as imaging data) generated at facility 1302 (and stored on one or more Picture Archiving and Communication System (PACS) servers at facility 1302), or may be trained using imaging or sequencing data 1308 or a combination thereof from another (one or more) facility. In at least one embodiment, training system 1304 may be used to provide applications, services, and / or other resources for generating jobs, deployable machine learning models, for deployment system 1306.
[0153] In at least one embodiment, the model registry 1324 may be supported by an object storage that supports versioning and object metadata. In at least one embodiment, the object storage may be, for example, cloud storage (e.g., Figure 14 The cloud platform (1426)-compatible application programming interface (API) is accessed from within the cloud platform. In at least one embodiment, machine learning models within the model registry 1324 can be uploaded, listed, modified, or deleted by the developer or partner of the system integrated with the API. In at least one embodiment, the API can provide access to methods that allow a user to associate a model with an application using appropriate credentials, enabling the model to be executed as part of the execution of the application's containerized instantiation.
[0154] In at least one embodiment, training pipeline 1404 ( Figure 14This could include a scenario where facility 1302 is training its own machine learning model, or has an existing machine learning model that needs to be optimized or updated. In at least one embodiment, imaging data 1308 generated by one or more imaging devices, sequencing devices, and / or other device types can be received. In at least one embodiment, once the imaging data 1308 is received, AI-assisted annotation 1310 can be used to help generate annotations corresponding to the imaging data 1308 for use as ground-based data for machine learning models. In at least one embodiment, AI-assisted annotation 1310 can include one or more machine learning models (e.g., convolutional neural networks (CNNs)) that can be trained to generate annotations corresponding to certain types of imaging data 1308 (e.g., from certain devices). In at least one embodiment, AI-assisted annotation 1310 can then be used directly, or it can be adjusted or fine-tuned using annotation tools to generate ground-based data. In at least one embodiment, AI-assisted annotation 1310, labeled clinical data 1312, or a combination thereof can be used as ground-based data for training machine learning models. In at least one embodiment, the trained machine learning model may be referred to as output model 1316 and may be used by deployment system 1306 as described herein.
[0155] In at least one embodiment, training pipeline 1404 ( Figure 14This could include scenarios where facility 1302 requires a machine learning model to perform one or more processing tasks for one or more applications in deployment system 1306, but facility 1302 may not currently have such a machine learning model (or may not have an optimized, efficient, or effective model for such a purpose). In at least one embodiment, an existing machine learning model can be selected from model registry 1324. In at least one embodiment, model registry 1324 can include machine learning models trained to perform various inference tasks on imaging data. In at least one embodiment, the machine learning models in model registry 1324 may have already been trained on imaging data from facilities other than facility 1302 (e.g., remote facilities). In at least one embodiment, the machine learning model may have already been trained on imaging data from one location, two locations, or any number of locations. In at least one embodiment, when training on imaging data from a particular location, training may occur at that location, or at least in a manner that protects the confidentiality of the imaging data or restricts the imaging data from being transmitted in the field. In at least one embodiment, once the model has been trained or partially trained at a location, the machine learning model can be added to model registry 1324. In at least one embodiment, the machine learning model can then be retrained or updated at any number of other facilities, and the retrained or updated model can be made available in model registry 1324. In at least one embodiment, the machine learning model can then be selected from model registry 1324—and referred to as output model 1316—and can be used in deployment system 1306 to perform one or more processing tasks for one or more applications of the deployment system.
[0156] In at least one embodiment, training pipeline 1404 ( Figure 14The scenario may include facility 1302, which requires a machine learning model to perform one or more processing tasks for one or more applications in deployment system 1306, but facility 1302 may not currently have such a machine learning model (or may not have an optimized, efficient, or effective model for such purposes). In at least one embodiment, the machine learning model selected from model registry 1324 may not be fine-tuned or optimized for the imaging data 1308 generated at facility 1302 due to population differences, robustness of training data used to train the machine learning model, anomalous diversity of training data, and / or other problems with the training data. In at least one embodiment, AI-assisted annotation 1310 may be used to help generate annotations corresponding to imaging data 1308, which is used as ground-based data for retraining or updating the machine learning model. In at least one embodiment, labeled data 1312 may be used as ground-based data for training the machine learning model. In at least one embodiment, retraining or updating the machine learning model may be referred to as model training 1314. In at least one embodiment, model training 1314 (e.g., AI-assisted annotation 1310, labeled clinical data 1312, or a combination thereof) can be used as ground-based real-world data for retraining or updating the machine learning model. In at least one embodiment, the trained machine learning model may be referred to as output model 1316 and may be used by deployment system 1306 as described herein.
[0157] In at least one embodiment, deployment system 1306 may include software 1318, service 1320, hardware 1322, and / or other components, features, and functions. In at least one embodiment, deployment system 1306 may include a software "stack" such that software 1318 can be built on top of service 1320 and can be used to perform some or all of the processing tasks, and service 1320 and software 1318 can be built on top of hardware 1322 and use hardware 1322 to perform processing, storage, and / or other computational tasks of deployment system 1306. In at least one embodiment, software 1318 may include any number of different containers, each of which can perform an instantiation of an application. In at least one embodiment, each application can perform one or more processing tasks (e.g., inference, object detection, feature detection, segmentation, image enhancement, calibration, etc.) in a high-level processing and inference pipeline. In at least one embodiment, in addition to receiving and configuring imaging data for use by each container and / or by facility 1302 after processing through the pipeline, advanced processing and inference pipelines (e.g., to convert output back to available data types) can be defined based on the selection of different containers desired or required for processing imaging data 1308. In at least one embodiment, a combination of containers within software 1318 (e.g., constituting a pipeline) may be referred to as a virtual instrument (as described in more detail herein), and the virtual instrument may utilize service 1320 and hardware 1322 to perform some or all of the processing tasks of an application instantiated within the container.
[0158] In at least one embodiment, the data processing pipeline may receive input data (e.g., imaging data 1308) in a specific format in response to an inference request (e.g., a request from a user of deployment system 1306). In at least one embodiment, the input data may represent one or more images, videos, and / or other data representations generated by one or more imaging devices. In at least one embodiment, the data may be preprocessed as part of the data processing pipeline to prepare it for processing by one or more applications. In at least one embodiment, post-processing may be performed on the output of one or more inference tasks or other processing tasks of the pipeline to prepare output data for the next application and / or to prepare output data for user transmission and / or use (e.g., as a response to an inference request). In at least one embodiment, the inference task may be performed by one or more machine learning models, such as trained or deployed neural networks, which may include the output model 1316 of training system 1304.
[0159] In at least one embodiment, the tasks of the data processing pipeline can be encapsulated in containers, each container representing a discrete, fully functional instantiation of an application and a virtualized computing environment capable of referencing a machine learning model. In at least one embodiment, containers or applications can be published to a private (e.g., limited access) area of a container registry (described in more detail herein), and trained or deployed models can be stored in a model registry 1324 and associated with one or more applications. In at least one embodiment, an image of an application (e.g., a container image) can be used in the container registry, and once a user selects an image from the container registry for deployment in the pipeline, that image can be used to generate containers for instantiation of the application for use by the user's system.
[0160] In at least one embodiment, a developer (e.g., a software developer, clinician, physician, etc.) can develop, publish, and store an application (e.g., as a container) for performing image processing and / or inference on provided data. In at least one embodiment, a software development kit (SDK) associated with the system can be used to perform development, publication, and / or storage (e.g., to ensure that the developed application and / or container conforms to or is compatible with the system). In at least one embodiment, the developed application can be tested locally using the SDK (e.g., at a first facility, testing data from a first facility), the SDK serving as a system (e.g.,...). Figure 14 System 1400 may support at least some services 1320. In at least one embodiment, since DICOM objects may contain one to hundreds of images or other data types, and due to variations in the data, the developer may be responsible for managing (e.g., setting up constructs for preprocessing built into the application, etc.) the extraction and preparation of incoming data. In at least one embodiment, once verified by system 1400 (e.g., for accuracy), the application becomes available in the container registry for user selection and / or implementation to perform one or more processing tasks on data at the user's facility (e.g., a second facility).
[0161] In at least one embodiment, the developer can then share the application or container over a network for the system (e.g., Figure 14The system 1400 allows for user access and use. In at least one embodiment, completed and validated applications or containers may be stored in a container registry, and associated machine learning models may be stored in a model registry 1324. In at least one embodiment, a requesting entity (which provides an inference or image processing request) may browse the container registry and / or model registry 1324 to obtain applications, containers, datasets, machine learning models, etc., select desired combinations of elements to include in the data processing pipeline, and submit an image processing request. In at least one embodiment, the request may include input data necessary to execute the request (and, in some examples, patient-related data), and / or may include selections of applications and / or machine learning models to be executed when the request is processed. In at least one embodiment, the request may then be passed to one or more components of the deployment system 1306 (e.g., the cloud) to perform processing in the data processing pipeline. In at least one embodiment, processing performed by the deployment system 1306 may include referencing elements (e.g., applications, containers, models, etc.) selected from the container registry and / or model registry 1324. In at least one embodiment, once the results are generated through the pipeline, the results can be returned to the user for reference (e.g., for viewing in a suite of viewing applications executed locally, on a local workstation, or on a terminal).
[0162] In at least one embodiment, service 1320 may be utilized to assist in processing or executing applications or containers in the pipeline. In at least one embodiment, service 1320 may include computing services, artificial intelligence (AI) services, visualization services, and / or other service types. In at least one embodiment, service 1320 may provide functionality common to one or more applications in software 1318, thus abstracting functionality into services that can be invoked or utilized by applications. In at least one embodiment, the functionality provided by service 1320 can operate dynamically and more efficiently, while also allowing applications to process data in parallel (e.g., using...). Figure 14The parallel computing platform 1430 in the system can be scaled well. In at least one embodiment, it is not required that each application providing the same functionality as service 1320 must have a corresponding instance of service 1320, but service 1320 can be shared between and among various applications. In at least one embodiment, as a non-limiting example, the service may include an inference server or engine that can be used to perform detection or segmentation tasks. In at least one embodiment, a model training service may be included, which can provide the ability to train and / or retrain machine learning models. In at least one embodiment, a data augmentation service may be further included, which can provide GPU-accelerated data (e.g., DICOM, RIS, CIS, conforming to REST, RPC, raw, etc.) extraction, resizing, scaling, and / or other enhancements. In at least one embodiment, a visualization service may be used, which can add image rendering effects (e.g., ray tracing, rasterization, denoising, sharpening, etc.) to add realism to two-dimensional (2D) and / or three-dimensional (3D) models. In at least one embodiment, a virtual instrument service may be included, which provides beamforming, segmentation, inference, imaging, and / or support for other applications within the virtual instrument pipeline.
[0163] In at least one embodiment, where service 1320 includes an AI service (e.g., an inference service), as part of application execution, one or more machine learning models can be executed by invoking (e.g., as an API call) the inference service (e.g., an inference server) to execute one or more machine learning models or their processing. In at least one embodiment, where another application includes one or more machine learning models for a segmentation task, the application can invoke the inference service to execute the machine learning models for performing one or more processing operations associated with the segmentation task. In at least one embodiment, software 1318 implementing advanced processing and inference pipelines, including a segmentation application and an anomaly detection application, can be pipelined because each application can invoke the same inference service to execute one or more inference tasks.
[0164] In at least one embodiment, hardware 1322 may include a GPU, CPU, graphics card, AI / deep learning system (e.g., an AI supercomputer, such as NVIDIA's DGX), cloud platform, or a combination thereof. In at least one embodiment, different types of hardware 1322 may be used to provide efficient, specially built support for software 1318 and services 1320 in deployment system 1306. In at least one embodiment, GPU processing may be used to perform local processing (e.g., at facility 1302) within the AI / deep learning system, in the cloud system, and / or other processing components of deployment system 1306 to improve the efficiency, accuracy, and performance of image processing and generation. In at least one embodiment, as a non-limiting example, software 1318 and / or services 1320 may be optimized for GPU processing in relation to deep learning, machine learning, and / or high-performance computing. In at least one embodiment, at least some of the computing environment of deployment system 1306 and / or training system 1304 may be executed in a data center, one or more supercomputers, or high-performance computing systems with GPU-optimized software (e.g., a hardware and software combination of an NVIDIA DGX system). In at least one embodiment, as described herein, hardware 1322 may include any number of GPUs that can be invoked to perform data processing in parallel. In at least one embodiment, the cloud platform may also include GPU-optimized execution for deep learning tasks, GPU processing for machine learning tasks, or other computational tasks. In at least one embodiment, an AI / deep learning supercomputer and / or GPU-optimized software (e.g., as provided on NVIDIA's DGX systems) may be used as a hardware abstraction and scaling platform to execute the cloud platform (e.g., NVIDIA's NGC). In at least one embodiment, the cloud platform may integrate application container cluster systems or coordination systems (e.g., Kubernetes) across multiple GPUs to achieve seamless scaling and load balancing.
[0165] Figure 14 This is a system diagram of an example system 1400 for generating and deploying an imaging deployment pipeline according to at least one embodiment. In at least one embodiment, system 1400 can be used to implement Figure 13 The process 1300 and / or other processes include advanced processing and inference pipelines. In at least one embodiment, system 1400 may include training system 1304 and deployment system 1306. In at least one embodiment, training system 1304 and deployment system 1306 may be implemented using software 1318, service 1320 and / or hardware 1322, as described herein.
[0166] In at least one embodiment, system 1400 (e.g., training system 1304 and / or deployment system 1306) may be implemented in a cloud computing environment (e.g., using cloud 1426). In at least one embodiment, system 1400 may be implemented locally (in relation to a healthcare facility) or as a combination of cloud computing resources and local computing resources. In at least one embodiment, access to the API in cloud 1426 may be restricted to authorized users by establishing security measures or protocols. In at least one embodiment, the security protocol may include a network token, which may be signed by an authentication service (e.g., AuthN, AuthZ, Gluecon, etc.) and may carry appropriate authorization. In at least one embodiment, the API of the virtual instrument (described herein) or other instances of system 1400 may be restricted to a set of public IPs that have been audited or authorized for interaction.
[0167] In at least one embodiment, the various components of system 1400 may communicate with each other using any of a variety of different network types, including but not limited to local area networks (LANs) and / or wide area networks (WANs) via wired and / or wireless communication protocols. In at least one embodiment, communication between facilities and components of system 1400 (e.g., for sending inference requests, for receiving the results of inference requests, etc.) may be transmitted via one or more data buses, wireless data protocols (Wi-Fi), wired data protocols (e.g., Ethernet), etc.
[0168] In at least one embodiment, similar to the description herein. Figure 13 As described, training system 1304 can execute training pipeline 1404. In at least one embodiment, where deployment system 1306 uses one or more machine learning models in deployment pipeline 1410, training pipeline 1404 can be used to train or retrain one or more (e.g., pre-trained) models, and / or implement one or more pre-trained models 1406 (e.g., without retraining or updating). In at least one embodiment, as a result of training pipeline 1404, output model 1316 can be generated. In at least one embodiment, training pipeline 1404 can include any number of processing steps, such as, but not limited to, transformation or adaptation of imaging data (or other input data). In at least one embodiment, different training pipelines 1404 can be used for different machine learning models used by deployment system 1306. In at least one embodiment, similar to the description of... Figure 13 The training pipeline 1404 described in the first example can be used for the first machine learning model, similar to the one described above. Figure 13 The training pipeline 1404 described in the second example can be used for a second machine learning model, similar to the one described above. Figure 13The training pipeline 1404 of the third example described can be used for a third machine learning model. In at least one embodiment, any combination of tasks within the training system 1304 can be used according to the requirements of each respective machine learning model. In at least one embodiment, one or more machine learning models may have already been trained and are ready for deployment, so the training system 1304 may not perform any processing on the machine learning models, and one or more machine learning models may be implemented by the deployment system 1306.
[0169] In at least one embodiment, depending on the implementation or embodiment, the output model 1316 and / or the pre-trained model 1406 may include any type of machine learning model. In at least one embodiment, and not limited thereto, the machine learning model used by system 1400 may include models using linear regression, logistic regression, decision trees, support vector machines (SVM), Naive Bayes, k-nearest neighbors (Knn), k-means clustering, random forests, dimensionality reduction algorithms, gradient boosting algorithms, neural networks (e.g., autoencoders, convolutions, recursion, perceptrons, long / short-term memory (LSTM), Hopfield, Boltzmann, deep belief, deconvolution, generative adversarial, liquid state machines, etc.), and / or other types of machine learning models.
[0170] In at least one embodiment, the training pipeline 1404 may include AI-assisted annotations, as described herein regarding at least Figure 15BMore specifically, in at least one embodiment, the labeled data 1312 can be generated using any number of techniques (e.g., conventional annotations). In at least one embodiment, in some examples, labels or other annotations can be generated by drawing programs (e.g., annotation programs), computer-aided design (CAD) programs, tagging programs, another type of application suitable for generating annotations or labels for ground reality, and / or can be hand-drawn. In at least one embodiment, ground reality data can be synthetically generated (e.g., generated from computer models or renderings), realistically generated (e.g., designed and generated from real-world data), automatically generated by machines (e.g., extracting features from data using feature analysis and learning, and then generating labels), manually annotated (e.g., taggers or annotation experts, defining the placement of labels), and / or combinations thereof. In at least one embodiment, for each instance of imaging data 1308 (or other data types used by machine learning models), there may be corresponding ground reality data generated by training system 1304. In at least one embodiment, AI-assisted annotation can be performed as part of deployment pipeline 1410; supplementing or replacing AI-assisted annotations included in training pipeline 1404. In at least one embodiment, system 1400 may include a multi-layer platform, which may include a software layer (e.g., software 1318) of a diagnostic application (or other application type) capable of performing one or more medical imaging and diagnostic functions. In at least one embodiment, system 1400 may be communicatively coupled (e.g., via an encrypted link) to a network of PACS servers in one or more facilities. In at least one embodiment, system 1400 may be configured to access and reference data from PACS servers to perform operations such as training machine learning models, deploying machine learning models, image processing, inference, and / or other operations.
[0171] In at least one embodiment, the software layer may be implemented as a secure, encrypted, and / or certified API that can invoke (e.g., call) an application or container from an external environment (e.g., facility 1302). In at least one embodiment, the application may then invoke or execute one or more services 1320 to perform computational, AI, or visualization tasks associated with their respective applications, and the software 1318 and / or service 1320 may utilize the hardware 1322 to perform processing tasks efficiently and effectively.
[0172] In at least one embodiment, deployment system 1306 may execute deployment pipeline 1410. In at least one embodiment, deployment pipeline 1410 may include any number of applications, which may be sequential, non-sequential, or otherwise applied to imaging data (and / or other data types) – including AI-assisted annotation, the imaging data being generated by imaging devices, sequencing devices, genomics devices, etc., as described above. In at least one embodiment, as described herein, deployment pipeline 1410 for an individual device may be referred to as a virtual instrument for the device (e.g., a virtual ultrasound instrument, a virtual CT scanner, a virtual sequencing instrument, etc.). In at least one embodiment, for a single device, more than one deployment pipeline 1410 may exist, depending on the desired information from the data generated from the device. In at least one embodiment, a first deployment pipeline 1410 may exist if it is desired to detect an anomaly from an MRI machine, and a second deployment pipeline 1410 may exist if it is desired to perform image enhancement from the output of the MRI machine.
[0173] In at least one embodiment, the image generation application may include processing tasks that include using a machine learning model. In at least one embodiment, a user may wish to use their own machine learning model or select a machine learning model from the model registry 1324. In at least one embodiment, a user may implement their own machine learning model or select a machine learning model to be included in the application performing the processing task. In at least one embodiment, the application may be optional and customizable, and by defining the construction of the application, the deployment and implementation of the application for a particular user is presented as a more seamless user experience. In at least one embodiment, by leveraging other features of system 1400 (e.g., service 1320 and hardware 1322), deployment pipeline 1410 can be more user-friendly, provide easier integration, and produce more accurate, efficient, and timely results.
[0174] In at least one embodiment, deployment system 1306 may include user interface 1414 (e.g., graphical user interface, web interface, etc.) which may be used to select applications to be included in deployment pipeline 1410, deploy applications, modify or change applications or their parameters or configurations, use and interact with deployment pipeline 1410 during setup and / or deployment, and / or otherwise interact with deployment system 1306. In at least one embodiment, although not shown with respect to training system 1304, user interface 1414 (or different user interfaces) may be used to select models to be used in deployment system 1306, to select models to be trained or retrained in training system 1304, and / or to otherwise interact with training system 1304.
[0175] In at least one embodiment, in addition to the application coordination system 1428, a pipeline manager 1412 may also be used to manage interactions between applications or containers deploying pipeline 1410 and services 1320 and / or hardware 1322. In at least one embodiment, the pipeline manager 1412 may be configured to facilitate interactions from application to application, from application to service 1320, and / or from application or service to hardware 1322. In at least one embodiment, although shown as included in software 1318, this is not intended to be limiting, and in some examples (e.g., as...) Figure 14 As shown, pipeline manager 1412 may be included in service 1320. In at least one embodiment, application coordination system 1428 (e.g., Kubernetes, DOCKER, etc.) may include container coordination system that can group applications into containers as logical units for coordination, management, scaling, and deployment. In at least one embodiment, by associating applications (e.g., rebuilding applications, splitting applications, etc.) from deployment pipeline 1410 with individual containers, each application can execute in a self-contained environment (e.g., at the kernel level) to improve speed and efficiency.
[0176] In at least one embodiment, each application and / or container (or its image) can be developed, modified, and deployed independently (e.g., a first user or developer can develop, modify, and deploy a first application, and a second user or developer can develop, modify, and deploy a second application separate from the first user or developer). This allows focus on the tasks of a single application and / or container without being hindered by the tasks of another application or container. In at least one embodiment, the pipeline manager 1412 and the application coordination system 1428 can facilitate communication and collaboration between different containers or applications. In at least one embodiment, the application coordination system 1428 and / or the pipeline manager 1412 can facilitate communication and resource sharing between and within each application or container, provided that the expected inputs and / or outputs of each container or application are known to the system (e.g., based on the construction of the application or container). In at least one embodiment, since one or more applications or containers in the deployment pipeline 1410 can share the same services and resources, the application coordination system 1428 can coordinate, load balance, and determine the sharing of services or resources between and within the various applications or containers. In at least one embodiment, the scheduler can be used to track the resource requirements of applications or containers, the current or planned use of these resources, and resource availability. Therefore, in at least one embodiment, the scheduler can allocate resources to different applications and distribute resources between and among applications, taking into account the system's needs and availability. In some examples, the scheduler (and / or other components of the application coordination system 1428) can determine resource availability and distribution based on constraints imposed on the system (e.g., user constraints), such as Quality of Service (QoS), the urgency of data output (e.g., to determine whether to perform real-time processing or delayed processing), etc.
[0177] In at least one embodiment, service 1320, utilized and shared by applications or containers in deployment system 1306, may include computing service 1416, AI service 1418, visualization service 1420, and / or other service types. In at least one embodiment, an application may invoke (e.g., execute) one or more services 1320 to perform processing operations for the application. In at least one embodiment, an application may utilize computing service 1416 to perform supercomputing or other high-performance computing (HPC) tasks. In at least one embodiment, one or more computing services 1416 may be utilized to perform parallel processing (e.g., using parallel computing platform 1430) to process data substantially simultaneously through one or more applications and / or one or more tasks of a single application. In at least one embodiment, parallel computing platform 1430 (e.g., NVIDIA's CUDA) may implement general-purpose computing on a GPU (GPGPU) (e.g., GPU 1422). In at least one embodiment, the software layer of parallel computing platform 1430 may provide access to the GPU's virtual instruction set and parallel computing elements to execute computing kernels. In at least one embodiment, the parallel computing platform 1430 may include memory, and in some embodiments, memory may be shared between and within multiple containers, and / or between and within different processing tasks within a single container. In at least one embodiment, inter-process communication (IPC) calls may be generated for multiple containers and / or multiple processes within containers to use the same data from a shared memory segment of the parallel computing platform 1430 (e.g., where multiple different stages of one or more applications are processing the same information). In at least one embodiment, instead of copying data and moving it to different locations in memory (e.g., read / write operations), the same data in the same memory location can be used for any number of processing tasks (e.g., at the same time, at different times, etc.). In at least one embodiment, this information about the new location of the data can be stored and shared between applications because the resulting data from processing is used to generate new data. In at least one embodiment, the location of the data, and the location of the updated or modified data, may be part of the definition of how the payload in the container is understood.
[0178] In at least one embodiment, AI service 1418 may be used to perform an inference service for executing a machine learning model associated with the application (e.g., a task to perform one or more processing tasks of the application). In at least one embodiment, AI service 1418 may utilize AI system 1424 to execute a machine learning model (e.g., a neural network such as a CNN) for segmentation, reconstruction, object detection, feature detection, classification, and / or other inference tasks. In at least one embodiment, the application deploying pipeline 1410 may use one or more output models 1316 from training system 1304 and / or other models of the application to perform inference on imaging data. In at least one embodiment, two or more examples of using application coordination system 1428 (e.g., a scheduler) for inference may be available. In at least one embodiment, a first category may include a high-priority / low-latency path that can implement a higher service level protocol, such as for performing inference on urgent requests in emergency situations or for radiologists during diagnostic procedures. In at least one embodiment, a second category may include a standard priority path that can be used for requests that may not be urgent or for situations where analysis can be performed at a later time. In at least one embodiment, the application coordination system 1428 may allocate resources (e.g., services 1320 and / or hardware 1322) based on priority paths for different inference tasks of the AI service 1418.
[0179] In at least one embodiment, shared memory may be installed into AI service 1418 in system 1400. In at least one embodiment, shared memory may operate as a cache (or other storage device type) and may be used to process inference requests from applications. In at least one embodiment, when an inference request is submitted, a set of API instances of deployment system 1306 may receive the request and may select one or more instances (e.g., for best fit, for load balancing, etc.) to process the request. In at least one embodiment, to process the request, the request may be fed into a database, and if not already in the cache, a machine learning model may be located from model registry 1324. A verification step may ensure that an appropriate machine learning model is loaded into the cache (e.g., shared memory), and / or a copy of the model may be saved to the cache. In at least one embodiment, if the application is not already running or there are not enough instances of the application, a scheduler (e.g., the scheduler of pipeline manager 1412) may be used to start the application referenced in the request. In at least one embodiment, if an inference server has not yet been started to execute the model, an inference server may be started. Any number of inference servers may be started for each model. In at least one embodiment, in a pull model that clusters inference servers, the model can be cached whenever load balancing is favorable. In at least one embodiment, the inference servers can be statically loaded into the corresponding distributed servers.
[0180] In at least one embodiment, an inference server running in a container can be used to perform inference. In at least one embodiment, an instance of the inference server can be associated with a model (and optionally multiple versions of the model). In at least one embodiment, if an instance of the inference server does not exist when a request to perform inference on the model is received, a new instance can be loaded. In at least one embodiment, when the inference server is started, a model can be passed to the inference server, allowing the same container to be used to serve different models, as long as the inference server runs as different instances.
[0181] In at least one embodiment, during application execution, an inference request for a given application can be received, and a container (e.g., an instance of a hosted inference server) can be loaded (if not already loaded), and a launcher can be invoked. In at least one embodiment, preprocessing logic within the container can (e.g., using a CPU and / or GPU) load, decode, and / or perform any additional preprocessing on the incoming data. In at least one embodiment, once the data is ready for inference, the container can infer the data as needed. In at least one embodiment, this can include a single inference call for an image (e.g., a hand X-ray) or can request inference for hundreds of images (e.g., a chest CT scan). In at least one embodiment, the application can summarize the results before completion, which may include, but is not limited to, a single confidence score, pixel-level segmentation, voxel-level segmentation, generating visualizations, or generating text to summarize the results. In at least one embodiment, different priorities can be assigned to different models or applications. For example, some models may have a real-time (TAT less than 1 minute) priority, while other models may have a lower priority (e.g., TAT less than 10 minutes). In at least one embodiment, model execution time can be measured from the requesting agency or entity, and may include cooperative network traversal time and inference service execution time.
[0182] In at least one embodiment, the transfer of requests between service 1320 and the inference application can be hidden behind a software development kit (SDK) and robust transfer can be provided via queues. In at least one embodiment, requests are placed in queues via an API for individual application / tenant ID combinations, and the SDK pulls requests from the queues and provides them to the application. In at least one embodiment, the name of the queue can be provided in the environment where the SDK picks up the queue. In at least one embodiment, asynchronous communication via queues may be useful because it allows any instance of the application to pick up work when it becomes available. Results can be sent back via queues to ensure no data loss. In at least one embodiment, queues can also provide the ability to partition work, as the highest priority work can go into a queue connected to a majority of instances of the application, while the lowest priority work can go into a queue connected to a single instance that processes tasks in the order they are received. In at least one embodiment, the application can run on a GPU-accelerated instance generated in cloud 1426, and the inference service can perform inference on the GPU.
[0183] In at least one embodiment, visualization service 1420 can be used to generate visualizations for viewing the output of application and / or deployment pipeline 1410. In at least one embodiment, visualization service 1420 can utilize GPU 1422 to generate visualizations. In at least one embodiment, visualization service 1420 can implement rendering effects such as ray tracing to generate higher quality visualizations. In at least one embodiment, visualizations can include, but are not limited to, 2D image rendering, 3D volume rendering, 3D volume reconstruction, 2D tomographic slicing, virtual reality display, augmented reality display, etc. In at least one embodiment, a virtualized environment can be used to generate virtual interactive displays or environments (e.g., virtual environments) for system users (e.g., doctors, nurses, radiologists, etc.) to interact with. In at least one embodiment, visualization service 1420 can include an internal visualizer, cinematic and / or other rendering or image processing capabilities or functions (e.g., ray tracing, rasterization, internal optics, etc.).
[0184] In at least one embodiment, hardware 1322 may include GPU 1422, AI system 1424, cloud 1426, and / or any other hardware for performing training system 1304 and / or deployment system 1306. In at least one embodiment, GPU 1422 (e.g., NVIDIA's TESLA and / or QUADRO GPUs) may include any number of GPUs that can be used to perform processing tasks for any feature or function of computing service 1416, AI service 1418, visualization service 1420, other services, and / or software 1318. For example, for AI service 1418, GPU 1422 may be used to perform preprocessing on imaging data (or other data types used by machine learning models), postprocessing on the output of machine learning models, and / or perform inference (e.g., to execute machine learning models). In at least one embodiment, cloud 1426, AI system 1424, and / or other components of system 1400 may use GPU 1422. In at least one embodiment, cloud 1426 may include a GPU-optimized platform for deep learning tasks. In at least one embodiment, AI system 1424 may use a GPU, and one or more AI systems 1424 may be used to perform cloud 1426 (or at least part of a task for deep learning or inference). Similarly, although hardware 1322 is shown as a discrete component, this is not intended to be limiting, and any component of hardware 1322 may be combined with or utilized by any other component of hardware 1322.
[0185] In at least one embodiment, AI system 1424 may include a specially built computing system (e.g., a supercomputer or HPC) configured for inference, deep learning, machine learning, and / or other artificial intelligence tasks. In at least one embodiment, in addition to CPU, RAM, memory, and / or other components, features, or functions, AI system 1424 (e.g., NVIDIA's DGX) may also include GPU-optimized software (e.g., a software stack) that can be executed using multiple GPUs 1422. In at least one embodiment, one or more AI systems 1424 may be implemented in a cloud 1426 (e.g., in a data center) to perform some or all of the AI-based processing tasks of system 1400.
[0186] In at least one embodiment, cloud 1426 may include GPU-accelerated infrastructure (e.g., NVIDIA's NGC) that can provide a GPU-optimized platform for performing processing tasks of system 1400. In at least one embodiment, cloud 1426 may include AI system 1424 for performing one or more AI-based tasks of system 1400 (e.g., as a hardware abstraction and scaling platform). In at least one embodiment, cloud 1426 may be integrated with application coordination system 1428 utilizing multiple GPUs to achieve seamless scaling and load balancing between and within applications and services 1320. In at least one embodiment, as described herein, cloud 1426 may be responsible for performing at least some of the services 1320 of system 1400, including computing service 1416, AI service 1418, and / or visualization service 1420. In at least one embodiment, cloud 1426 may perform large and small batch inference (e.g., perform NVIDIA's TENSOR RT), provide accelerated parallel computing APIs and platform 1430 (e.g., NVIDIA's CUDA), perform application coordination system 1428 (e.g., KUBERNETES), provide graphics rendering APIs and platform (e.g., for ray tracing, 2D graphics, 3D graphics and / or other rendering techniques to produce higher quality cinematic effects), and / or provide other functionalities for system 1400.
[0187] Figure 15A A data flow diagram of a process 1500 for training, retraining, or updating a machine learning model according to at least one embodiment is shown. In at least one embodiment, a non-limiting example can be used. Figure 14System 1400 executes process 1500. In at least one embodiment, process 1500 may utilize service 1320 and / or hardware 1322 of system 1400, as described herein. In at least one embodiment, the refined model 1512 generated by process 1500 may be executed by deployment system 1306 for one or more containerized applications in deployment pipeline 1410.
[0188] In at least one embodiment, model training 1314 may include retraining or updating the initial model 1504 (e.g., a pre-trained model) using new training data (e.g., new input data, such as customer dataset 1506, and / or new ground reality data associated with the input data). In at least one embodiment, to retrain or update the initial model 1504, the output or loss layer of the initial model 1504 may be reset or deleted, and / or replaced with an updated or new output or loss layer. In at least one embodiment, the initial model 1504 may have previously finely tuned parameters (e.g., weights and / or biases) retained from previous training, so training or retraining 1314 may not require as much time or processing as training the model from scratch. In at least one embodiment, during model training 1314, by resetting or replacing the output or loss layer of the initial model 1504, on a new customer dataset 1506 (e.g., new input data, such as customer dataset 1506, and / or new ground reality data associated with the input data), the initial model 1504 may be retrained or updated. Figure 13 When generating predictions on image data (1308), the parameters of the new dataset can be updated and readjusted based on the loss calculation associated with the accuracy of the output or loss layer.
[0189] In at least one embodiment, the pre-trained model 1406 may be stored in a data storage or registry (e.g., Figure 13(Model registry 1324). In at least one embodiment, the pre-trained model 1406 may have been trained at least partially at one or more facilities other than the facility executing process 1500. In at least one embodiment, to protect the privacy and rights of patients, subjects, or customers at different facilities, the pre-trained model 1406 may have been trained locally using locally generated customer or patient data. In at least one embodiment, the pre-trained model 1406 may be trained using cloud 1426 and / or other hardware 1322, but confidential, privacy-protected patient data may not be transferred to, used by, or accessed by any component of cloud 1426 (or other non-local hardware). In at least one embodiment, if the pre-trained model 1406 is trained using patient data from more than one facility, the pre-trained model 1406 may have been trained separately for each facility before training on patient or customer data from another facility. In at least one embodiment, such as when customer or patient data has been published for privacy reasons (e.g., by abandonment, for experimental purposes, etc.), or where customer or patient data is included in a public dataset, customer or patient data from any number of facilities can be used to train a pre-trained model 1406 locally and / or externally, such as in a data center or other cloud computing infrastructure.
[0190] In at least one embodiment, when selecting an application for use in deployment pipeline 1410, the user may also select a machine learning model for a specific application. In at least one embodiment, the user may not have a model available, so the user may select a pre-trained model 1406 to use with the application. In at least one embodiment, the pre-trained model 1406 may not be optimized to generate accurate results on the user facility's customer dataset 1506 (e.g., based on patient diversity, demographics, type of medical imaging equipment used, etc.). In at least one embodiment, the pre-trained model 1406 may be updated, retrained, and / or fine-tuned for use at various facilities before being deployed to deployment pipeline 1410 for use with one or more applications.
[0191] In at least one embodiment, a user may select a pre-trained model 1406 to be updated, retrained, and / or fine-tuned, and the pre-trained model 1406 may be referred to as the initial model 1504 of the training system 1304 in process 1500. In at least one embodiment, a client dataset 1506 (e.g., imaging data, genomic data, sequencing data, or other data types generated by equipment at the facility) may be used to perform model training 1314 (which may include, but is not limited to, transfer learning) on the initial model 1504 to generate a refined model 1512. In at least one embodiment, ground-based data corresponding to the client dataset 1506 may be generated by the training system 1304. In at least one embodiment, ground-based data (e.g., such as...) may be generated at the facility at least in part by clinicians, scientists, physicians, practitioners, etc. Figure 13 Clinical data marked in 1312).
[0192] In at least one embodiment, AI-assisted annotation 1310 may be used in some examples to generate ground reality data. In at least one embodiment, AI-assisted annotation 1310 (e.g., implemented using an AI-assisted annotation SDK) may leverage machine learning models (e.g., neural networks) to generate suggested or predicted ground reality data for a customer dataset. In at least one embodiment, user 1510 may use the annotation tool within a user interface (graphical user interface (GUI)) on computing device 1508.
[0193] In at least one embodiment, user 1510 can interact with the GUI via computing device 1508 to edit or fine-tune (automatic) annotations. In at least one embodiment, polygon editing features can be used to move the vertices of a polygon to more precise or fine-tuned positions.
[0194] In at least one embodiment, once the customer dataset 1506 has associated ground-based data, the ground-based data (e.g., from AI-assisted annotations, manual labeling, etc.) can be used to generate a refined model 1512 during model training 1314. In at least one embodiment, the customer dataset 1506 can be applied to the initial model 1504 an arbitrary number of times, and the ground-based data can be used to update the parameters of the initial model 1504 until an acceptable level of accuracy is achieved for the refined model 1512. In at least one embodiment, once the refined model 1512 is generated, it can be deployed within one or more deployment pipelines 1410 at the facility to perform one or more processing tasks related to medical imaging data.
[0195] In at least one embodiment, the refined model 1512 can be uploaded to the pre-trained model 1406 in the model registry 1324 for selection by another facility. In at least one embodiment, this process can be completed at any number of facilities, allowing the refined model 1512 to be further refined any number of times on a new dataset to generate a more general model.
[0196] Figure 15B This is an example illustration of a client-server architecture 1532 for enhancing an annotation tool using a pre-trained annotation model, according to at least one embodiment. In at least one embodiment, an AI-assisted annotation tool 1536 may be instantiated based on the client-server architecture 1532. In at least one embodiment, the annotation tool 1536 in an imaging application can assist radiologists, for example, in identifying organs and abnormalities. In at least one embodiment, the imaging application may include software tools, as a non-limiting example, that help user 1510 identify several extreme points on a specific organ of interest in a raw image 1534 (e.g., in a 3D MRI or CT scan) and receive automatic annotation results for all 2D slices of that specific organ. In at least one embodiment, the results may be stored in a data store as training data 1538 and used as (e.g., but not limited to) ground-based data for training. In at least one embodiment, when computing device 1508 sends extreme points for AI-assisted annotation 1310, for example, a deep learning model may receive this data as input and return inference results for segmenting organs or abnormalities. In at least one embodiment, a pre-instantiated annotation tool (e.g., Figure 15B The AI-assisted annotation tool 1536B can be enhanced by making API calls (e.g., API call 1544) to a server (such as annotation assistant server 1540), which may include a set of pre-trained models 1542 stored, for example, in an annotation model registry. In at least one embodiment, the annotation model registry may store pre-trained models 1542 (e.g., machine learning models, such as deep learning models) that have been pre-trained to perform AI-assisted annotation on specific organs or abnormalities. These models can be further updated using training pipeline 1404. In at least one embodiment, the pre-installed annotation tool can be improved over time as new labeled clinical data 1312 is added.
[0197] Such components can be used to generate synthetic data that simulates failures during network training, which can help improve network performance while limiting the amount of synthetic data to avoid overfitting.
[0198] autonomous vehicles
[0199] Figure 16AAn example of an autonomous vehicle 1600 according to at least one embodiment is shown. In at least one embodiment, the autonomous vehicle 1600 (which may alternatively be referred to herein as "vehicle 1600") may be, but is not limited to, a passenger vehicle, such as a car, truck, bus, and / or another type of vehicle capable of accommodating one or more passengers. In at least one embodiment, vehicle 1600 may be a semi-tractor-trailer for hauling goods. In at least one embodiment, vehicle 1600 may be an aircraft, robotic vehicle, or other type of vehicle.
[0200] Autonomous vehicles can be described according to the levels of automation defined by the National Highway Traffic Safety Administration (“NHTSA”) and the Society of Automotive Engineers (“SAE”) of the U.S. Department of Transportation in their standard “Taxonomy and Definitions for Terms Related to Driving Automation Systems for On-Road Motor Vehicles” (e.g., standard number J3016-201806, published June 15, 2018; standard number J3016-201609, published September 30, 2016; and previous and future versions of this standard). In one or more embodiments, vehicle 1600 may be able to function according to one or more of the levels of autonomous driving from Level 1 to Level 5. For example, in at least one embodiment, vehicle 1600 may be able to perform conditional automation (Level 3), high automation (Level 4), and / or full automation (Level 5).
[0201] In at least one embodiment, vehicle 1600 may include, but is not limited to, components such as chassis, body, wheels (e.g., 2, 4, 6, 8, 18, etc.), tires, axles, and other vehicle components. In at least one embodiment, vehicle 1600 may include, but is not limited to, propulsion system 1650, such as an internal combustion engine, a hybrid powertrain, an all-electric motor, and / or another type of propulsion system. In at least one embodiment, propulsion system 1650 may be connected to the drivetrain of vehicle 1600, which may include, but is not limited to, a transmission, to enable propulsion of vehicle 1600. In at least one embodiment, propulsion system 1650 may be controlled in response to receiving a signal from throttle / accelerator 1652.
[0202] In at least one embodiment, when the propulsion system 1650 is operating (e.g., when the vehicle 1600 is traveling), the steering system 1654 (which may include, but is not limited to, a steering wheel) is used to steer the vehicle 1600 (e.g., along a desired path or route). In at least one embodiment, the steering system 1654 may receive signals from the steering actuator 1656. The steering wheel may be optional for fully automated (Level 5) functionality. In at least one embodiment, the brake sensor system 1646 may be used to operate the vehicle brakes in response to signals received from the brake actuator 1648 and / or brake sensors.
[0203] In at least one embodiment, controller 1636 may include, but is not limited to, one or more system-on-chips (“SoCs”). Figure 16A A controller 1636 (not shown) and / or a graphics processing unit (“GPU”) provides signals (e.g., representing commands) to one or more components and / or systems of vehicle 1600. For example, in at least one embodiment, controller 1636 may send signals to operate vehicle braking via brake actuator 1648, to operate steering system 1654 via one or more steering actuators 1656, and / or to operate propulsion system 1650 via one or more throttles / accelerators 1652. In at least one embodiment, one or more controllers 1636 may include one or more onboard (e.g., integrated) computing devices (e.g., supercomputers) that process sensor signals and output operating commands (e.g., signals representing commands) to enable autonomous driving and / or assist a driver in driving vehicle 1600. In at least one embodiment, one or more controllers 1636 may include a first controller for autonomous driving functions, a second controller for functional safety functions, a third controller for artificial intelligence functions (e.g., computer vision), a fourth controller for infotainment functions, a fifth controller for redundancy in emergency situations, and / or other controllers. In at least one embodiment, a single controller may handle two or more of the functions described above, and two or more controllers may handle a single function and / or any combination thereof.
[0204] In at least one embodiment, one or more controllers 1636 provide signals for controlling one or more components and / or systems of vehicle 1600 in response to sensor data received from one or more sensors (e.g., sensor inputs). In at least one embodiment, sensor data can be received from sensors, including but not limited to one or more Global Navigation Satellite System (“GNSS”) sensors 1658 (e.g., one or more Global Positioning System sensors), one or more RADAR sensors 1660, one or more ultrasonic sensors 1662, one or more LIDAR sensors 1664, one or more inertial measurement unit (IMU) sensors 1666 (e.g., one or more accelerometers, one or more gyroscopes, one or more magnetic compasses, one or more magnetometers, etc.), one or more microphones 1696, one or more stereo cameras 1668, one or more wide-angle cameras 1670 (e.g., fisheye cameras), one or more infrared cameras 1672, one or more surround cameras 1674 (e.g., 360-degree cameras), and remote cameras (…). Figure 16A (not shown in the image), medium-range camera ( Figure 16A (Not shown in the diagram) One or more speed sensors 1644 (e.g., for measuring the speed of vehicle 1600), one or more vibration sensors 1642, one or more steering sensors 1640, one or more brake sensors (e.g., as part of brake sensor system 1646) and / or other sensor types are received.
[0205] In at least one embodiment, one or more controllers 1636 may receive input (e.g., represented by input data) from the instrument panel 1632 of the vehicle 1600 and provide output (e.g., represented by output data, display data, etc.) via a human-machine interface (“HMI”) display 1634, a voice signaler, a speaker, and / or other components of the vehicle 1600. In at least one embodiment, the output may include information such as vehicle speed, velocity, time, map data (e.g., high-definition map). Figure 16A The HMI display 1634 may display information such as (not shown in the image), location data (e.g., the location of vehicle 1600, for example on a map), direction, the location of other vehicles (e.g., occupancy raster), information about objects, and the state of objects sensed by one or more controllers 1636. For example, in at least one embodiment, the HMI display 1634 may display information about the presence of one or more objects (e.g., road signs, warning signs, traffic light changes, etc.) and / or information about driving operations that have been, are being, or will be made (e.g., changing lanes now, exiting exit 34B within two miles, etc.).
[0206] In at least one embodiment, vehicle 1600 also includes a network interface 1624, which can communicate over one or more networks using one or more wireless antennas 1626 and / or one or more modems. For example, in at least one embodiment, network interface 1624 may be able to communicate via Long Term Evolution (“LTE”), Wideband Code Division Multiple Access (“WCDMA”), Universal Mobile Telecommunications System (“UMTS”), Global System for Mobile Communications (“GSM”), IMT-CDMA Multicarrier (“CDMA2000”), etc. In at least one embodiment, one or more wireless antennas 1626 may also enable communication between objects in the environment (e.g., vehicles, mobile devices) using one or more local area networks (e.g., Bluetooth, Bluetooth Low Energy (LE), Z-Wave, ZigBee, etc.) and / or one or more low-power wide area networks (hereinafter referred to as “LPWAN”) (e.g., LoRaWAN, SigFox, etc.).
[0207] Inference and / or training logic 715 is used to perform inference and / or training operations associated with one or more embodiments. In at least one embodiment, inference and / or training logic 715 can be in the system Figure 16A The operation is used to infer or predict the operation based at least in part on weight parameters calculated using neural network training operations, neural network functions and / or architectures or neural network use cases described herein.
[0208] Such components can be used to generate synthetic data that simulates failure scenarios during network training, which can help improve network performance while limiting the amount of synthetic data to avoid overfitting.
[0209] Figure 16B The illustration shows an embodiment according to at least one of the embodiments. Figure 16A Examples of camera positions and fields of view for an autonomous vehicle 1600. In at least one embodiment, the camera and its respective field of view are an example embodiment and are not intended to be limiting. For example, in at least one embodiment, additional and / or alternative cameras may be included and / or the cameras may be located at different positions on the vehicle 1600.
[0210] In at least one embodiment, the camera type used for the camera may include, but is not limited to, a digital camera suitable for use with components and / or systems of vehicle 1600. In at least one embodiment, one or more cameras may operate at Automotive Safety Integrity Level (“ASIL”) B and / or other ASILs. In at least one embodiment, the camera type may have any image capture rate, such as 60 frames per second (fps), 120 fps, 240 fps, etc. In at least one embodiment, the camera may be able to use a rolling shutter, a global shutter, another type of shutter, or a combination thereof. In at least one embodiment, the color filter array may include a red-to-clear (“RCCC”) color filter array, a red-to-clear-blue (“RCCB”) color filter array, a red-blue-green (“RBGC”) color filter array, a Foveon X3 color filter array, a Bayer sensor (“RGGB”) color filter array, a monochrome sensor color filter array, and / or other types of color filter arrays. In at least one embodiment, a transparent pixel camera, such as a camera with an array of RCCC, RCCB and / or RBGC color filters, may be used to improve photosensitivity.
[0211] In at least one embodiment, one or more cameras may be used to perform advanced driver assistance system (“ADAS”) functions (e.g., as part of a redundancy or fail-safe design). For example, in at least one embodiment, a multi-function mono camera may be installed to provide functions including lane departure warning, traffic sign assist, and intelligent headlight control. In at least one embodiment, one or more cameras (e.g., all cameras) may simultaneously record and provide image data (e.g., video).
[0212] In at least one embodiment, one or more cameras may be mounted in a mounting assembly, such as a custom-designed (3D-printed) assembly, to cut out stray light and reflections from within the vehicle (e.g., dashboard reflections reflected in the windshield mirror), which may interfere with the camera's image data capture capabilities. Regarding the rearview mirror mounting assembly, in at least one embodiment, the rearview mirror assembly may be 3D-printed custom-made such that the camera mounting plate conforms to the shape of the rearview mirror.
[0213] In at least one embodiment, one or more cameras may be integrated into the rearview mirror. For side-view cameras, one or more cameras may also be integrated into the four pillars at each corner of the cabin.
[0214] In at least one embodiment, a camera (e.g., a forward-facing camera) having a field of view including a portion of the environment in front of the vehicle 1600 can be used for surround view and, with the assistance of one or more controllers 1636 and / or control SoCs, to help identify the forward path and obstacles, thereby providing information crucial for generating an occupancy grid and / or determining a preferred vehicle path. In at least one embodiment, the forward-facing camera can be used to perform many of the same ADAS functions as LIDAR, including but not limited to emergency braking, pedestrian detection, and collision avoidance. In at least one embodiment, the forward-facing camera can also be used for ADAS functions and systems, including but not limited to lane departure warning (“LDW”), adaptive cruise control (“ACC”), and / or other functions (e.g., traffic sign recognition).
[0215] In at least one embodiment, various cameras can be used in a forward-facing configuration, including, for example, a monocular camera platform including a CMOS (“complementary metal-oxide-semiconductor”) color imager. In at least one embodiment, a wide-angle camera 1670 can be used to sense objects entering from the periphery (e.g., pedestrians, people crossing the street, or bicycles). Although in Figure 16B Only one wide-angle camera 1670 is shown; however, in other embodiments, the vehicle 1600 may have any number (including zero) of wide-angle cameras. In at least one embodiment, any number of remote cameras 1698 (e.g., a pair of remote stereo cameras) can be used for depth-based object detection, especially for objects for which a neural network has not yet been trained. In at least one embodiment, the remote camera 1698 can also be used for object detection and classification, as well as basic object tracking.
[0216] In at least one embodiment, any number of stereo cameras 1668 may also be included in a forward configuration. In at least one embodiment, one or more stereo cameras 1668 may include an integrated control unit comprising a scalable processing unit that may provide programmable logic (“FPGA”) and a multi-core microprocessor with a controller area network (“CAN”) or Ethernet interface integrated on a single chip. In at least one embodiment, such a unit may be used to generate a 3D map of the environment of the vehicle 1600, including distance estimates for all points in the image. In at least one embodiment, one or more stereo cameras 1668 may include, but are not limited to, a compact stereo vision sensor, which may include, but is not limited to, two camera lenses (one on the left and one on the right) and an image processing chip that can measure the distance from the vehicle 1600 to a target object and use the generated information (e.g., metadata) to activate autonomous emergency braking and lane departure warning functions. In at least one embodiment, other types of stereo cameras 1668 may also be used in addition to those described herein.
[0217] In at least one embodiment, a camera (e.g., a side-view camera) having a field of view including a portion of the environment on the side of the vehicle 1600 can be used for surround viewing, thereby providing information for creating and updating the occupied grid, and generating a side collision warning. For example, in at least one embodiment, a surround camera 1674 (e.g., as...) Figure 16B The four surround cameras shown can be positioned on vehicle 1600. In at least one embodiment, one or more surround cameras 1674 can include, but are not limited to, any number and combination of wide-angle cameras, one or more fisheye cameras, one or more 360-degree cameras, and / or similar cameras. For example, in at least one embodiment, four fisheye cameras can be located at the front, rear, and sides of vehicle 1600. In at least one embodiment, vehicle 1600 can use three surround cameras 1674 (e.g., left, right, and rear) and can utilize one or more other cameras (e.g., forward-facing cameras) as a fourth surround-view camera.
[0218] In at least one embodiment, a camera (e.g., a rear-view camera) having a field of view including a portion of the environment behind the vehicle 1600 can be used for parking assistance, surround view, rear collision warning, and creating and updating occupancy raster. In at least one embodiment, a wide variety of cameras can be used, including but not limited to cameras that are also suitable as one or more forward-facing cameras (e.g., long-range camera 1698 and / or one or more mid-range cameras 1676, one or more stereo cameras 1668, one or more infrared cameras 1672, etc.), as described herein.
[0219] Inference and / or training logic 715 is used to perform inference and / or training operations associated with one or more embodiments. Details regarding inference and / or training logic 715 are provided below. In at least one embodiment, inference and / or training logic 715 may be... Figure 16B Used in systems for reasoning or predicting operations based at least in part on weight parameters computed using neural network training operations, neural network functions and / or architectures, or neural network use cases described herein.
[0220] Such components can be used to generate synthetic data that simulates failure scenarios during network training, which can help improve network performance while limiting the amount of synthetic data to avoid overfitting.
[0221] Figure 16C The illustration shows an embodiment according to at least one of the embodiments. Figure 16A A block diagram of an example system architecture for an autonomous vehicle 1600. In at least one embodiment, Figure 16CEach of one or more components, one or more features, and one or more systems of vehicle 1600 is shown as connected via bus 1602. In at least one embodiment, bus 1602 may include, but is not limited to, a CAN data interface (which may alternatively be referred to herein as “CAN bus”). In at least one embodiment, the CAN bus may be a network within vehicle 1600 used to help control various features and functions of vehicle 1600, such as brake actuation, acceleration, braking, steering, windshield wipers, etc. In one embodiment, bus 1602 may be configured to have dozens or even hundreds of nodes, each node having its own unique identifier (e.g., CAN ID). In at least one embodiment, bus 1602 can be read to find steering wheel angle, ground speed, engine revolutions per minute (“RPM”), button positions, and / or other vehicle status indicators. In at least one embodiment, bus 1602 may be an ASIL B compliant CAN bus.
[0222] In at least one embodiment, FlexRay and / or Ethernet may be used in addition to or from CAN. In at least one embodiment, there may be any number of buses 1602, which may include, but are not limited to, zero or more CAN buses, zero or more FlexRay buses, zero or more Ethernet buses, and / or zero or more other types of buses using other protocols. In at least one embodiment, two or more buses may be used to perform different functions and / or may be used for redundancy. For example, a first bus may be used for a collision avoidance function, and a second bus may be used for actuation control. In at least one embodiment, each bus 1602 may communicate with any component of vehicle 1600, and two or more buses 1602 may communicate with the same component. In at least one embodiment, each of any number of system-on-chip (“SoC”) 1604, each of one or more controllers 1636, and / or each computer within the vehicle may access the same input data (e.g., input from sensors of vehicle 1600) and may be connected to a common bus, such as a CAN bus.
[0223] In at least one embodiment, vehicle 1600 may include one or more controllers 1636, such as those described herein. Figure 16A The controller 1636 can be used for a variety of functions. In at least one embodiment, the controller 1636 can be coupled to any of the various other components and systems of the vehicle 1600 and can be used to control the vehicle 1600, the artificial intelligence of the vehicle 1600, the infotainment and / or other functions of the vehicle 1600.
[0224] In at least one embodiment, vehicle 1600 may include any number of SoCs 1604. Each of the SoCs 1604 may include, but is not limited to, a central processing unit (“one or more CPUs”) 1606, a graphics processing unit (“one or more GPUs”) 1608, one or more processors 1610, one or more caches 1612, one or more accelerators 1614, one or more data storage 1616, and / or other components and features not shown. In at least one embodiment, one or more SoCs 1604 may be used to control vehicle 1600 on various platforms and systems. For example, in at least one embodiment, one or more SoCs 1604 may be combined with a high-definition (“HD”) map 1622 in a system (e.g., the system of vehicle 1600), the high-definition map 1622 being available from one or more servers via a network interface 1624. Figure 16C (Not shown in the image) Get map refresh and / or update.
[0225] In at least one embodiment, one or more CPUs 1606 may include CPU clusters or CPU complexes (which may alternatively be referred to herein as “CCPLEX”). In at least one embodiment, one or more CPUs 1606 may include multiple cores and / or a secondary (“L2”) cache. For example, in at least one embodiment, one or more CPUs 1606 may include eight cores in an intercoupled multiprocessor configuration. In at least one embodiment, one or more CPUs 1606 may include four dual-core clusters, each cluster having a dedicated L2 cache (e.g., 2MB L2 cache). In at least one embodiment, one or more CPUs 1606 (e.g., CCPLEX) may be configured to support simultaneous cluster operation, such that any combination of clusters of one or more CPUs 1606 can be active at any given time.
[0226] In at least one embodiment, one or more CPUs 1606 may implement power management functions, including but not limited to one or more of the following features: automatic clock gating of individual hardware modules to conserve dynamic power when idle; clock gating of each core when the core is not actively executing instructions due to executing Wait for Interrupt (“WFI”) / Event Wait (“WFE”) instructions; independent power supply for each core; independent clock gating for each core cluster when all cores are clock-gated or power-gated; and / or independent power gating for each core cluster when all cores are power-gated. In at least one embodiment, one or more CPUs 1606 may further implement an enhanced algorithm for managing power states, wherein allowed power states and expected wake-up times are specified, and the hardware / microcode determines the optimal power state for cores, clusters, and CCPLEX inputs. In at least one embodiment, the processing core may support a simplified power state input sequence in software, wherein the work is offloaded to the microcode.
[0227] In at least one embodiment, one or more GPUs 1608 may include integrated GPUs (or "iGPUs" herein). In at least one embodiment, one or more GPUs 1608 may be programmable and efficient for parallel workloads. In at least one embodiment, one or more GPUs 1608 may use an enhanced tensor instruction set. In at least one embodiment, one or more GPUs 1608 may include one or more streaming microprocessors, wherein each streaming microprocessor may include a Level 1 ("L1") cache (e.g., an L1 cache with at least 96KB of storage capacity), and two or more streaming microprocessors may share an L2 cache (e.g., an L2 cache with 512KB of storage capacity). In at least one embodiment, one or more GPUs 1608 may include at least eight streaming microprocessors. In at least one embodiment, one or more GPUs 1608 may use a computational application programming interface (API). In at least one embodiment, one or more GPUs 1608 may use one or more parallel computing platforms and / or programming models (e.g., NVIDIA's CUDA).
[0228] In at least one embodiment, one or more GPU 1608s may be power-optimized for optimal performance in automotive and embedded use cases. For example, in one embodiment, one or more GPU 1608s may be fabricated on a FinFET (“FinFET”). In at least one embodiment, each streaming microprocessor may include multiple mixed-precision processing cores divided into multiple blocks. For example, but not limited to, 64 PF32 cores and 32 PF64 cores may be divided into four processing blocks. In at least one embodiment, each processing block may be allocated 16 FP32 cores, 8 FP64 cores, 16 INT32 cores, two mixed-precision NVIDIA Tensor cores for deep learning matrix arithmetic, a level-zero (“L0”) instruction cache, a thread bundle scheduler, a dispatch unit, and / or a 64KB register file. In at least one embodiment, the streaming microprocessor may include independent parallel integer and floating-point data paths to provide efficient execution of workloads that mix computation and addressing operations. In at least one embodiment, the streaming microprocessor may include independent thread scheduling capabilities to enable finer-grained synchronization and cooperation between parallel threads. In at least one embodiment, the streaming microprocessor may include a combined L1 data cache and shared memory unit to improve performance while simplifying programming.
[0229] In at least one embodiment, one or more GPUs 1608 may include high-bandwidth memory (“HBM”) and / or a 16GB HBM2 memory subsystem to provide a peak storage bandwidth of approximately 900GB / s in some examples. In at least one embodiment, in addition to or instead of HBM memory, synchronous graphics random access memory (“SGRAM”) may be used, such as graphics double data rate type five synchronous random access memory (“GDDR5”).
[0230] In at least one embodiment, one or more GPUs 1608 may include unified memory technology. In at least one embodiment, address translation service (“ATS”) support can be used to allow one or more GPUs 1608 to directly access the page tables of one or more CPUs 1606. In at least one embodiment, when one or more GPUs 1608 memory management units (“MMUs”) experience a miss, an address translation request can be sent to one or more CPUs 1606. In response, in at least one embodiment, two CPUs of the one or more CPUs 1606 can look up the virtual-physical mapping of the address in their page tables and transfer the translation back to one or more GPUs 1608. In at least one embodiment, unified memory technology can allow a single unified virtual address space to be used for the memory of both one or more CPUs 1606 and one or more GPUs 1608, thereby simplifying the programming of one or more GPUs 1608 and the porting of applications to one or more GPUs 1608.
[0231] In at least one embodiment, one or more GPUs 1608 may include any number of access counters that can track the frequency of memory accesses by one or more GPUs 1608 to other processors. In at least one embodiment, one or more access counters can help ensure that memory pages are moved to the physical memory of the processor that accesses the pages most frequently, thereby improving the efficiency of shared memory ranges between processors.
[0232] In at least one embodiment, one or more SoCs 1604 may include any number of caches 1612, including those described herein. For example, in at least one embodiment, one or more caches 1612 may include a Level 3 (“L3”) cache available for one or more CPUs 1606 and one or more GPUs 1608 (e.g., connected to CPUs 1606 and GPUs 1608). In at least one embodiment, one or more caches 1612 may include a write-back cache that can, for example, track the state of a line using a cache coherence protocol (e.g., MEI, MESI, MSI, etc.). In at least one embodiment, although a smaller cache size may be used, the L3 cache may include 4 MB or more, depending on the embodiment.
[0233] In at least one embodiment, one or more SoCs 1604 may include one or more accelerators 1614 (e.g., hardware accelerators, software accelerators, or combinations thereof). In at least one embodiment, one or more SoCs 1604 may include a hardware acceleration cluster, which may include optimized hardware accelerators and / or large on-chip memory. In at least one embodiment, large on-chip memory (e.g., 4MB of SRAM) enables the hardware acceleration cluster to accelerate neural networks and other computations. In at least one embodiment, the hardware acceleration cluster may be used to supplement one or more GPUs 1608 and offload some tasks from one or more GPUs 1608 (e.g., freeing up more cycles from one or more GPUs 1608 to perform other tasks). In at least one embodiment, one or more accelerators 1614 may be used for a target workload (e.g., perceptual, convolutional neural network (“CNN”), recurrent neural network (“RNN”), etc.) that is sufficiently stable to withstand acceleration testing. In at least one embodiment, the CNN may include region-based or region convolutional neural networks (“RCNN”) and fast RCNN (e.g., for object detection) or other types of CNNs.
[0234] In at least one embodiment, one or more accelerators 1614 (e.g., a hardware acceleration cluster) may include one or more deep learning accelerators (“DLAs”). One or more DLAs may include, but are not limited to, one or more Tensor Processing Units (“TPUs”), which may be configured to provide an additional 10 trillion operations per second for deep learning applications and inference. In at least one embodiment, the TPU may be an accelerator configured and optimized for performing image processing functions (e.g., for CNNs, RCNNs, etc.). One or more DLAs may be further optimized for specific sets of neural network types and floating-point operations and inference. In at least one embodiment, one or more DLAs are designed to provide higher performance per millimeter than typical general-purpose GPUs and typically significantly outperform CPUs. In at least one embodiment, one or more TPUs may perform several functions, including single-instance convolution functions supporting, for example, INT8, INT16, and FP16 data types for features and weights, as well as post-processor functions. In at least one embodiment, one or more DLAs can execute neural networks, particularly CNNs, quickly and efficiently on processed or unprocessed data for any of the various functions, including, but not limited to: CNNs for object recognition and detection using data from camera sensors; CNNs for distance estimation using data from camera sensors; CNNs for emergency vehicle detection, recognition, and identification using data from microphone 1696; CNNs for face recognition and vehicle owner recognition using data from camera sensors; and / or CNNs for safety and / or safety-related events.
[0235] In at least one embodiment, the DLA can perform any function of one or more GPUs 1608, and by using inference accelerators, for example, the designer can target one or more DLAs or one or more GPUs 1608 for any function. For example, in at least one embodiment, the designer can concentrate the CNN processing and floating-point operations on one or more DLAs, leaving other functions to one or more GPUs 1608 and / or one or more accelerators 1614.
[0236] In at least one embodiment, one or more accelerators 1614 (e.g., hardware acceleration clusters) may include one or more programmable vision accelerators (“PVAs”), which may alternatively be referred to herein as computer vision accelerators. In at least one embodiment, one or more PVAs may be designed and configured to accelerate computer vision algorithms for advanced driver assistance systems (“ADAS”) 1638, autonomous driving, augmented reality (“AR”) applications, and / or virtual reality (“VR”) applications. One or more PVAs may strike a balance between performance and flexibility. For example, in at least one embodiment, each of one or more PVAs may include, for example, but not limited to, any number of reduced instruction set computer (“RISC”) cores, direct memory access (“DMA”), and / or any number of vector processors.
[0237] In at least one embodiment, the RISC core can interact with an image sensor (e.g., the image sensor of any camera described herein), an image signal processor, and / or other processors. In at least one embodiment, each RISC core may include any number of memories. In at least one embodiment, the RISC core may use any of a variety of protocols, depending on the embodiment. In at least one embodiment, the RISC core may execute a real-time operating system (“RTOS”). In at least one embodiment, the RISC core may be implemented using one or more integrated circuit devices, application-specific integrated circuits (“ASICs”), and / or memory devices. For example, in at least one embodiment, the RISC core may include an instruction cache and / or tightly coupled RAM.
[0238] In at least one embodiment, DMA enables components of (one or more) PVAs to access system memory independently of one or more CPUs 1606. In at least one embodiment, DMA can support any number of features for providing optimization to the PVA, including but not limited to, support for multidimensional addressing and / or circular addressing. In at least one embodiment, DMA can support up to six or more addressing dimensions, which may include, but are not limited to, block width, block height, block depth, horizontal block step, vertical block step, and / or depth step.
[0239] In at least one embodiment, the vector processor may be a programmable processor designed to efficiently and flexibly execute programming for computer vision algorithms and provide signal processing capabilities. In at least one embodiment, the PVA may include a PVA core and two vector processing subsystem partitions. In at least one embodiment, the PVA core may include a processor subsystem, a DMA engine (e.g., two DMA engines), and / or other peripherals. In at least one embodiment, the vector processing subsystem may serve as the main processing engine of the PVA and may include a vector processing unit (“VPU”), an instruction cache, and / or a vector memory (e.g., “VMEM”). In at least one embodiment, the VPU core may include a digital signal processor, such as a Single Instruction Multiple Data (“SIMD”) or Very Long Instruction Word (“VLIW”) digital signal processor. In at least one embodiment, the combination of SIMD and VLIW can improve throughput and speed.
[0240] In at least one embodiment, each vector processor may include an instruction cache and may be coupled to dedicated memory. As a result, in at least one embodiment, each vector processor may be configured to execute independently of other vector processors. In at least one embodiment, vector processors included in a particular PVA may be configured to employ data parallelism. For example, in at least one embodiment, multiple vector processors included in a single PVA may execute the same computer vision algorithm, except on different regions of an image. In at least one embodiment, vector processors included in a particular PVA may execute different computer vision algorithms simultaneously on the same image, or even execute different algorithms on a sequence of images or portions of images. In at least one embodiment, among others, any number of PVAs may be included in the hardware-accelerated cluster, and any number of vector processors may be included in each PVA. In at least one embodiment, one or more PVAs may include additional error-correcting code (“ECC”) memory to enhance overall system security.
[0241] In at least one embodiment, one or more accelerators 1614 (e.g., hardware acceleration clusters) may include an on-chip computer vision network and static random access memory (“SRAM”) for providing high-bandwidth, low-latency SRAM to one or more accelerators 1614. In at least one embodiment, the on-chip memory may include at least 4 MB of SRAM, comprising, for example, but not limited to, eight field-configurable memory blocks accessible to both the PVA and DLA. In at least one embodiment, each pair of memory blocks may include an Advanced Peripheral Bus (“APB”) interface, configuration circuitry, a controller, and a multiplexer. In at least one embodiment, any type of memory may be used. In at least one embodiment, the PVA and DLA may access the memory via a backbone providing high-speed access to the memory for both the PVA and DLA. In at least one embodiment, the backbone may include an on-chip computer vision network that interconnects the PVA and DLA to the memory (e.g., using an APB).
[0242] In at least one embodiment, the on-chip computer vision network may include an interface that determines that both the PVA and DLA provide ready and valid signals before transmitting any control signals / addresses / data. In at least one embodiment, the interface may provide separate phases and separate channels for transmitting control signals / addresses / data, as well as bursty communication for continuous data transmission. In at least one embodiment, although other standards and protocols may be used, the interface may conform to the International Organization for Standardization (“ISO”) 26262 or the International Electrotechnical Commission (“IEC”) 61508 standard.
[0243] In at least one embodiment, one or more SoCs 1604 may include a real-time eye-tracking hardware accelerator. In at least one embodiment, the real-time eye-tracking hardware accelerator may be used to quickly and efficiently determine the location and extent of an object (e.g., within a world model) to generate real-time visualization simulations for RADAR signal interpretation, for sound propagation synthesis and / or analysis, for simulation of SONAR systems, for general wave propagation simulation, for comparison with LIDAR data for localization and / or other functions, and / or for other purposes.
[0244] In at least one embodiment, one or more accelerators 1614 (e.g., hardware acceleration clusters) have broad applications for autonomous driving. In at least one embodiment, the PVA can be a programmable vision accelerator that can be used in critical processing stages in ADAS and autonomous vehicles. In at least one embodiment, the capabilities of the PVA at low power and low latency are well-matched to algorithmic domains requiring predictable processing. In other words, the PVA performs well in semi-intensive or intensive conventional computations, even on small datasets that may require predictable runtimes with low latency and low power consumption. In at least one embodiment, for autonomous vehicles, such as vehicle 1600, the PVA is designed to run classical computer vision algorithms because they are efficient in object detection and integer mathematical operations.
[0245] For example, according to at least one embodiment of the technology, PVA is used to perform computer stereo vision. In at least one embodiment, a semi-global matching-based algorithm may be used in some examples, although this is not intended to be limiting. In at least one embodiment, applications for Level 3-5 autonomous driving use dynamic estimation / stereo matching during operation (e.g., structure recovery from motion, pedestrian recognition, lane detection, etc.). In at least one embodiment, PVA can perform computer stereo vision functions on input from two monocular cameras.
[0246] In at least one embodiment, the PVA can be used to perform intensive optical flow. For example, in at least one embodiment, the PVA can process raw RADAR data (e.g., using 4D Fast Fourier Transform) to provide processed RADAR data. In at least one embodiment, the PVA is used for time-of-flight depth processing, for example, by processing raw time-of-flight data to provide processed time-of-flight data.
[0247] In at least one embodiment, the DLA can be used to run any type of network to enhance control and driving safety, including, but not limited to, neural networks whose output is used for a confidence score for each object detection. In at least one embodiment, the confidence score can be represented or interpreted as a probability, or as providing a relative “weight” for each detection relative to other detections. In at least one embodiment, the confidence score enables the system to make further decisions about which detections should be considered true positives rather than false positives. For example, in at least one embodiment, the system can set a threshold for the confidence score and only consider detections exceeding the threshold as true positives. In embodiments using an Automatic Emergency Braking (“AEB”) system, false positives would cause the vehicle to automatically perform emergency braking, which is obviously undesirable. In at least one embodiment, a highly confident detection can be considered a trigger for AEB. In at least one embodiment, the DLA can run a neural network for regressing the confidence score value. In at least one embodiment, the neural network may take at least a subset of parameters as its input, such as bounding box size, obtained ground plane estimate (e.g., from another subsystem), and outputs of one or more IMU sensors 1666 related to the vehicle 1600 orientation, distance, and 3D position estimate of the object obtained from the neural network and / or other sensors (e.g., one or more LiDAR sensors 1664 or one or more RADAR sensors 1660).
[0248] In at least one embodiment, one or more SoCs 1604 may include one or more data storage units 1616 (e.g., memory). In at least one embodiment, one or more data storage units 1616 may be on-chip memory of one or more SoCs 1604, which may store neural networks to be executed on one or more GPUs 1608 and / or DLAs. In at least one embodiment, one or more data storage units 1616 may have a sufficiently large capacity to store multiple instances of the neural network for redundancy and security. In at least one embodiment, one or more data storage units 1616 may include L2 or L3 caches.
[0249] In at least one embodiment, one or more SoCs 1604 may include any number of processors 1610 (e.g., embedded processors). In at least one embodiment, one or more processors 1610 may include a boot and power management processor, which may be a dedicated processor and subsystem for handling boot power and management functions, as well as associated security implementations. In at least one embodiment, the boot and power management processor may be part of a boot sequence of one or more SoCs 1604 and may provide runtime power management services. In at least one embodiment, the boot power and management processor may provide clock and voltage programming, assist system low-power state transitions, thermal and temperature sensor management of one or more SoCs 1604s, and / or power state management of one or more SoCs 1604s. In at least one embodiment, each temperature sensor may be implemented with its output frequency proportional to temperature, and one or more SoCs 1604s may use the ring oscillator to detect the temperature of one or more CPUs 1606s, one or more GPUs 1608s, and / or one or more accelerators 1614s. In at least one embodiment, if it is determined that the temperature exceeds a threshold, the startup and power management processor may enter a temperature fault routine and place one or more SoCs 1604s into a lower power state and / or place the vehicle 1600 into a safe stopping pattern for the driver (e.g., bring the vehicle 1600 to a safe stop).
[0250] In at least one embodiment, one or more processors 1610 may further include a set of embedded processors that can be used as an audio processing engine. In at least one embodiment, the audio processing engine may be an audio subsystem capable of providing full hardware support for multi-channel audio through multiple interfaces and a wide and flexible range of audio I / O interfaces. In at least one embodiment, the audio processing engine is a dedicated processor core with a digital signal processor with dedicated RAM.
[0251] In at least one embodiment, one or more processors 1610 may also include an always-on processor engine that can provide the necessary hardware features to support low-power sensor management and wake-up use cases. In at least one embodiment, the processor on the always-on processor engine may include, but is not limited to, a processor core, tightly coupled RAM, peripheral support (e.g., timers and interrupt controllers), various I / O controller peripherals, and routing logic.
[0252] In at least one embodiment, one or more processors 1610 may further include a secure clustering engine, which includes, but is not limited to, a dedicated processor subsystem for handling security management of automotive applications. In at least one embodiment, the secure clustering engine may include, but is not limited to, two or more processor cores, tightly coupled RAM, supporting peripheral devices (e.g., timers, interrupt controllers, etc.) and / or routing logic. In secure mode, in at least one embodiment, the two or more cores may operate in lockstep mode and may be used as a single core with comparison logic for detecting any differences between their operations. In at least one embodiment, one or more processors 1610 may further include a real-time camera engine, which may include, but is not limited to, a dedicated processor subsystem for handling real-time camera management. In at least one embodiment, one or more processors 1610 may further include a high dynamic range signal processor, which may include, but is not limited to, an image signal processor, which is a hardware engine as part of the camera processing pipeline.
[0253] In at least one embodiment, one or more processors 1610 may include a video image synthesizer, which may be a processing block (e.g., implemented on a microprocessor) that implements video post-processing functions required by the video playback application to produce the final image for the player window. In at least one embodiment, the video image synthesizer may perform lens distortion correction on one or more wide-angle cameras 1670, one or more surround cameras 1674, and / or one or more cabin monitoring camera sensors. In at least one embodiment, preferably, the cabin monitoring camera sensors are monitored by a neural network running on another instance of the SoC 1604, the neural network being configured to recognize cabin events and respond accordingly. In at least one embodiment, the cabin system may perform, but is not limited to, lip reading to activate cellular service and make phone calls, instruct emails, change the vehicle's destination, activate or change the vehicle's infotainment system and settings, or provide voice-activated web browsing. In at least one embodiment, certain functions are available to the driver when the vehicle is operating in autonomous mode, and are otherwise disabled.
[0254] In at least one embodiment, the video image synthesizer may include enhanced temporal denoising for simultaneous spatial and temporal denoising. For example, in at least one embodiment, when motion occurs in the video, denoising appropriately weights spatial information, thereby reducing the weight of information provided by adjacent frames. In at least one embodiment, when the image or a portion of the image does not contain motion, temporal denoising performed by the video image synthesizer may use information from previous images to reduce noise in the current image.
[0255] In at least one embodiment, the video image compositor can also be configured to perform stereoscopic correction on the input stereo lens frames. In at least one embodiment, when using an operating system desktop, the video image compositor can also be used for user interface compositing and does not require one or more GPUs 1608 to continuously render new surfaces. In at least one embodiment, when one or more GPUs 1608 are powered and actively performing 3D rendering, the video image compositor can be used to offload one or more GPUs 1608 to improve performance and responsiveness.
[0256] In at least one embodiment, one or more SoCs of SoC 1604 may also include a Mobile Industrial Processor Interface (“MIPI”) camera serial interface, a high-speed interface, and / or a video input block that can be used for receiving video and input from a camera and associated pixel input functions. In at least one embodiment, one or more SoCs of SoC 1604 may also include an input / output controller that can be software controlled and can be used to receive I / O signals not assigned to a specific role.
[0257] In at least one embodiment, one or more SoCs of SoC 1604 may also include extensive peripheral interfaces to enable communication with peripheral devices, audio encoders / decoders (“codecs”), power management and / or other devices. One or more SoCs of SoC 1604 may be used to process data from (e.g., connected via gigabit multimedia serial links and Ethernet channels) cameras, sensors (e.g., one or more LiDAR sensors 1664, one or more RADAR sensors 1660, etc., which may be connected via Ethernet), data from bus 1602 (e.g., vehicle 1600 speed, steering wheel position, etc.), data from one or more GNSS sensors 1658 (e.g., connected via Ethernet bus or CAN bus), etc. In at least one embodiment, one or more SoCs of SoC 1604 may also include a dedicated high-performance mass storage controller, which may include its own DMA engine and may be used to free one or more CPUs 1606 from routine data management tasks.
[0258] In at least one embodiment, one or more SoCs 1604 can be an end-to-end platform with a flexible architecture spanning automation levels 3-5, providing a comprehensive functional safety architecture that leverages and effectively utilizes computer vision and ADAS technologies to achieve diversity and redundancy. This provides a platform offering a flexible and reliable driving software stack as well as deep learning tools. In at least one embodiment, one or more SoCs 1604 can be faster, more reliable, and even more energy and space efficient than conventional systems. For example, in at least one embodiment, one or more accelerators 1614, when combined with one or more CPUs 1606, one or more GPUs 1608, and one or more data storage 1616, can provide a fast and efficient platform for Level 3-5 autonomous vehicles.
[0259] In at least one embodiment, the computer vision algorithm can be executed on a CPU, which can be configured using a high-level programming language (e.g., C) to execute multiple processing algorithms on a variety of visual data. However, in at least one embodiment, the CPU typically cannot meet the performance requirements of many computer vision applications, such as performance requirements related to execution time and power consumption. In at least one embodiment, many CPUs cannot execute complex object detection algorithms in real time, algorithms used in automotive ADAS applications and practical Level 3-5 autonomous vehicles.
[0260] The embodiments described herein allow multiple neural networks to be executed simultaneously and / or sequentially, and allow the results to be combined to achieve Level 3-5 autonomous driving capabilities. For example, in at least one embodiment, a CNN executed on a DLA or discrete GPU (e.g., one or more GPU 1620s) may include text and word recognition, thereby allowing a supercomputer to read and understand traffic signs, including signs for which the neural network has not yet been specifically trained. In at least one embodiment, the DLA may also include a neural network capable of recognizing, interpreting, and providing semantic understanding of symbols, and passing this semantic understanding to a path planning module running on a CPU Complex.
[0261] In at least one embodiment, for drives of levels 3, 4, or 5, multiple neural networks can run simultaneously. For example, in at least one embodiment, a warning sign consisting of a light bulb accompanied by the warning sign “Caution: flashing lights indicate icy conditions” can be interpreted independently or jointly by multiple neural networks. In at least one embodiment, the warning sign itself can be recognized as a traffic sign by a first deployed neural network (e.g., a trained neural network), and the text “flashing lights indicate icy conditions” can be interpreted by a second deployed neural network, which informs the vehicle’s path planning software (preferably executed on the CPU Complex) that icing conditions exist when flashing lights are detected. In at least one embodiment, flashing lights can be identified by operating a third deployed neural network across multiple frames, informing the vehicle’s path planning software of the presence (or absence) of flashing lights. In at least one embodiment, all three neural networks can run simultaneously, for example within the DLA and / or on one or more GPUs 1608.
[0262] In at least one embodiment, the CNN for facial recognition and vehicle owner identification can use data from camera sensors to identify the presence of an authorized driver and / or the owner of vehicle 1600. In at least one embodiment, a normally open sensor processor engine can be used to unlock the vehicle when the owner approaches the driver's door and turns on the lights, and, in security mode, can be used to disable the vehicle when the owner leaves it. In this way, one or more SoCs 1604 provide protection against theft and / or carjacking.
[0263] In at least one embodiment, the CNN for emergency vehicle detection and identification can use data from microphone 1696 to detect and identify emergency vehicle sirens. In at least one embodiment, one or more SoCs 1604 use the CNN to classify environmental and urban sounds, as well as visual data. In at least one embodiment, the CNN running on DLA is trained to identify the relative approach speed of emergency vehicles (e.g., by using the Doppler effect). In at least one embodiment, the CNN can also be trained to identify emergency vehicles in the area where the vehicle is operating, as identified by one or more GNSS sensors 1658. In at least one embodiment, when operating in Europe, the CNN will seek to detect European sirens, while in the United States the CNN will seek to identify only North American sirens. In at least one embodiment, once an emergency vehicle is detected, a control program can be used, with the assistance of one or more ultrasonic sensors 1662, to execute emergency vehicle safety routines, slow down the vehicle, pull the vehicle to the side of the road, stop, and / or leave the vehicle idle until (one or more) emergency vehicles have passed.
[0264] In at least one embodiment, vehicle 1600 may include one or more CPUs 1618 (e.g., one or more discrete CPUs or one or more dCPUs) that may be coupled to one or more SoCs 1604 via high-speed interconnects (e.g., PCIe). In at least one embodiment, one or more CPUs 1618 may include x86 processors. For example, one or more CPUs 1618 may be used to perform any of the various functions, such as arbitrating the results of potential inconsistencies between ADAS sensors and one or more SoCs 1604, and / or monitoring the status and health of one or more monitoring controllers 1636 and / or on-chip information systems (“information SoCs”) 1630.
[0265] In at least one embodiment, vehicle 1600 may include one or more GPUs 1620 (e.g., one or more discrete GPUs or one or more dGPUs) that may be coupled to one or more SoCs 1604 via a high-speed interconnect (e.g., NVIDIA's NVLINK). In at least one embodiment, one or more GPUs 1620 may provide additional artificial intelligence capabilities, such as by executing redundant and / or different neural networks, and may be used to train and / or update the neural networks based at least in part on inputs from sensors of vehicle 1600 (e.g., sensor data).
[0266] In at least one embodiment, vehicle 1600 may also include a network interface 1624, which may include, but is not limited to, one or more wireless antennas 1626 (e.g., one or more wireless antennas for different communication protocols, such as cellular antennas, Bluetooth antennas, etc.). In at least one embodiment, network interface 1624 may be used to enable wireless connectivity with other vehicles and / or computing devices (e.g., passenger client devices) via an internet cloud (e.g., employing servers and / or other network devices). In at least one embodiment, for communication with other vehicles, a direct link and / or an indirect link (e.g., via a network and the internet) may be established between vehicle 1600 and other vehicles. In at least one embodiment, a vehicle-to-vehicle communication link may be used to provide a direct link. The vehicle-to-vehicle communication link may provide vehicle 1600 with information about vehicles near vehicle 1600 (e.g., vehicles in front, to the side, and / or behind vehicle 1600). In at least one embodiment, the foregoing functionality may be part of a cooperative adaptive cruise control function of vehicle 1600.
[0267] In at least one embodiment, network interface 1624 may include a System-on-Chip (SoC) that provides modulation and demodulation functions and enables one or more controllers 1636 to communicate over a wireless network. In at least one embodiment, network interface 1624 may include a radio frequency (RF) front-end for up-conversion from baseband to RF and down-conversion from RF to baseband. In at least one embodiment, frequency conversion may be performed in any technically feasible manner. For example, frequency conversion may be performed using known processes and / or using a superheterodyne process. In at least one embodiment, the RF front-end functionality may be provided by a separate chip. In at least one embodiment, the network interface may include wireless functions for communication via LTE, WCDMA, UMTS, GSM, CDMA2000, Bluetooth, Bluetooth LE, Wi-Fi, Z-Wave, ZigBee, LoRaWAN, and / or other wireless protocols.
[0268] In at least one embodiment, vehicle 1600 may also include one or more data storage units 1628, which may include, but are not limited to, off-chip (e.g., one or more SoC 1604) memories. In at least one embodiment, one or more data storage units 1628 may include, but are not limited to, one or more storage elements, including RAM, SRAM, dynamic random access memory (“DRAM”), video random access memory (“VRAM”), flash memory, hard disk and / or other components and / or devices capable of storing at least one bit of data.
[0269] In at least one embodiment, the vehicle 1600 may also include one or more GNSS sensors 1658 (e.g., GPS and / or auxiliary GPS sensors) to assist in map creation, perception, occupancy raster generation, and / or path planning functions. In at least one embodiment, any number of GNSS sensors 1658 may be used, including, for example, but not limited to, GPS sensors connected to a serial interface (e.g., RS-232) bridge using a USB connector with Ethernet.
[0270] In at least one embodiment, vehicle 1600 may also include one or more RADAR sensors 1660. One or more RADAR sensors 1660 can be used by vehicle 1600 for remote vehicle detection, even in dark and / or inclement weather conditions. In at least one embodiment, the RADAR functional safety level may be ASIL B. One or more RADAR sensors 1660 may use a CAN bus and / or bus 1602 (e.g., to transmit data generated by one or more RADAR sensors 1660) for control and access to object tracking data, and in some examples, may access Ethernet to access raw data. In at least one embodiment, a wide variety of RADAR sensor types can be used. For example, but not limited to, one or more RADAR sensors 1660 may be suitable for front, rear, and side RADAR use. In at least one embodiment, one or more RADAR sensors 1660 are pulse Doppler RADAR sensors.
[0271] In at least one embodiment, one or more RADAR sensors 1660 may include different configurations, such as long-range with a narrow field of view, short-range with a wide field of view, short-range side coverage, etc. In at least one embodiment, the long-range RADAR can be used for adaptive cruise control functions. In at least one embodiment, the long-range RADAR system can provide a wide field of view achieved through two or more independent scans (e.g., within a 250m range). In at least one embodiment, one or more RADAR sensors 1660 can help distinguish between stationary and moving objects and can be used by the ADAS system 1638 for emergency braking assistance and forward collision warning. One or more sensors 1660 included in the long-range RADAR system may include, but are not limited to, a monostatic multimode RADAR with multiple (e.g., six or more) fixed RADAR antennas and high-speed CAN and FlexRay interfaces. In at least one embodiment, having six antennas, with the four central antennas, can create a focused beammap designed to record the surrounding environment of the vehicle 1600 at a high speed while minimizing traffic interference from adjacent lanes. In at least one embodiment, the other two antennas can expand the field of view, thereby enabling rapid detection of vehicles entering or leaving the lane 1600.
[0272] In at least one embodiment, as an example, a mid-range RADAR system may include, for example, a range of up to 160m (front) or 80m (rear), and a field of view of up to 42 degrees (front) or 150 degrees (rear). In at least one embodiment, a short-range RADAR system may include, but is not limited to, any number of RADAR sensors 1660 designed to be mounted at both ends of the rear bumper. When mounted at both ends of the rear bumper, in at least one embodiment, the RADAR sensor system may generate two beams that continuously monitor the blind spots at and near the rear of the vehicle. In at least one embodiment, the short-range RADAR system may be used in ADAS system 1638 for blind spot detection and / or lane change assistance.
[0273] In at least one embodiment, the vehicle 1600 may also include one or more ultrasonic sensors 1662. One or more ultrasonic sensors 1662, which may be positioned at the front, rear, and / or sides of the vehicle 1600, can be used for parking assistance and / or creating and updating occupancy detectors. In at least one embodiment, a wide variety of ultrasonic sensors 1662 can be used, and different ultrasonic sensors 1662 can be used for different detection ranges (e.g., 2.5m, 4m). In at least one embodiment, the ultrasonic sensors 1662 can operate at the ASIL B functional safety level.
[0274] In at least one embodiment, vehicle 1600 may include one or more LiDAR sensors 1664. The one or more LiDAR sensors 1664 may be used for object and pedestrian detection, emergency braking, collision avoidance, and / or other functions. In at least one embodiment, the one or more LiDAR sensors 1664 may be of functional safety level ASIL B. In at least one embodiment, vehicle 1600 may include multiple (e.g., two, four, six, etc.) LiDAR sensors 1664 that can use Ethernet (e.g., providing data to a Gigabit Ethernet switch).
[0275] In at least one embodiment, one or more LiDAR sensors 1664 may be able to provide a list of objects and their distances for a 360-degree field of view. In at least one embodiment, one or more commercially available LiDAR sensors 1664 may, for example, have an advertising range of approximately 100m, an accuracy of 2cm-3cm, and support a 100Mbps Ethernet connection. In at least one embodiment, one or more non-protruding LiDAR sensors may be used. In such embodiments, one or more LiDAR sensors 1664 may be implemented as small devices embedded in the front, rear, sides, and / or corners of a vehicle 1600. In at least one embodiment, one or more LiDAR sensors 1664, in such embodiments, can provide a horizontal field of view of up to 120 degrees and a vertical field of view of 35 degrees, even for objects with low reflectivity, and have a range of 200m.
[0276] In at least one embodiment, one or more forward-facing LIDAR sensors 1664 may be configured for a horizontal field of view between 45 degrees and 135 degrees.
[0277] In at least one embodiment, LIDAR technology (such as 3D flash LIDAR) may also be used. 3D flash LIDAR uses a laser flash as a transmission source to illuminate approximately 200m around the vehicle 1600. In at least one embodiment, the flash LIDAR unit includes, but is not limited to, a receiver that records the laser pulse propagation time and reflected light on each pixel, which in turn corresponds to the range from the vehicle 1600 to the object. In at least one embodiment, flash LIDAR can allow the generation of highly accurate and distortion-free images of the surrounding environment using each laser flash. In at least one embodiment, four flash LIDAR sensors may be deployed, one on each side of the vehicle 1600. In at least one embodiment, the 3D flash LIDAR system includes, but is not limited to, a solid-state 3D line-of-sight array LIDAR camera with no moving parts other than a fan (e.g., a non-scanning LIDAR device). In at least one embodiment, the flash LIDAR device can use a 5-nanosecond Class I (eye-safe) laser pulse per frame and can capture reflected laser light in the form of a 3D ranging point cloud and co-registered intensity data.
[0278] In at least one embodiment, vehicle 1600 may further include one or more IMU sensors 1666. In at least one embodiment, one or more IMU sensors 1666 may be located at the center of the rear axle of vehicle 1600. In at least one embodiment, one or more IMU sensors 1666 may include, for example, but not limited to, one or more accelerometers, one or more magnetometers, one or more gyroscopes, a magnetic compass, multiple magnetic compasses, and / or other sensor types. In at least one embodiment, for example in a six-axis application, one or more IMU sensors 1666 may include, but are not limited to, accelerometers and gyroscopes. In at least one embodiment, for example in a nine-axis application, one or more IMU sensors 1666 may include, but are not limited to, accelerometers, gyroscopes, and magnetometers.
[0279] In at least one embodiment, one or more IMU sensors 1666 can be implemented as a miniature, high-performance GPS-assisted inertial navigation system (“GPS / INS”) combining a microelectromechanical system (“MEMS”) inertial sensor, a high-sensitivity GPS receiver, and an advanced Kalman filtering algorithm to provide position, velocity, and attitude estimations; in at least one embodiment, one or more IMU sensors 1666 enable vehicle 1600 to estimate heading without input from a magnetic sensor obtained by directly observing and correlating velocity changes from GPS to one or more IMU sensors 1666. In at least one embodiment, one or more IMU sensors 1666 and one or more GNSS sensors 1658 can be combined in a single integrated unit.
[0280] In at least one embodiment, vehicle 1600 may include one or more microphones 1696 placed inside and / or around vehicle 1600. In at least one embodiment, in addition, one or more microphones 1696 may be used for emergency vehicle detection and identification.
[0281] In at least one embodiment, vehicle 1600 may also include any number of camera types, including one or more stereo cameras 1668, one or more wide-angle cameras 1670, one or more infrared cameras 1672, one or more surround cameras 1674, one or more long-range cameras 1698, one or more mid-range cameras 1676, and / or other camera types. In at least one embodiment, the cameras can be used to capture image data around the entire perimeter of vehicle 1600. In at least one embodiment, the type of camera used depends on vehicle 1600. In at least one embodiment, any combination of camera types can be used to provide the necessary coverage around vehicle 1600. In at least one embodiment, the number of cameras deployed may vary depending on the embodiment. For example, in at least one embodiment, vehicle 1600 may include six cameras, seven cameras, ten cameras, twelve cameras, or other numbers of cameras. The cameras may be examples, but are not limited to, supporting Gigabit Multimedia Serial Link (“GMSL”) and / or Gigabit Ethernet. In at least one embodiment, references previously made herein... Figure 16A and Figure 16B Each camera can be described in more detail.
[0282] In at least one embodiment, the vehicle 1600 may also include one or more vibration sensors 1642. In at least one embodiment, the one or more vibration sensors 1642 can measure vibrations of components of the vehicle 1600 (e.g., axles). For example, in at least one embodiment, changes in vibration can indicate changes in road surface conditions. In at least one embodiment, when two or more vibration sensors 1642 are used, differences between vibrations can be used to determine road surface friction or slippage (e.g., when there is a vibration difference between a power drive axle and a free-rotating axle).
[0283] In at least one embodiment, vehicle 1600 may include ADAS system 1638. ADAS system 1638 may include, but is not limited to, SoC. In at least one embodiment, ADAS system 1638 may include, but is not limited to, any number of autonomous / adaptive / automatic cruise control (“ACC”) systems, cooperative adaptive cruise control (“CACC”) systems, forward collision warning (“FCW”) systems, automatic emergency braking (“AEB”) systems, lane departure warning (“LDW”) systems, lane keeping assist (“LKA”) systems, blind spot warning (“BSW”) systems, rear cross traffic warning (“RCTW”) systems, collision warning (“CW”) systems, lane centering (“LC”) systems, and / or other systems, features, and / or functions, and combinations thereof.
[0284] In at least one embodiment, the ACC system may use one or more RADAR sensors 1660, one or more LIDAR sensors 1664, and / or any number of cameras. In at least one embodiment, the ACC system may include a longitudinal ACC system and / or a lateral ACC system. In at least one embodiment, the longitudinal ACC system monitors and controls the distance to vehicles adjacent to vehicle 1600 and automatically adjusts the speed of vehicle 1600 to maintain a safe distance from the vehicle ahead. In at least one embodiment, the lateral ACC system performs distance holding and suggests that vehicle 1600 change lanes when necessary. In at least one embodiment, lateral ACC is associated with other ADAS applications, such as LC and CW.
[0285] In at least one embodiment, the CACC system uses information from other vehicles, which may be received from other vehicles via network interface 1624 and / or one or more wireless antennas 1626 via a wireless link or indirectly via a network connection (e.g., via the Internet). In at least one embodiment, the direct link may be provided by a vehicle-to-vehicle (“V2V”) communication link, while the indirect link may be provided by an infrastructure-to-vehicle (“I2V”) communication link. Typically, the V2V communication concept provides information about the vehicle immediately preceding it (e.g., a vehicle immediately in front of vehicle 1600 and in the same lane as it), while the I2V communication concept provides information about traffic further ahead. In at least one embodiment, the CACC system may include one or both of the I2V and V2V information sources. In at least one embodiment, given information about vehicles preceding vehicle 1600, the CACC system can be more reliable and has the potential to improve traffic flow smoothness and reduce road congestion.
[0286] In at least one embodiment, the FCW system is designed to warn the driver of danger so that the driver can take corrective action. In at least one embodiment, the FCW system uses a forward-facing camera and / or one or more RADAR sensors 1660, coupled to a dedicated processor, DSP, FPGA, and / or ASIC, electrically coupled to driver feedback, such as a display, speaker, and / or vibration components. In at least one embodiment, the FCW system can provide warnings, for example, in the form of audible, visual warnings, vibrations, and / or rapid braking pulses.
[0287] In at least one embodiment, the AEB system detects an impending forward collision with another vehicle or other object and can automatically apply brakes if the driver does not take corrective action within a specified time or distance parameter. In at least one embodiment, the AEB system may use one or more forward-facing cameras and / or one or more RADAR sensors 1660 coupled to a dedicated processor, DSP, FPGA, and / or ASIC. In at least one embodiment, when the AEB system detects a hazard, it typically first warns the driver to take corrective action to avoid a collision, and if the driver does not take corrective action, the AEB system may automatically apply brakes to attempt to prevent or at least mitigate the effects of the predicted collision. In at least one embodiment, the AEB system may include techniques such as dynamic braking to support and / or brakes for impending collisions.
[0288] In at least one embodiment, when vehicle 1600 crosses lane markings, the LDW system provides visual, auditory, and / or tactile warnings, such as steering wheel or seat vibrations, to alert the driver. In at least one embodiment, the LDW system is inactive when the driver indicates intentional lane departure, such as by activating turn signals. In at least one embodiment, the LDW system may use a front-facing camera coupled to a dedicated processor, DSP, FPGA, and / or ASIC, which is electrically coupled to provide driver feedback such as a display, speaker, and / or vibration components. The LKA system is a variant of the LDW system. In at least one embodiment, if vehicle 1600 begins to leave the lane, the LKA system provides steering input or braking to correct vehicle 1600.
[0289] In at least one embodiment, the BSW system detects and warns the driver of a vehicle in the blind spot. In at least one embodiment, the BSW system can provide visual, auditory, and / or tactile alerts to indicate that merging or changing lanes is unsafe. In at least one embodiment, the BSW system can provide additional warnings when the driver uses the turn signal. In at least one embodiment, the BSW system can use one or more rear-facing cameras and / or one or more RADAR sensors 1660 coupled to a dedicated processor, DSP, FPGA, and / or ASIC, electrically coupled to driver feedback, such as a display, speaker, and / or vibration assembly.
[0290] In at least one embodiment, the RCTW system can provide visual, auditory, and / or tactile notifications when an object is detected outside the range of the rear camera while the vehicle 1600 is reversing. In at least one embodiment, the RCTW system includes an AEB system to ensure the application of the vehicle brakes to avoid a collision. In at least one embodiment, the RCTW system may use one or more rear-facing RADAR sensors 1660 coupled to a dedicated processor, DSP, FPGA, and / or ASIC, which are electrically coupled to driver feedback such as a display, speaker, and / or vibration assembly.
[0291] In at least one embodiment, conventional ADAS systems may be prone to generating false alarms, which can be annoying and distracting to the driver, but are generally not catastrophic because conventional ADAS systems warn the driver and allow the driver to determine whether a safe situation truly exists and take appropriate action. In at least one embodiment, in the event of conflicting results, the vehicle 1600 itself decides whether to follow the result of the primary computer or the secondary computer (e.g., the first controller 1636 or the second controller 1636). For example, in at least one embodiment, ADAS system 1638 may be a backup and / or auxiliary computer for providing perception information to a backup computer rationality module. In at least one embodiment, the backup computer rationality monitor may run redundant software on hardware components to detect faults in perception and dynamic driving tasks. In at least one embodiment, the output from ADAS system 1638 may be provided to a monitoring MCU. In at least one embodiment, if the outputs from the primary computer and the auxiliary computer conflict, the monitoring MCU decides how to reconcile the conflict to ensure safe operation.
[0292] In at least one embodiment, the master computer may be configured to provide a confidence score to the supervisory MCU to indicate the master computer's confidence in the selected result. In at least one embodiment, if the confidence score exceeds a threshold, the supervisory MCU may follow the master computer's instructions regardless of whether the auxiliary computer provides conflicting or inconsistent results. In at least one embodiment, if the confidence score does not meet the threshold, and if the master computer and the auxiliary computer indicate different results (e.g., conflicting), the supervisory MCU may arbitrate between the computers to determine the appropriate result.
[0293] In at least one embodiment, the supervisory MCU may be configured to run a neural network trained and configured to determine, at least in part, the conditions under which the auxiliary computer provides a false alarm based on outputs from both the host computer and the auxiliary computer. In at least one embodiment, the neural network in the supervisory MCU may learn when the output of the auxiliary computer can be trusted and when it cannot. For example, in at least one embodiment, when the auxiliary computer is a RADAR-based FCW system, the neural network in the supervisory MCU may learn when the FCW system recognizes a metallic object that is not actually dangerous, such as a drain grat or manhole cover that would trigger an alarm. In at least one embodiment, when the auxiliary computer is a camera-based LDW system, the neural network in the supervisory MCU may learn to override the LDW when a cyclist or pedestrian is present and lane departure is actually the safest operation. In at least one embodiment, the supervisory MCU may include at least one of a DLA or GPU suitable for running a neural network with associated memory. In at least one embodiment, the supervisory MCU may include and / or be included as a component of one or more SoC 1604s.
[0294] In at least one embodiment, the ADAS system 1638 may include an auxiliary computer that performs ADAS functions using conventional computer vision rules. In at least one embodiment, the auxiliary computer may use classic computer vision rules (if-then), and the presence of a neural network in the supervisory MCU can improve reliability, security, and performance. For example, in at least one embodiment, diverse implementations and intentional non-identity make the entire system more fault-tolerant, especially for failures caused by software (or software-hardware interface) functionality. For example, in at least one embodiment, if a software vulnerability or bug exists in the software running on the host computer, and different software code running on the auxiliary computer provides the same overall result, the supervisory MCU can more confidently assume that the overall result is correct and that the vulnerability in the software or hardware on the host computer will not lead to a significant error.
[0295] In at least one embodiment, the output of the ADAS system 1638 can be input to the perception module and / or the dynamic driving task module of the host computer. For example, in at least one embodiment, if the ADAS system 1638 indicates a forward collision warning due to an object directly ahead, the perception block can use this information when identifying the object. In at least one embodiment, as described herein, the assistance computer can have its own neural network trained to reduce the risk of false alarms.
[0296] In at least one embodiment, vehicle 1600 may also include an infotainment SoC 1630 (e.g., an in-vehicle infotainment system (IVI)). Although shown and described as an SoC, in at least one embodiment, the infotainment system SoC 1630 may not be an SoC and may include, but is not limited to, two or more discrete components. In at least one embodiment, the infotainment SoC 1630 may include, but is not limited to, a combination of hardware and software that can be used to provide audio (e.g., music, personal digital assistant, navigation instructions, news, radio, etc.), video (e.g., television, movies, streaming media, etc.), telephone (e.g., hands-free calling), network connectivity (e.g., LTE, WiFi, etc.) and / or information services (e.g., navigation system, rear parking assist, radio data system, vehicle-related information such as fuel level, total coverage distance, brake fuel level, fuel level, door opening / closing, air filter information, etc.) to vehicle 1600. For example, the infotainment SoC 1630 may include a radio, disk player, navigation system, video player, USB and Bluetooth connectivity, automobile, in-vehicle entertainment system, WiFi, steering wheel audio controls, hands-free voice control, head-up display (“HUD”), HMI display 1634, telematics device, control panel (e.g., for controlling and / or interacting with various components, features and / or systems) and / or other components. In at least one embodiment, the infotainment SoC 1630 may further be used to provide information (e.g., visual and / or auditory) to a user of vehicle 1600, such as information from ADAS system 1638, autonomous driving information (such as planned vehicle maneuvers), trajectory, surrounding environment information (e.g., intersection information, vehicle information, road information, etc.) and / or other information.
[0297] In at least one embodiment, the infotainment SoC 1630 may include any number and type of GPU functionality. In at least one embodiment, the infotainment SoC 1630 may communicate with other devices, systems, and / or components of the vehicle 1600 via a bus 1602 (e.g., CAN bus, Ethernet, etc.). In at least one embodiment, the infotainment SoC 1630 may be coupled to a monitoring MCU, enabling the GPU of the infotainment system to perform some autonomous driving functions in the event of a failure of the main controller 1636 (e.g., the main computer and / or backup computer of the vehicle 1600). In at least one embodiment, the infotainment SoC 1630 may cause the vehicle 1600 to enter a driver-to-safe-stop mode, as described herein.
[0298] In at least one embodiment, vehicle 1600 may also include instrument panel 1632 (e.g., digital instrument panel, electronic instrument panel, digital instrument control panel, etc.). In at least one embodiment, instrument panel 1632 may include, but is not limited to, controllers and / or supercomputers (e.g., discrete controllers or supercomputers). In at least one embodiment, instrument panel 1632 may include, but is not limited to, any number and combination of a set of instruments, such as speedometer, fuel level, oil pressure, tachometer, odometer, turn indicator, shift position indicator, one or more seatbelt warning lights, one or more parking brake warning lights, one or more engine malfunction lights, auxiliary restraint system (e.g., airbag) information, lighting controls, safety system controls, navigation information, etc. In some examples, information may be displayed and / or shared between infotainment SoC 1630 and instrument panel 1632. In at least one embodiment, instrument panel 1632 may be included as part of infotainment SoC 1630, or vice versa.
[0299] Inference and / or training logic 715 is used to perform inference and / or training operations associated with one or more embodiments. Details regarding inference and / or training logic 715 are provided below. In at least one embodiment, inference and / or training logic 715 can be implemented in the system. Figure 16C The operation is used to infer or predict the operation based at least in part on weight parameters calculated using neural network training operations, neural network functions and / or architectures or neural network use cases described herein.
[0300] Such components can be used to generate synthetic data that simulates failure scenarios during network training, which can help improve network performance while limiting the amount of synthetic data to avoid overfitting.
[0301] Figure 16D It is based on at least one embodiment in a cloud-based server and Figure 16AA diagram of a system 1676 for communication between autonomous vehicles 1600. In at least one embodiment, system 1676 may include, but is not limited to, one or more servers 1678, one or more networks 1690, and any number and type of vehicles, including vehicle 1600. In at least one embodiment, one or more servers 1678 may include, but is not limited to, multiple GPUs 1684(A)-1684(H) (collectively referred to herein as GPU 1684), PCIe switches 1682(A)-1682(D) (collectively referred to herein as PCIe switch 1682), and / or CPUs 1680(A)-1680(B) (collectively referred to herein as CPU 1680). GPU 1684, CPU 1680, and PCIe switch 1682 may be interconnected with high-speed interconnects, such as, but not limited to, NVLink interface 1688 developed by NVIDIA and / or PCIe connection 1686. The GPU 1684 is connected via NVLink and / or NVSwitchSoC, and the GPU 1684 and PCIe switch 1682 are connected via PCIe interconnect. In at least one embodiment, although eight GPUs 1684, two CPUs 1680, and four PCIe switches 1682 are shown, this is not intended to be limiting. In at least one embodiment, each of one or more servers 1678 may include, but is not limited to, any combination of any number of GPUs 1684, CPUs 1680, and / or PCIe switches 1682. For example, in at least one embodiment, one or more servers 1678 may each include eight, sixteen, thirty-two, and / or more GPUs 1684.
[0302] In at least one embodiment, one or more servers 1678 may receive image data representing an image from a vehicle via one or more networks 1690, the image showing unexpected or changed road conditions, such as recently commenced roadworks. In at least one embodiment, one or more servers 1678 may transmit an updated neural network 1692 and / or map information 1694, including but not limited to information about traffic and road conditions, to the vehicle via one or more networks 1690. In at least one embodiment, updating the map information 1694 may include, but is not limited to, updating the HD map 1622, such as information about construction sites, potholes, sidewalks, floods, and / or other obstacles. In at least one embodiment, the neural network 1692, the updated neural network 1692, and / or the map information 1694 may be generated from new training and / or experience represented by data received from any number of vehicles in the environment, and / or at least based on training performed in a data center (e.g., using one or more servers 1678 and / or other servers).
[0303] In at least one embodiment, one or more servers 1678 can be used to train a machine learning model (e.g., a neural network) at least in part based on training data. In at least one embodiment, the training data can be generated by the vehicle, and / or can be generated in a simulation (e.g., using a game engine). In at least one embodiment, any amount of training data is labeled (e.g., where the associated neural network benefits from supervised learning) and / or undergoes other preprocessing. In at least one embodiment, no amount of training data is labeled and / or preprocessed (e.g., where the associated neural network does not require supervised learning). In at least one embodiment, once the machine learning model is trained, the machine learning model can be used by the vehicle (e.g., transmitted to the vehicle via one or more networks 1690), and / or the machine learning model can be used by one or more servers 1678 to remotely monitor the vehicle.
[0304] In at least one embodiment, one or more servers 1678 may receive data from the vehicle and apply the data to a state-of-the-art real-time neural network for real-time intelligent inference. In at least one embodiment, one or more servers 1678 may include a deep learning supercomputer and / or a dedicated AI computer powered by one or more GPUs 1684, such as the DGX and DGX Station machines developed by NVIDIA. However, in at least one embodiment, one or more servers 1678 may include a deep learning infrastructure in a data center using CPU power.
[0305] In at least one embodiment, the deep learning infrastructure of one or more servers 1678 may be capable of fast, real-time inference and can use this capability to assess and verify the health of the processor, software, and / or associated hardware in vehicle 1600. For example, in at least one embodiment, the deep learning infrastructure may receive periodic updates from vehicle 1600, such as image sequences and / or objects located by vehicle 1600 in the image sequence (e.g., via computer vision and / or other machine learning object classification techniques). In at least one embodiment, the deep learning infrastructure may run its own neural network to identify objects and compare them with objects identified by vehicle 1600, and if the results do not match and the deep learning infrastructure determines that the AI in vehicle 1600 is malfunctioning, one or more servers 1678 may signal to vehicle 1600 to instruct the fail-safe computer of vehicle 1600 to take control, notify passengers, and complete a safe stopping operation.
[0306] In at least one embodiment, one or more servers 1678 may include one or more GPUs 1684 and one or more programmable inference accelerators (e.g., NVIDIA's TensorRT 3 devices). In at least one embodiment, the combination of GPU-driven servers and inference acceleration can enable real-time response. In at least one embodiment, for example, where performance is less critical, servers driven by CPUs, FPGAs, and other processors may be used for inference. In at least one embodiment, inference and / or training logic 715 is used to execute one or more embodiments. Details regarding the inference and / or training logic 715 are provided elsewhere herein.
[0307] Other variations are within the spirit of this disclosure. Therefore, although the disclosed technology is readily adaptable to various modifications and alternative constructions, certain embodiments thereof are illustrated in the accompanying drawings and have been described in detail above. However, it should be understood that the disclosure is not intended to be limited to one or more specific forms disclosed, but rather, it is intended to cover all modifications, alternative constructions, and equivalents falling within the spirit and scope of this disclosure as defined in the appended claims.
[0308] Unless otherwise stated or obviously contradicted by the context, the terms “a,” “an,” and “the,” and similar references, used in the context of describing the disclosed embodiments (particularly in the context of the appended claims), should be interpreted as encompassing both singular and plural forms, rather than as definitions of the terms. Unless otherwise stated, the terms “comprising,” “having,” “including,” and “containing” should be interpreted as open-ended terms (meaning “including, but not limited to”). The term “connection” (referring to a physical connection where not modified) should be interpreted as partially or wholly contained, attached to, or joined together, even with some intervention. Unless otherwise indicated herein, references to numerical ranges herein are intended only as a way of abbreviating each individual value falling within that range, and each individual value is incorporated into the specification as if it were separately described herein. Unless otherwise indicated or contradicted by the context, the use of the terms “set” (e.g., “item set”) or “subset” should be interpreted as a non-empty set comprising one or more members. Furthermore, unless otherwise indicated or contradicted by the context, the term “subset” of the corresponding set does not necessarily mean an appropriate subset of the corresponding set, but rather that the subset and the corresponding set can be equal.
[0309] Unless otherwise explicitly stated or clearly contradicted by the context, connective phrases such as “at least one of A, B, and C” or “at least one of A, B, and C” are understood in the context to generally refer to items, terms, etc., which can be A or B or C, or any non-empty subset of the set A, B, and C. For example, in an illustrative example of a set with three members, the connective phrases “at least one of A, B, and C” and “at least one of A, B, and C” refer to any of the following sets: {A}, {B}, {C}, {A, B}, {A, C}, {B, C}, {A, B, C}. Therefore, such connective language is generally not intended to imply that some embodiments require the presence of at least one of A, at least one of B, and at least one of C. Additionally, unless otherwise stated or contradicted by the context, the term “multiple” indicates a plural state (e.g., “multiple items” means multiple items). Multiple means at least two items, but more may be indicated if explicitly stated or by the context. Furthermore, unless otherwise stated or clearly understood from the context, the phrase “based on” means “at least partially based on” rather than “based on only”.
[0310] Unless otherwise indicated herein or clearly contradicted by the context, the operations of the processes described herein may be performed in any suitable order. In at least one embodiment, processes such as those described herein (or variations thereof and / or combinations thereof) are executed under the control of one or more computer systems configured with executable instructions and are implemented as code (e.g., executable instructions, one or more computer programs, or one or more applications) that is executed jointly on one or more processors via hardware or a combination thereof. In at least one embodiment, the code is stored on a computer-readable storage medium, for example, in the form of a computer program comprising a plurality of instructions executable by one or more processors. In at least one embodiment, the computer-readable storage medium is a non-transitory computer-readable storage medium that excludes transient signals (e.g., propagating transient electrical or electromagnetic transmissions) but includes non-transitory data storage circuitry (e.g., buffers, caches, and queues). In at least one embodiment, code (e.g., executable code or source code) is stored on one or more non-transitory computer-readable storage media (or other memory for storing executable instructions) on which executable instructions are stored, which, when executed by one or more processors of a computer system (i.e., as a result of execution), cause the computer system to perform the operations described herein. In at least one embodiment, the set of non-transitory computer-readable storage media comprises multiple non-transitory computer-readable storage media, and one or more of the individual non-transitory storage media lack all the code, but the multiple non-transitory computer-readable storage media collectively store all the code. In at least one embodiment, the executable instructions are executed such that different instructions are executed by different processors; for example, the non-transitory computer-readable storage media store the instructions, and the main central processing unit (“CPU”) executes some instructions while the graphics processing unit (“GPU”) executes other instructions. In at least one embodiment, different components of the computer system have separate processors, and the different processors execute different subsets of the instructions.
[0311] Therefore, in at least one embodiment, the computer system is configured to implement one or more services that perform the operations of the processes described herein, either individually or collectively, and such a computer system is configured with suitable hardware and / or software to enable the implementation of the operations. Furthermore, the computer system implementing at least one embodiment of this disclosure is a single device, and in another embodiment it is a distributed computer system comprising multiple devices operating in different ways, such that the distributed computer system performs the operations described herein, and that a single device does not perform all the operations.
[0312] The use of any and all examples or exemplary language (e.g., “such as”) provided herein is intended only to better illustrate embodiments of this disclosure and does not constitute a limitation on the scope of the disclosure unless otherwise required. No language in the specification should be construed as indicating that any unclaimed element is essential to the practice of the disclosure.
[0313] All references cited in this article, including publications, patent applications and patents, are incorporated herein by reference as if each reference were individually and specifically indicated to be incorporated herein by reference and the entire contents of which are described herein.
[0314] The terms “coupled” and “connected”, and their derivatives, may be used in the specification and claims. It should be understood that these terms may not be intended to be synonyms with each other. Rather, in certain examples, “connected” or “coupled” may be used to indicate that two or more elements are in direct or indirect physical or electrical contact with each other. “Coupled” may also mean that two or more elements are not in direct contact with each other, but still cooperate or interact with each other.
[0315] Unless otherwise expressly stated, it will be understood that throughout this specification, terms such as “processing,” “computing,” “determining,” etc., refer to the actions and / or processes of a computer or computing system or similar electronic computing device that process and / or convert data represented as physical quantities (e.g., electrons) in the registers and / or memory of the computing system into other data represented as physical quantities in the memory, registers, or other such information storage, transmission, or display devices of the computing system.
[0316] In a similar manner, the term "processor" can refer to any device or part of memory that processes electronic data from registers and / or memory and converts that electronic data into other electronic data that can be stored in registers and / or memory. As a non-limiting example, a "processor" can be a CPU or a GPU. A "computing platform" can include one or more processors. As used herein, a "software" process can include, for example, software and / or hardware entities that perform work over time, such as tasks, threads, and intelligent agents. Similarly, each process can refer to multiple processes that execute instructions sequentially or intermittently, sequentially, or in parallel. The terms "system" and "method" are used interchangeably herein, provided that a system can embody one or more methods, and a method can be considered a system.
[0317] This document refers to the process of acquiring, obtaining, receiving, or inputting analog or digital data into a subsystem, computer system, or computer-implemented machine. Acquiring, obtaining, receiving, or inputting analog and digital data can be accomplished in various ways, such as by receiving data as a parameter to a function call or an application programming interface (API) call. In some implementations, the process of acquiring, obtaining, receiving, or inputting analog or digital data can be accomplished by transmitting data via a serial or parallel interface. In another implementation, the process of acquiring, obtaining, receiving, or inputting analog or digital data can be accomplished by transmitting data from a providing entity to an acquiring entity via a computer network. Reference can also be made to providing, outputting, transmitting, sending, or presenting analog or digital data. In various examples, the process of providing, outputting, transmitting, sending, or presenting analog or digital data can be implemented by transmitting data as an input or output parameter to a function call, an API call, or an inter-process communication mechanism.
[0318] While the discussion above illustrates example implementations of the described technologies, other architectures can be used to implement the described functionality and are intended to fall within the scope of this disclosure. Furthermore, although specific assignments of responsibilities have been defined above for discussion purposes, various functions and responsibilities can be assigned and divided in different ways depending on the circumstances.
[0319] Furthermore, although the subject matter has been described in language specific to structural features and / or methodological actions, it should be understood that the subject matter claimed in the appended claims is not necessarily limited to the specific features or actions described. Rather, specific features and actions are disclosed as exemplary forms for implementing the claims.
Claims
1. A method comprising: Identify a first set of images containing multiple objects across multiple categories; The first image set is provided as input to a first machine learning model, which is trained to detect the presence of one or more objects of at least one of the plurality of categories depicted in the given input image for a given input image, and to predict at least mask data including one or more pixels indicating one or more of the detected objects depicted in the given input image. Object data associated with each of the first image set is determined from one or more first outputs of the first machine learning model, wherein the object data for each corresponding image in the first image set includes mask data indicating one or more pixels in each corresponding image that depict each object detected in the corresponding image; A second machine learning model is trained to detect objects of a target category in a second image set, wherein the second machine learning model is trained using at least a subset of the first image set and a target output for the at least subset of the first image set, wherein the target output includes mask data associated with each object detected in the at least subset of the first image set, and an indication of whether the category associated with each object detected in the at least subset of the first image set corresponds to the target category, wherein the trained second machine learning model includes a plurality of model heads, the plurality of model heads including at least one model head associated with predicting bounding boxes associated with objects detected in a given image and at least one model head associated with predicting mask data associated with objects detected in a given image; and After training the second machine learning model, the second machine learning model is updated to remove at least one model head associated with the mask data that predicts objects detected in a given image.
2. The method of claim 1, wherein the first machine learning model is further trained to predict, for each of one or more detected objects, a specific category among the plurality of categories associated with the corresponding detected object.
3. The method according to claim 2, further comprising: Generating the target output, wherein generating the target output includes: Determine whether the specific category associated with the corresponding detected object corresponds to the target category.
4. The method according to claim 1, further comprising: Use indications of one or more bounding boxes associated with the image to identify truth data associated with the corresponding object depicted in the image.
5. The method of claim 4, wherein at least one of the one or more bounding boxes is provided by at least one of the users of an accredited bounding box authority or platform.
6. The method according to claim 1, further comprising: A third set of images is provided as input to the second machine learning model; Obtain one or more second outputs from the second machine learning model; as well as Based on the one or more second outputs, additional object data associated with each of the third image set is determined, wherein the additional object data for each corresponding image in the second image set includes a region in the corresponding image that includes an object detected in the corresponding image and an indication of the category associated with the detected object.
7. The method according to claim 1, further comprising: The updated second machine learning model is sent over the network to at least one of the edge device or endpoint device.
8. A system comprising: Memory devices; as well as A processing device coupled to the memory device, wherein the processing device is configured to perform operations including the following: Generating training data for a machine learning model, wherein generating the training data includes: Generate training input including images depicting the objects; and Generate a target output for the training input, wherein the target output includes a bounding box associated with the depicted object, mask data including an indication of one or more pixels in the image depicting the object, and an indication of a category associated with the depicted object, wherein the bounding box, the mask data, and the indication of a category associated with the depicted object are obtained based on the outputs of one or more additional machine learning models; The training data is provided to train the machine learning model on (i) the set of training inputs including generated training inputs and (ii) the set of target outputs including generated target outputs, wherein the trained machine learning model includes a plurality of model heads, the plurality of model heads including at least one model head associated with predicting bounding boxes associated with objects detected in a given image and at least one model head associated with predicting mask data associated with detected objects; and; The trained machine learning model is updated to remove at least one model head associated with the mask data that predicts the object being detected.
9. The system of claim 8, wherein the operation further comprises: Provide a set of images as input to an updated, trained machine learning model; Obtain one or more outputs of the updated trained machine learning model; as well as Object data associated with each of the image sets is determined from one or more outputs, wherein the object data for each corresponding image in the image set includes a region in the corresponding image that includes an object detected in the corresponding image and an indication of a category associated with the detected object.
10. The system of claim 8, wherein the operation further comprises: The updated, trained machine learning model is deployed using at least one of an edge device or an endpoint device.
11. The system of claim 8, wherein generating the target output for the training input comprises: The image depicting the object is provided as input to the additional machine learning model, wherein the additional machine learning model is trained to detect the presence of one or more objects depicted in the given input image for a given input image, and to predict at least mask data associated with one or more of the detected objects; as well as Object data associated with the image is determined from one or more outputs of the additional machine learning model, wherein the object data of the image includes mask data associated with the depicted object.
12. The system of claim 11, wherein the additional machine learning model is further trained to predict a category associated with each of one or more detected objects, and wherein the object data of the image further includes the indication of the category associated with the depicted object.
13. The system of claim 8, wherein generating the target output for the training input comprises: Obtain truth data associated with the image, wherein the truth data includes the bounding box associated with the depicted object.
14. The system of claim 13, wherein the truth data is obtained from a database, the database including indications of one or more bounding boxes associated with objects depicted in an image set, wherein the images are included in the image set, and wherein the one or more bounding boxes are provided by a user of an accredited bounding box authority or platform.
15. A non-transitory computer-readable storage medium comprising instructions that, when executed by a processing device, cause the processing device to perform operations including: The current image set is provided as input to a first machine learning model, wherein the first machine learning model is trained to use... (i) including the training input of the training image set and (ii) Detecting objects of a target category in a given set of images for the training input, wherein for each corresponding training image in the training image set, the target output includes ground truth data associated with each object depicted in the corresponding training image, mask data including an indication of one or more pixels in the corresponding training image depicting each detected object, and an indication of whether the category associated with each object depicted in the corresponding training image corresponds to the target category, wherein the ground truth data indicates a region in the corresponding training image that includes the corresponding object. Furthermore, after training the first machine learning model, the mask head is removed from the multiple model heads of the trained first machine learning model; Obtain one or more outputs of the first machine learning model; and Based on the one or more outputs of the first machine learning model, object data associated with each of the current image sets is determined, wherein the object data for each corresponding current image in the current image set includes an indication of a region in the corresponding current image that includes an object detected in the corresponding current image, and an indication of whether the detected object corresponds to the target category.
16. The non-transitory computer-readable storage medium of claim 15, wherein the object data further includes mask data associated with the object detected in the corresponding current image.
17. The non-transitory computer-readable storage medium of claim 15, wherein determining the object data associated with each of the current image set comprises: Extract one or more object data sets from the one or more outputs of the first machine learning model, wherein each of the one or more object data sets is associated with a confidence level corresponding to the object data and an object detected in the corresponding current image; and Determine whether the confidence level associated with the corresponding object data set meets the confidence level criteria.
18. The non-transitory computer-readable storage medium of claim 15, further comprising training the first machine learning model by: The training image set is provided as input to a second machine learning model, wherein the second machine learning model is trained to detect one or more objects of at least one of a plurality of categories depicted in the given input image for a given input image, and for each of the one or more detected objects, predict at least mask data associated with the corresponding detected object; and From one or more outputs of the second machine learning model, determine object data associated with each of the training image sets, wherein the object data for each corresponding training image in the training image set includes mask data associated with each object detected in the corresponding image.
19. The non-transitory computer-readable storage medium of claim 15, wherein the truth data is obtained using a database, the database including indications of one or more bounding boxes associated with the training image set, wherein each of the one or more bounding boxes is provided by at least one of the users of an accredited bounding box authority or platform.
Citation Information
Patent Citations
Model generation method, target detection method, device, electronic equipment and medium
CN112257815A
Instance segmentation inferred from machine-learning model output
CN112334906A