Constrained device positioning

Through vehicle-mounted cameras and image processing technology, the wearing status of seat belts in the vehicle is analyzed, and the problem of difficulty in accurately identifying and correcting incorrect wear in the prior art is solved, and real-time and accurate monitoring of the wearing status of seat belts for passengers is achieved.

CN114684062BActive Publication Date: 2025-06-27NVIDIA CORP
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202111628623.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-12-29
Filing Date
2021-12-28
Publication Date
2025-06-27
Estimated Expiration
2041-12-28

AI Technical Summary

Technical Problem

The prior art has challenges in detecting and ensuring that vehicle occupants wear seat belts correctly, especially in the difficulty of effectively identifying and correcting incorrect wear.

Method used

By capturing images using an onboard camera, combined with a local predictor and a global assembler, pixels in the image are analyzed to identify the position and shape of the seat belt, modeling using advanced polynomial curves to determine whether the seat belt is properly worn.

Benefits of technology

Real-time monitoring and accurate identification of the wear status of vehicle occupants' seat belts is realized, traffic safety is improved, and physical modification of the vehicle's existing system is not required.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114684062B_ABST
    Figure CN114684062B_ABST
Patent Text Reader

Abstract

Disclosed is restraint device positioning, and specifically disclosed are systems and methods related to the positioning of restraint devices (e.g., seat belts). In one embodiment, the present disclosure relates to systems and methods for seat belt detection and modeling. A vehicle may be occupied by one or more occupants wearing one or more seat belts. A camera or other sensor is placed within the vehicle to capture images of the one or more occupants. The system analyzes the images to detect and model the seat belts depicted in the images. Specifically, the system may scan the images and image regions that may correspond to seat belts. The system may aggregate candidate image regions that may correspond to seat belts and refine the candidate regions based on various constraints. The system may build a model based on the refined candidate regions indicating seat belts. The system may use the images to visualize the model indicating seat belts.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross - Reference to Related Applications

[0002] For all purposes, this application incorporates by reference the entire disclosures of the following U.S. patent applications: co-pending U.S. patent application No. 17 / 005,914, filed on August 28, 2020, entitled "NEURAL NETWORK BASED DETERMINATION OF GAZE DIRECTION USING SPATIAL MODELS"; co-pending U.S. patent application No. 17 / 004,252, filed on August 27, 2020, entitled "NEURAL NETWORK BASED FACIAL ANALYSIS USING FACIAL LANDMARKS AND ASSOCIATED CONFIDENCE VALUES"; and co-pending U.S. patent application No. 16 / 905,418, filed on June 18, 2020, entitled "MACHINE LEARNING-BASED SEATBELT DETECTION AND USAGE RECOGNITION USING FIDUCIAL MARKING". BACKGROUND OF THE DISCLOSURE

[0003] Automobiles and other vehicles and machines that transport or house passengers or operators typically have various safety features, such as seat belts or other restraint devices. In many cases, if not used correctly, the safety devices can be less effective or even ineffective. For example, an incorrectly worn seat belt can be significantly less effective than a correctly worn seat belt. Various attempts have been made to improve the use of such safety devices. Some devices involve retrofitting machines, which can be difficult and expensive, especially considering the variations between machines of the same type. Generally, detecting the correct use of restraint devices, such as seat belts in vehicles, has many challenges and various degrees of success in terms of accuracy. TECHNICAL FIELD

[0004] At least one embodiment relates to processing resources for identifying and modeling one or more restraint devices from one or more images. For example, at least one embodiment relates to a processor or computing system for identifying and modeling one or more restraint devices from one or more images of the one or more restraint devices according to the various novel techniques described herein. BRIEF DESCRIPTION OF THE DRAWINGS

[0005] The system and method for constraining device positioning according to the present disclosure will be described in detail below with reference to the accompanying drawings, where:

[0006] Figure 1 An example of a system for seat belt positioning according to some embodiments of the present disclosure is shown;

[0007] Figure 2 An example of a seat belt pitch for one or more pixels having one or more pixel neighbors in one or more target directions according to some embodiments of the present disclosure is shown;

[0008] Figure 3 An example of a non-seat belt pitch for one or more pixels having one or more pixel neighbors in one or more target directions according to some embodiments of the present disclosure is shown;

[0009] Figure 4 A seat belt surface pitch according to some embodiments of the present disclosure is shown

[0010] Figure 5 An example of the positioning of a structured edge according to some embodiments of the present disclosure is shown;

[0011] Figure 6 An example of a candidate map according to some embodiments of the present disclosure is shown;

[0012] Figure 7 Examples of multiple seat belt positionings according to some embodiments of the present disclosure are shown;

[0013] Figure 8 An example of a seat belt positioning according to some embodiments of the present disclosure is shown;

[0014] Figure 9 An example of a seat belt positioning according to some embodiments of the present disclosure is shown;

[0015] Figure 10 An example of a seat belt positioning according to some embodiments of the present disclosure is shown;

[0016] Figure 11 An example of a seat belt positioning according to some embodiments of the present disclosure is shown;

[0017] Figure 12 A flowchart showing a method for generating a model for an input image according to some embodiments of the present disclosure;

[0018] Figure 13A An illustration of an example autonomous vehicle according to some embodiments of the present disclosure;

[0019] Figure 13B Examples of camera positions and fields of view of exemplary autonomous vehicles according to some embodiments of the present disclosure Figure 13A ;

[0020] Figure 13C Block diagram of an exemplary system architecture of an exemplary autonomous vehicle according to some embodiments of the present disclosure Figure 13A ;

[0021] Figure 13D System diagram for communication between one or more cloud - based servers and an exemplary autonomous vehicle according to some embodiments of the present disclosure; and Figure 13A ;

[0022] Figure 14 Block diagram of an exemplary computing device suitable for implementing some embodiments of the present disclosure DETAILED DESCRIPTION

[0023] The techniques described herein provide ways to passively detect seat belts and other restraint devices for different functions, such as detecting whether an occupant in a vehicle is wearing a seat belt and wearing it correctly. Cameras disposed in the vehicle capture images and analyze the images to determine whether the occupant is wearing the seat belt correctly. Typically, a vehicle's system uses a sensing system that detects whether the seat belt is in a locked position. However, there are situations where the sensing system can be deceived or permanently disabled. In addition, there are situations where a passenger has fastened the seat belt but is wearing it incorrectly (e.g., the seat belt is worn behind the back). To counter this and improve traffic safety, the techniques described herein use video / images captured by cameras mounted on the vehicle to determine whether the seat belt is being worn correctly. This technique is performed without having to physically modify the seat belt itself or any existing locking sensing system

[0024] More specifically, the system first analyzes the captured image to classify whether a pixel is part of the seatbelt or part of the background. This classification results in generating multiple pixel candidates that are most likely part of the seatbelt. To more accurately identify the pixels that are part of the seatbelt, the system uses multiple pieces of information to correct misidentifications of pixels identified as part of the seatbelt and vice versa. Such multiple pieces of information can include: parameters / constraints such as how the seatbelt is constructed inside the vehicle, the direction in which the seatbelt should extend when locked, the physical properties of the seatbelt, and the configuration of the camera that captured the image to assemble a more accurate set of pixels. After filtering the pixels to generate a more accurate set of pixels, the system then parameterizes the seatbelt and models the shape of the seatbelt using a high-order polynomial curve that further removes any pixels that are outliers with respect to the modeled shape and retrieves back pixels that were previously incorrectly filtered out or occluded. The final seatbelt curve can be used to enhance passenger safety, such as determining whether the seatbelt is being worn correctly (e.g., the seatbelt is in the locked position and is worn diagonally across the passenger's chest).

[0025] The techniques described herein extend methods for improving safety features in all types of vehicles without having to modify existing vehicle systems. The techniques are capable of identifying which occupants in a vehicle are correctly wearing their seatbelts based on video / images captured by the vehicle's in-vehicle systems and / or cameras. The system can identify not only whether the driver is correctly wearing their seatbelt, but also can determine whether all other occupants in the vehicle are correctly wearing their seatbelts. Although the system described herein is applied to seatbelts in passenger vehicles, the system can also be applied to other areas that require seatbelts, harnesses, or other straps (e.g., construction equipment, amusement park rides, 4D cinema seats). The system can be a component of an in-vehicle occupant monitoring system (OMS).

[0026] The system can be robust and applicable to different types of camera sensor configurations, such as color or infrared cameras, conventional field of view or fisheye cameras. The system can operate on and process images captured under different lighting conditions, such as low lighting conditions (e.g., night-time conditions), variable lighting conditions (e.g., day-time conditions), and / or variants thereof. The system can determine novel local parallel line patterns and detect such patterns at a micro-scale within an image through one or more techniques described herein. The system can utilize parallel computing techniques using one or more general-purpose graphics processing units (GPGPUs) to efficiently determine all seat belts for all frames of a video simultaneously. The system can provide real-time monitoring of seat belts and seat belt usage to improve the safety of one or more occupants of a vehicle. The system can provide a binary classification function for any patch of an image and can utilize different algorithms that can remove noise or false positive seat belt part candidates. The system can utilize a seat belt shape modeling function based on higher-order curves, which can even locate seat belts using occlusions or other obstacles. The system can locate seat belts corresponding to a subject and can be applicable to subjects with any suitable body appearance, including body appearances with different sizes, shapes, and attributes (such as clothing type and hair type) and / or other various body appearance attributes.

[0027] Referring Figure 1 , Figure 1 is an example of a system for seat belt localization according to some embodiments of the present disclosure. It should be understood that such and other arrangements described herein are set forth only as examples. Other arrangements and elements (e.g., machines, interfaces, functions, orders, function groupings, etc.) may be used in addition to or instead of those shown, and some elements may be omitted altogether. Further, many of the elements described herein can be implemented as discrete or distributed components or as functional entities combined with other components and implemented in any suitable combination and location. The various functions described herein as being performed by an entity can be performed by hardware, firmware, and / or software. For example, the various functions can be performed by a processor executing instructions stored in a memory.

[0028] Figure 1Example 100 of a system for seatbelt positioning according to at least one embodiment is shown. In various embodiments, seatbelt positioning refers to one or more processes of identifying the position, size, shape, orientation, and / or other characteristics of one or more seatbelts depicted in one or more images (e.g., an image of a person in the driver's or passenger's seat of a vehicle). It should be noted that although Example 100 depicts seatbelt positioning, any suitable portion of a restraint device of any suitable system can be positioned. The restraint device can include devices such as seatbelts (e.g., two-point seatbelts, three-point seatbelts, four-point seatbelts, etc.), safety harnesses, or other vehicle safety devices.

[0029] A system for seatbelt positioning (which may be referred to as a seatbelt locator) can include a local predictor 104, a global assembler 108, and a shape modeler 116. Various processes of the system for seatbelt positioning can be performed by one or more graphics processing units (GPUs) (such as parallel processing units (PPUs)). One or more processes of the system for seatbelt positioning can be performed by any suitable processing system or unit (e.g., GPU, PPU, central processing unit (CPU)) and in any suitable manner (including sequentially, in parallel, and / or their variants). The local predictor 104, the global assembler 108, and the shape modeler 116 can be software modules of one or more computer systems loaded on a vehicle. In some examples, the local predictor 104, the global assembler 108, and the shape modeler 116 are software programs executed on a computer server accessible via one or more networks, where the input image 102 is provided to the computer server from one or more computing systems of the vehicle via one or more networks, and the results of seatbelt positioning of one or more seatbelts of the input image 102 are provided back to the vehicle via one or more networks. The input to the system for seatbelt positioning can include the input image 102. In an embodiment, the input image 102 is an image of an entity in a vehicle having a restraint device such as a seatbelt. See Figure 1 , the input image 102 can be an image of a person wearing a seatbelt sitting in the driver's seat of a vehicle. The system for seatbelt positioning can determine the position and orientation of the seatbelt of the input image 102, and can further determine whether the seatbelt of the input image 102 is correctly applied.

[0030] The input image 102 can be an image captured from one or more image capture devices such as a camera or other device. The input image 102 can be a frame of a video captured from one or more cameras. The input image 102 can be captured from one or more cameras placed inside the vehicle. In some embodiments, the input image 102 is captured from one or more cameras outside the vehicle, such as through a surveillance system, a mobile phone, a drone, a handheld imaging device, or other imaging systems separated from the vehicle, and can capture images of the vehicle and the vehicle's occupants. In some examples, the input image is data captured from an image capture device and is further processed by adjusting one or more color properties, upsampling or downsampling, cropping, and / or otherwise processing the data. The vehicle can be a vehicle such as an autonomous vehicle, a semi-autonomous vehicle, a manual vehicle, and / or variants thereof. In some examples, the vehicle is an amusement ride vehicle (e.g., a roller skater), a construction vehicle, or other vehicle that requires one or more restraint devices. The input image 102 can be captured from one or more cameras (such as those cameras described in conjunction with Figures 13A to 13D the cameras described). The input image 102 can be captured from a camera in the vehicle that faces the vehicle occupants. The input image 102 can be a color image, a grayscale image, a black / white image, etc. The input image 102 can be captured from one or more cameras with different FOVs. The input image 102 can be an image with a minimum color contrast, which can result in a low contrast of the seatbelt and the background of the input image 102. The input image 102 can depict a seatbelt that may be occluded by one or more entities such as hands, arms, clothing, or other objects. The input image 102 can depict a seatbelt that can be moved or extended to various positions. The input image 102 can depict a seatbelt that can be positioned in a straight line or other irregular forms such as a curve. The input image 102 may be distorted due to object or vehicle movement. The input image 102 can depict a subject (e.g., a vehicle driver or passenger) and a restraint device (e.g., a seatbelt) applied to the subject. In some examples, the input image 102 depicts a subject and a restraint device (e.g., the input image 102 depicts a subject in a built environment with a restraint device such as a safety harness) in an environment such as a built environment.

[0031] The input image 102 can be received by or otherwise obtained by the local predictor 104. In some examples, the input image 102 is processed by one or more preprocessing operations to condition the input image 102 for processing; the preprocessing operations can include operations such as feature / contrast enhancement, denoising, morphological operations, resizing / rescaling, and / or variations thereof. The local predictor 104 can be a collection of one or more hardware and / or software computing resources having executable instructions that, when executed, scan the input image and predict regions of the input image that are part of a seat belt. The input image 102 can include a collection of pixels, where the seat belt depicted in the input image 102 can occupy more than one pixel. For any given pixel, the local predictor 104 can use the pixel and information from its neighboring pixels to predict whether the pixel belongs to the seat belt or is a background pixel. Pixels determined to be part of the seat belt can be identified as part of the foreground of the input image 102, and pixels determined not to be part of the seat belt can be identified as part of the background of the input image 102. The local predictor 104 can perform an initial classification of regions of the input image 102 to determine regions of the input image 102 that depict or otherwise represent a restraint device (e.g., a seat belt). The regions can include pixels, collections of pixels, etc. The regions can include groups of pixels that are in close proximity to each other. The regions can include contiguous groups or areas of pixels of the image.

[0032] The local predictor 104 can generate true results (which can be indicated by a value of 1) or false results (which can be indicated by a value of 0) representing a seat belt or not a seat belt, respectively. In various embodiments, the input image 102 includes noise, and the classification results can initially be correct or not correct. In an embodiment, there are four possible cases, but there can also be additional cases: (1) true positive, where the pixel is part of the seat belt and the local predictor 104 returns 1, (2) false positive, where the pixel is not part of the seat belt and the local predictor returns 1, (3) false negative, where the pixel is part of the seat belt and the local predictor 104 returns 0, and (4) true negative, where the pixel is not part of the seat belt and the local predictor 104 returns 0. The local predictor 104 can prioritize high recall over high precision (e.g., the local predictor 104 can reduce false negatives, although it may increase false positives).

[0033] The local predictor 104 can utilize nearby pixel neighbors for additional information. A pixel of the input image 102 located at the y-th row and the x-th column of the input image 102 can be represented as p(x, y) or any variation thereof. The dimensions of the input image 102 can be represented as a width of W, a height of H, and K channels (e.g., for a color image K = 3 (or 4), and for a grayscale image K = 1). p(x, y) and its neighboring pixels can form a set S and can be represented by the following equation, although any variation thereof can be utilized:

[0034] S = {p ij | f dist (p ij , p) ≤ ε, 0 < i ≤ W, 0 < j ≤ H} (1)

[0035] where f dist can be a pixel distance function (e.g., L1 distance), and ε can be a threshold that determines the boundary of the neighborhood.

[0036] In some examples, candidates for the neighborhood shape (also referred to as a patch) include circles, rectangles, squares, and / or their variants. In an embodiment, the local predictor 104 utilizes a square with a candidate pixel located at its center as the neighborhood shape. The length of the square can be represented by L = 2k + 1, where k = 1, 2, 3..., and where the set S can be simplified by the following equation, although any variation thereof can be utilized:

[0037] S = {p ij | 0 ≤ |i - y| ≤ k, 0 ≤ |j - x| ≤ k} (2)

[0038] The local predictor 104 can perform isotropic checks along various directions. The local predictor 104 can output a binary prediction result vector which can be represented by the following equation, although any variation thereof can be utilized:

[0039]

[0040] where D can be the total number of directions for prediction, θ0, θ1, … θ D-1 can be angles uniformly spaced within the range of [0, π), and can be defined by the following equation, although any variation thereof can be utilized:

[0041]

[0042] can be a criteria function represented by f criteria (x, y, θ i ) for any pixel p(x, y) within the input image 102 and a specified direction θi Determine whether the tile image generated using P and its neighborhood S in this direction is a seat belt component. The local predictor (which can be represented by f predictor (x,y) means) can be defined as is a scaling function of , and can be represented by the following equation, although any variation thereof may be used:

[0043]

[0044] in There may be weights assigned to the different directions.The weights may be determined from a statistical analysis of the distribution of the seat belt shapes when worn by one or more entities. Figure 2 and Figure 3 Examples of seatbelt tiles and non-seatbelt tiles for one or more pixels having one or more pixel neighbors in one or more target directions (eg, tile directions) are shown in accordance with at least one embodiment.

[0045] See also Figure 2 , example 200 may include (a) a sample seatbelt tile depicted in the upper left, (b) an enhanced seatbelt tile depicted in the upper right, (c) a seatbelt tile projected curve depicted in the lower left, and (d) a smoothed seatbelt curve depicted in the lower right. Figure 2 , (c) the seat belt tile projection curve may correspond to (b) the enhanced seat belt tile, where the x-axis of (c) may correspond to the x-values ​​or horizontal values ​​of (b), and the y-axis of (c) may correspond to the values ​​of the pixel intensities of the vertical pixel columns of (b). The pixel intensity of a particular pixel column may be the sum of all pixel intensities (also referred to as pixel intensity levels) of the pixels of the pixel column. The pixel intensity or pixel intensity level of a particular pixel may correspond to the brightness of the pixel (e.g., a pixel with a high intensity may appear white and a pixel with a low intensity may appear black). See Figure 2 , (d) the smoothed seat belt curve can be a smoothed version of (c) the seat belt patch projection curve. For example, see Figure 2 , the two peaks of (c) the seat belt patch projection curve and (d) the smoothed seat belt curve correspond to the white vertical lines of (b) the enhanced seat belt patch, which may correspond to the boundaries of the seat belt.

[0046] See also Figure 3 , example 300 may include (e) a non-seatbelt tile sample depicted at the upper left, (f) an enhanced non-seatbelt tile depicted at the upper right, (g) a non-seatbelt tile projected curve depicted at the lower left, and (h) a smoothed non-seatbelt curve depicted at the lower right. Figure 3, (g) The non - seat - belt tile projection curve can correspond to the (f) enhanced non - seat - belt tile, where the x - axis of (g) can correspond to the x - value or the horizontal value of (f), and the y - axis of (g) can correspond to the value of the pixel intensity of the vertical pixel column of (f). The pixel intensity of a specific pixel column can be the sum of the pixel intensities of all the pixels in the pixel column. See Figure 3 , (h) The smoothed non - seat - belt curve can be a smoothed version of the (f) enhanced non - seat - belt tile.

[0047] It should be noted that Figure 2 and Figure 3 depict examples of potential seat - belt tiles / curves and non - seat - belt tiles / curves, and the seat - belt tiles / curves and / or non - seat - belt tiles / curves can be any variation thereof. The seat - belt tile sample can correspond to any tile or region of an image (e.g., input image 102) that can depict one or more seat - belts, and the corresponding seat - belt tile curve can have any suitable shape that is at least partially based on the seat - belt tile sample.

[0048] Various information can be determined from Figure 2 and Figure 3 In one example, structured edges are observed because the seat - belt edges can generate two generally parallel lines in the tile direction based on the scale of the tile (e.g., if the tile is small enough), where, based on the camera - seat - belt relative geometry, the structured edges can be parameterized in terms of the edge - to - edge distance range, and its variations.

[0049] The intensity and / or saturation range can also be observed because the seat - belt can have different colors, such as black, gray, tan, etc., where although the pixel values can change with different environments (e.g., lighting), the seat - belt pixel values can have a certain intensity range (e.g., seat - belt pixels are rarely pure white).

[0050] In an embodiment, surface smoothness is also observed because the seat - belt can have a similar texture along the belt, where the pixels within each seat - belt region can have limited diversity in terms of smoothness.

[0051] In an embodiment, f structure , f intensity and f smoothnes are represented as binary functions, where:

[0052] f criteria (x, y, θ i ) = f structure (x, y, θ i ) ∩ f intensity (x, y, θ i ) ∩ f smoothness (x, y, θ i) (6)

[0054] Structural criterion f structure It can be characterized in combination with a task, and the task can be defined as how to represent parallel edges of a seat belt boundary within the neighborhood of a pixel. The local predictor 104 can identify and locate candidate seat belt edges and use the candidate seat belt edges to distinguish seat belt pixels from non-seat belt pixels (e.g., noise).

[0055] The seat belt tiles can have one or more directions, and for each tile, the local predictor 104 can determine whether there is a seat belt boundary along the tile direction. The tile directions can be evenly spaced into D categories within the range [0, π) (e.g., as depicted in Equation (3)), where regardless of how the seat belt is oriented, its orientation can be classified into one of these categories. For a given tile direction θ i , the local predictor 104 can check whether there is a seat belt edge parallel to θ i . The local predictor 104 can determine seat belt tiles in any suitable direction and the boundaries of the seat belt along the direction of the seat belt.

[0056] Figure 4 Example 400 showing the seat belt tile geometry according to at least one embodiment is shown. The seat belt tile geometry can refer to one or more geometries or geometric properties of a particular seat belt tile, such as tile direction, angle, orientation, size, boundary, etc. It should be noted that Figure 4 Examples of potential seat belt tile geometries are depicted, and the seat belt tile geometry can be any variant thereof. The seat belt tile can correspond to any tile or region of an image (e.g., input image 102) that can depict one or more seat belts, and the corresponding seat belt tile geometry can have any suitable geometry that is at least partially based on the shape or geometry of the seat belt tile. See Figure 4 , for each pixel (i, j) in the tile generated around the pixel p(x, y), the value from the pixel (x’, y’) can be retrieved from the image (e.g., input image 102) using the following equation, although any variant thereof can be utilized:

[0057]

[0058] where i = 1, 2, …, L and j = 1, 2, …, L.

[0059] The local predictor 104 can retrieve the seatbelt tiles and / or non-seatbelt tiles using one or more equations, such as equation (7) described above. The local predictor 104 can determine whether a tile has a structured edge pair. The edges in a tile can be one or more series of aligned pixel rows having a sudden pixel intensity change at the same position within each row. The change can be an increase or decrease in pixel intensity. The index of the sudden change can be the position where the line can be located. In some examples, in the case of two parallel lines, the tile has a structured edge.

[0060] The local predictor 104 can utilize a two-dimensional (2D) gradient operation to capture the magnitude of the sudden intensity change and sum along the tile direction to distinguish the seatbelt edge from background noise. patch (x,y,θ,j) can be defined as the resulting curve after projection and can be represented by the following equation, although any variant thereof can be utilized:

[0061]

[0062] where j can be a variable index tile column, i can be a row index, and f(x,y,θ,i,j) can be a tile intensity function. Figure 2 At (c) and Figure 3 At (g) an example of the curve can be depicted.

[0063] The curve can be affected by different perturbations, such as illumination changes and other noises. The local predictor 104 can apply an algorithm such as the Savtzsky-Golay algorithm to smooth the curve for further processing. patch (x,y,θ,j) can be defined as the filtered curve and can be represented by the following equation, although any variant thereof can be utilized:

[0064]

[0065] where t can be the number of convolution coefficients and C k can be the respective coefficients. Figure 2 At (d) and Figure 3 At (h) an example of the filtered curve can be depicted.

[0066] The seatbelt tiles having structured edges can have a filtered curve pattern similar to those depicted in Figure 5 , while the non-seatbelt tiles can have random curves. The local predictor 104 can extract different features from the curves to identify the seatbelt pattern, such as the distance d edges between the edges and the peak difference d peaks . In an embodiment, structure (x,y,θ i) is defined by the following equations, but any variant thereof can be utilized:

[0067] The local predictor 104 can identify two peaks of any given curve (e.g., peak left and peak right ) and their respective indices (e.g., idx left and idx right ) to determine d edges and d peaks .

[0068] Figure 5 Example 500 shows the positioning of structured edges according to at least one embodiment. Example 500 depicts determining the edges of a seatbelt from seatbelt tile curves (such as those depicted in Figure 2 and Figure 3 ). It should be noted that Figure 5 depicts an example of the potential positioning of structured edges from seatbelt tile curves, and the positioning of structured edges from seatbelt tile curves can be any variant thereof. The seatbelt tile curve can correspond to any suitable tile or region of an image (e.g., input image 102) that can depict one or more seatbelts, and the corresponding positioning of structured edges from the seatbelt tile curve can be at least partially based on the seatbelt tile curve indicating any suitable edge or feature.

[0069] See Figure 5 , the local predictor 104 may or may not select the two maximum points within a particular curve because the selection may be inaccurate since the two points may be near the highest peak. See Figure 5 , if the curve is cut into two parts, the two peaks can be separated, and the two peaks can be retrieved by the local predictor 104 by traversing each point to find the maximum points. A one-dimensional two-class classification task can be utilized by finding the optimal cutting position represented by idx opitmal . The inter-class distance can be maximized in combination with idx opitmal which can be defined by the following equations, but any variant thereof can be utilized:

[0070]

[0071] where, in an embodiment, and

[0072] The local predictor 104 can solve the non-linear optimization task by at least determining idx opitmal . The local predictor 104 can solve the non-linear optimization task by searching in the range [1, idxoptimal and (idx optimal , L] to search for the position of the maximum element to calculate idx left and idx rig . Their respective function values can be peak left and peak right .

[0073] The seat belt pixel intensity can change and vary in response to various ambient lighting. The ambient lighting can include various light sources, such as the interior lights of a vehicle. The local predictor 104 can learn the intensity distribution of the pixels. δ min and v max can be set to the lower and upper limits of the seat belt positioning instance. f patch The weighted intensity of f(x, y, θ) can be expressed as d intensity , and can be calculated by the following equation, although any variation thereof can be utilized:

[0074]

[0075] where can be the Gaussian distribution weight assigned to each pixel in the seat belt tile.

[0076] In an embodiment, f intensity is defined by the following equation, although any variation thereof can be utilized:

[0077]

[0078] In different embodiments, the seat belt surface is smooth, with minimal intensity variation. The intensity variance can be utilized as a criterion for evaluating smoothness within the region of interest ω. In an embodiment, d smoothnes is defined by the following equation, although any variation thereof can be utilized:

[0079]

[0080] In an embodiment, f smoothness is defined by the following equation, although any variation thereof can be utilized:

[0081]

[0082] The local predictor 104 can output pixel candidates 106, which can include one or more indications of one or more pixels in the input image 102 that may potentially correspond to the seat belt depicted in the input image 102. The pixel candidates 106 can be input to the global assembler 108, which can remove inaccurate candidates from the pixel candidates 106 and selectively assemble eligible seat belt pixel candidates.

[0083] The global assembler 108 can be a collection of one or more hardware and / or software computing resources having executable instructions that, when executed, pool positive seat belt segment candidates by removing false positives based on seat belt attributes such as shape and position. The global assembler 108 can apply a set of constraints to the pixel candidates 106 to refine the pixel candidates 106 to obtain eligible pixel candidates 114. The set of constraints can include characteristics such as parameters of the camera that captured the input image 102 (e.g., camera lens dimensions, focus, principal point, distortion parameters), standard seat belt dimension ranges (e.g., width and / or length ranges), standard seat belt characteristics and / or parameters (e.g., standard seat belt color, material), vehicle parameters (e.g., vehicle layout, vehicle components), and the like. The global assembler 108 can apply the set of constraints to the pixel candidates 106 such that pixel candidates in the pixel candidates 106 that do not comply with or otherwise do not conform to the set of constraints can be removed to determine the eligible pixel candidates 114 (e.g., pixel candidates indicating a seat belt with a width significantly longer than the standard seat belt width range).

[0084] The global assembler 108 can include an automatic position mask generation 110 and a width range estimation 112. The automatic position mask generation 110 can be a collection of one or more hardware and / or software computing resources having executable instructions that generate one or more image masks of an image such as the input image 102 indicating one or more seat belts. The width range estimation 112 can be a collection of one or more hardware and / or software computing resources having executable instructions that estimate the range of the width of the seat belt depicted in an image such as the input image 102.

[0085] In an embodiment, for a three-point seat belt, the seat belt has three anchors, called points: the upper right anchor represented by A tr the lower left anchor represented by A bl and the lower right anchor represented by A br The belts from the upper right anchor and the lower right anchor can be inserted into the lower left anchor using a buckle. The techniques described herein in connection with the upper right anchor and the lower right anchor can be similarly applied to the lower left anchor and the lower right anchor, or any suitable restraint device anchor. In some embodiments, the seat belt anchors are movable. That is, for example, one or more seat belt anchors of a three-point seat belt can be adjusted, rotated, and / or tilted in various directions. In some embodiments, one or more seat belt anchors are adjusted to a specific angle via tilting. In some cases, the seat belt anchor is connected to a mounting plate that is attached to the side door of a passenger vehicle. The seat belt anchor can move relative to the mounting plate (e.g., by sliding up, down, left, or right).

[0086] A constant force may exist within the anchor for cinching the belt using a spring; the position distribution of the seat belt pixels observed via a fixed camera may be within a limited area. In some examples, the global assembler 108 includes different machine learning algorithms for learning the position distribution of the seat belt pixels. In various embodiments, the seat belt typically does not appear on top of the steering wheel or other objects in the vehicle system (e.g., entertainment system, driver assistance system). In some examples, the global assembler 108 estimates the range of the seat belt width in the image frame for filtering seat belt candidates. The global assembler 108 may process multiple seat belts within a frame. In some embodiments, one or more processes of the global assembler 108 are executed in parallel using algorithm-level parallelization. Algorithm-level parallelization may refer to the process in which the processes that can be executed in parallel (e.g., the processes or operations performed by a system for seat belt positioning) are first identified and then executed in parallel. The global assembler 108 may execute different optimization techniques, such as image downsampling, pixel step adjustment, angular sampling, and so on.

[0087] The automatic position mask generation 110 may generate a seat belt position mask, which may be a visual indication of the position of the seat belt within an image (e.g., the input image 102). The automatic position mask generation 110 may include one or more software programs, which may analyze one or more aspects of the camera (e.g., the camera that captures the input image 102). The automatic position mask generation 110 may analyze various configurations of the camera to generate a seat belt position mask, such as the calibration of the camera, the camera pose, the type of the camera lens, etc. The automatic position mask generation 110 may include one or more components that can perform camera calibration, camera positioning, and 3D reconstruction. Camera calibration may include one or more functions and / or processes, which may calibrate the camera (e.g., an infrared (IR) camera, a red-green-blue (RGB) camera, a color camera, etc.) and process the images from the camera by processing the internal attributes of the camera, such as the focusing lens, the principal point, or the undistortion parameters. Camera positioning may include one or more functions and / or processes, which may provide six degrees of freedom (6DOF) position information of the camera in a predefined world coordinate system. 3D reconstruction may include one or more functions and / or processes, which may determine the coordinates of any point in the 3D space by utilizing the encoded tags. Camera positioning may include a process that can determine the 6DOF pose of the camera with respect to the vehicle coordinate system and can retrieve the coordinates or position of the seat belt anchor within the camera coordinate system. Camera positioning may utilize one or more models of the vehicle to determine the coordinates or position of the seat belt anchor. 3D reconstruction may include a process that can calculate the relative relationship between the seat belt anchor and the camera coordinate system.

[0088] In an embodiment, the camera (e.g., the camera that captures the input image 102) is calibrated with the internal parameter K and the external parameters R and T. The seat belt may have a tr (Xtr , Y tr , Z tr ), and A bl (X bl , Y bl , Z bl ) represents the anchor. The automatic position mask generation 110 can obtain the corresponding coordinates of the anchor in the image coordinate system, which is called the anchor position. In an embodiment, the image position of the anchor is represented by A i (x i , y i ), where i = tr, bl, and although any variation of the following equation can be utilized:

[0089]

[0090] Most of the position distribution of the seat belt can fall into an ellipse with two anchors as the ends of the major axis. In an embodiment, the major axis distance is represented by d major and can be defined by the following equation, although any variation thereof can be utilized:

[0091]

[0092] The automatic position mask generation 110 can obtain the minor axis distance through one or more machine learning algorithms, which can analyze different seat belt fastening processes. In an embodiment, the minor axis distance is represented by d minor and the set position mask is represented by S location and is defined by the following equation, but any variation thereof can be utilized:

[0093]

[0094] These points can be defined in the local coordinate system of the ellipse, where the origin O ellipse can be at the midpoint of the segment A with image coordinates bl A tr . The x-axis can point from A bl to A tr , and the y-axis can be obtained by rotating the x-axis 90 degrees clockwise. The corresponding coordinates in the image coordinates can be represented by (x global , y global ), where:

[0095]

[0096] Where:

[0097]

[0098] The automatic position mask generation 110 can obtain the position mask S in the image coordinateslocation Seat belt pixel candidates generated by the local predictor 104 outside the mask may be omitted.

[0099] The width range estimation 112 may estimate the width range (also referred to as the width size) of the seat belt in the image (e.g., the input image 102). The standard width of the seat belt may be in the range of 46 - 49 mm (e.g., 2 inches), where most seat belts may be approximately about 47 mm, or any appropriate value within any appropriate range. The width of the seat belt in the image may depend on the pose of the seat belt. If the seat belt is parallel to the camera imaging sensor or extends towards the camera, the seat belt width may increase. If the seat belt surface is perpendicular to the image sensor or further away from the camera, the seat belt width may decrease. The seat belt width τ min The lower limit of can be set at least in part based on one or more sensitivity requirements. In an embodiment, τ min is set to half of τ max to avoid dealing with noise patterns.

[0100] In an embodiment, the seat belt point in the middle of the two anchors is represented by P s (X s , Y s , Z s ), and the standard width of the seat belt is represented by d std , where the possible corresponding point (represented by S pair ) set for the other edge is defined by the following equation, although any variation thereof may be utilized:

[0101]

[0102] The regular passenger arm distance may be represented by d arm , where the maximum width τ max may occur when the passenger grabs the seat belt and pushes the seat belt towards the camera. In an embodiment, the new set of seat belt point coordinates is represented by S new and is defined by the following equation, although any variation thereof may be utilized:

[0103]

[0104] The width range estimation 112 may extract the envelope of the positions of the new points and estimate the width upper limit τ max . The seat belt width may be used to determine the structured edge threshold and may be used to determine the tile size L that may need to be greater than τ max . The global assembler 108 may determine the positions of multiple seat belts within a single image. The global assembler 108 may generate a position mask for each seat belt region.

[0105] The global assembler 108 may utilize one or more graphics processing units (GPUs) to process an image (e.g., the input image 102). The global assembler 108 may use multiple threads of one or more GPUs to process multiple tiles simultaneously. The global assembler 108 may utilize the workflows of one or more GPUs to process the seatbelt regions in parallel. The global assembler 108 may use one or more threads of one or more GPUs to process one or more regions of the image in parallel or simultaneously. The global assembler 108 may assign the processing of multiple frames to different computing units of one or more GPUs to achieve frame-level acceleration. The global assembler 108 may downsample the input image to reduce the use of computing resources when processing the input image. In addition to the per-pixel traversal of the image, the global assembler 108 may also process the image (e.g., the input image 102) in multiple directions in other ways. The global assembler 108 may adjust the discretization of the tile orientation angles to reduce the use of computing resources. The global assembler 108 may aggregate or otherwise determine eligible pixel candidates 114, which may include pixels of the input image 102 corresponding to one or more seatbelts depicted in the input image 102.

[0106] The eligible pixel candidates 114 may include a set of pixel candidates for the seatbelt for each seatbelt region of interest. The eligible pixel candidates 114 may also be referred to as candidate pixels, pixel candidates, candidates, and / or variants thereof. The shape modeler 116 may be a collection of one or more hardware and / or software computing resources having executable instructions that, when executed, build a geometric seatbelt shape model based on the output from the global assembler 108 (e.g., the eligible pixel candidates 114) and map the model onto the input image. The shape modeler 116 may remove noise candidates from the eligible pixel candidates 114 and model the seatbelt shape.

[0107] In an embodiment, the set of candidates from the local predictor 104 is represented by S predictor and the position mask is represented by S location where the candidates for shape modeling from within S predictor are defined by the following equation, although any variation thereof may be utilized: modeling S

[0108] S modeling ={(x,y)| (x,y)∈S location && f predictor (x,y)≥γ pre} (22) where γ pre may be a threshold for the prediction response number, which may be proportional to the total number of tile orientations.

[0109] The shape modeler 116 can model the shape of the seat belt through a high-order polynomial curve. A high-order polynomial curve can refer to a polynomial curve having a degree higher than 2. The shape modeler 116 can model the shape of the seat belt through a polynomial curve of any suitable order. The order of the polynomial curve can correspond to the complexity of the shape of the seat belt, where a higher order corresponds to a more complex shape (e.g., a shape having more curves, oscillations, etc.). In some examples, a passenger may push the seat belt away while driving, which results in multiple y-values at various positions along the x-axis; this can cause the polynomial model to be inapplicable. In an embodiment, the shape modeler 116 transforms the global image coordinates to the anchor local coordinate system through the following equation, although any variant thereof may be utilized:

[0110]

[0111] Figure 6 Example 600 of a candidate map according to at least one embodiment is shown. Example 600 may depict a candidate map after transformation from global image coordinates to the anchor local coordinate system. It should be noted that Figure 6 Examples depicting potential candidates and curve fitting to the candidates are shown, and the candidates and the fitting curves may be any variant thereof. The potential candidates may correspond to any suitable pixel or pixel region that may depict an image (e.g., input image 102) of one or more seat belts, and the corresponding curve fitting to the potential candidates may indicate any suitable seat belt shape.

[0112] Example 600 may depict candidate pixels of the seat belt and a curve (e.g., a high-order polynomial curve) fitted to the candidate pixels corresponding to the seat belt shape. High-order polynomial regression can be utilized to model the shape of the seat belt because different users may stretch the seat belt into different shapes. In an embodiment, the observation point is represented by (x i , y i ), where y i corresponds to the following equation, although any variant thereof may be utilized:

[0113]

[0114] where β0, β1, β2…β N can be coefficients and N can be the polynomial order. The coefficients of the polynomial curve can be determined such that the polynomial curve corresponds to or approximates the shape of the seat belt.

[0115] The shape modeler 116 can utilize different algorithms to reduce noise, such as the M - estimator sample consensus (MSAC) algorithm. The shape modeler 116 can transform the curve into the global image plane for visualization. The shape modeler 116 can determine a model of the seat - belt shape that is at least approximately the shape and / or position of the seat - belt from an image depicting the seat - belt (e.g., the input image 102). The shape modeler 116 can further determine whether the seat - belt identified in the image (e.g., the input image 102) is correctly applied. In some examples, the shape modeler 116 provides the model of the seat - belt to one or more neural networks that are trained to reason whether the seat - belt is correctly applied based on the model. One or more neural networks can determine whether the seat - belt is correctly applied based on the seat - belt model obtained from the shape modeler 116, and provide an indication to the shape modeler 116 as to whether the seat - belt is correctly applied. The shape modeler 116 can visualize the modeled seat - belt shape and whether the seat - belt corresponding to the modeled seat - belt shape is correctly applied as the modeled input image 118. In an embodiment, the modeled input image 118 is the input image 102 with the model of the visualized seat - belt shape. The modeled input image 118 can depict the input image 102 with the position and shape of the seat - belt of the input image 102 indicated by one or more visualizations (such as the visual boundary of the seat - belt, etc.). The modeled input image 118 can further include one or more visualizations that can indicate whether the seat - belt is correctly worn and / or applied.

[0116] For example, see Figure 1 , the modeled input image 118 includes the visual boundary of the seat - belt and an indication that the seat - belt is correctly worn and applied represented by "Seat - belt: ON". The indication that the seat - belt is correctly worn and applied can be represented in different ways, such as "Seat - belt: ON", "Seat - belt: Applied", etc. The indication that the seat - belt is worn but incorrectly applied can be represented in different ways, such as "Seat - belt: OFF", "Seat - belt: Incorrect", etc. The indication that the seat - belt is not worn can be represented in different ways, such as "Seat - belt: OFF", "Seat - belt: Not Applied", etc. A correctly applied or properly positioned seat - belt can be a seat - belt that is in the locked position and worn diagonally across the chest of the occupant. A seat - belt that is incorrectly applied or out of position can be a seat - belt that is not correctly applied or not in the proper position.

[0117] A system for seat belt positioning can provide a model of the seat belt and an indication of whether the seat belt is correctly applied to one or more systems of a vehicle. In some examples, the system for seat belt positioning is implemented by one or more computer servers, where the model of the seat belt and the indication of whether the seat belt is correctly applied are provided to one or more systems of the vehicle via one or more communication networks. As a result of obtaining the model of the seat belt and the indication of whether the seat belt is correctly applied, the systems of the vehicle can perform different actions. If the seat belt is not worn or is worn but incorrectly applied, the systems of the vehicle can provide a warning indication. The warning indication can be an audio indication such as a warning sound, a visual indication such as a warning light, a physical indication such as a warning vibration, etc. If the seat belt is not worn or is worn but incorrectly applied, the systems of the vehicle may cause one or more propulsion systems of the vehicle to stop or halt the propulsion of the vehicle. If the seat belt is not worn or is worn but incorrectly applied, the systems of the vehicle can provide an indication to different systems (e.g., a safety monitoring system) via one or more networks.

[0118] The system for seat belt positioning can be a passive system. The system may not need to modify the existing systems of the vehicle to detect the seat belt of the vehicle's occupant from one or more images of the seat belt. The system can passively detect the seat belt, which can refer to a detection that does not require input from one or more sensors of the vehicle or modification to one or more systems of the vehicle (e.g., the system may not require the vehicle's seat belt to have an identification mark). The system for seat belt positioning can combine models determined from multiple images to determine a final seat belt model. In some examples, the system for seat belt positioning locates and models a set of seat belts from a collection of images depicting the occupant and the corresponding seat belts, where the system combines the individual models to determine the final seat belt model for the corresponding seat belt of the occupant. The system for seat belt positioning can combine the determined seat belt models together or otherwise enhance them to determine the final seat belt model such that the confidence in the final seat belt model is increased. The system for seat belt positioning can combine the seat belt models together to determine the final seat belt model by evaluating the similarities and differences between the seat belt models. The system for seat belt positioning can combine the inferences made from one or more images to determine the final seat belt model and increase the confidence in the final seat belt model.

[0119] In some examples, a system for seatbelt positioning may be used in conjunction with one or more active seatbelt detection systems. An active seatbelt detection system may refer to one or more systems of a vehicle for directly detecting and positioning the seatbelt of the vehicle. The active seatbelt detection system may include using various sensors to detect and position the seatbelt, such as pressure sensors, weight sensors, motion sensors, or other sensors, and may require modification of the vehicle's existing systems (e.g., the active seatbelt detection system may require the vehicle's seatbelt to have an identification mark). The system for seatbelt positioning may position the seatbelt and utilize inputs from one or more active seatbelt detection systems to verify that the positioned seatbelt has been correctly determined. Alternatively, the system may receive one or more models of the positioned seatbelt from one or more active seatbelt detection systems and position the seatbelt through one or more images to verify that one or more models of the positioned seatbelt have been correctly determined. The system may receive or obtain inputs from one or more vehicle systems (such as an active seatbelt detection system) to position the seatbelt. For example, the system for seatbelt positioning obtains data from the vehicle's sensors, such as pressure sensor data, weight sensor data, motion sensor data, etc., and utilizes this data in combination with the above techniques to position and model the seatbelt from an image of the seatbelt in the vehicle. Sensor data may be used to determine constraints for filtering pixel candidates (e.g., the sensor data may provide information indicating the position or location of the seatbelt, which can be used to remove inaccurate candidates among the pixel candidates that do not conform to the information provided by the sensor data).

[0120] The system for seatbelt positioning may position and model the seatbelt through a tile matching and tile refinement process (such as those described above and / or different algorithms, such as tile matching algorithms, mixed resolution tile matching (MRPM) algorithms, and / or variants thereof). The tile matching process may obtain image tiles as inputs to one or more neural networks to extract tile features and evaluate tile similarity. The system for seatbelt positioning may define or otherwise be provided with the features, characteristics, behaviors, etc. of the seatbelt tiles and utilize different tile matching processes to construct an estimated seatbelt tile and refine the estimated seatbelt tile to determine the final seatbelt model. Tile matching may be in the spatial and / or temporal dimensions.

[0121] Figure 7 An example 700 of multiple seatbelt positionings according to at least one embodiment is shown. The input image 702, the seatbelt locator 704, and the modeled input image 706 may be in accordance with the combination Figures 1 to 6Those described. The seat belt locator 704 can be a collection of one or more hardware and / or software computing resources having executable instructions that, when executed, model one or more restraint devices from one or more images. The seat belt locator 704 can obtain or otherwise receive an input image 702, perform an initial classification of regions of the input image 702 that represent restraint devices (e.g., seat belts), apply a set of constraints to the regions of the input image 702 to refine the initial classification to obtain a refined classification, and generate a model of the restraint device based at least in part on the refined classification. The seat belt locator 704 can identify and model multiple seat belts depicted in an image (e.g., the input image 702).

[0122] Referring Figure 7 , the input image 702 can be an image captured from one or more cameras inside the vehicle. The input image 702 can depict the occupants of the vehicle, including a passenger that can be depicted on the left side and a driver that can be depicted on the right side. The input image 702 can depict the passenger wearing the seat belt correctly and the driver wearing the seat belt correctly. The seat belt locator 704 can analyze the input image 702 to identify and model the seat belts in the input image 702. The seat belt locator 704 can visualize the modeled seat belts via a modeled input image 706.

[0123] See Figure 7 , the modeled input image 706 can depict the input image 702 with the indicated seat belts and seat belt states (e.g., orientation / position). The seat belt locator 704 can perform an initial classification in conjunction with the passenger depicted in the input image 702, refine the initial classification to obtain a refined classification, and generate a model for the passenger's seat belt based on the refined classification, and perform a second initial classification in conjunction with the driver depicted in the input image 702, refine the second initial classification to obtain a second refined classification, and generate a second model based on the second refined classification for the driver's seat belt.

[0124] The modeled input image 706 can include a first visual indication of the passenger's seat belt having a corresponding label indicating the state of the passenger's seat belt. See Figure 7 , the modeled input image 706 can depict on the left side a box indicating the passenger's seat belt and the label "Seat Belt: On" indicating that the passenger's seat belt is correctly worn and applied. The modeled input image 706 can include a second visual indication of the driver's seat belt having a corresponding label indicating the state of the driver's seat belt. See Figure 7 , the modeled input image 706 can depict on the right side a box indicating the driver's seat belt and the label "Seat Belt: On" indicating that the driver's seat belt is correctly worn and applied.

[0125] Figure 8 An example 800 of seatbelt positioning according to at least one embodiment is shown. The input image 802, the seatbelt locator 804, and the modeled input image 806 may be according to those described in conjunction with Figures 1 to 6 . The seatbelt locator 804 may be a collection of one or more hardware and / or software computing resources having executable instructions that, when executed, model one or more constraints for one or more images. The seatbelt locator 804 may obtain or otherwise receive the input image 802, perform an initial classification of regions of the input image 802 that represent a constraint device (e.g., a seatbelt), apply a set of constraints to the regions of the input image 802 to refine the initial classification to obtain a refined classification, and generate a model of the constraint device based at least in part on the refined classification. The seatbelt locator 804 may identify and model the seatbelt depicted in the image (e.g., the input image 802).

[0126] Referring to Figure 8 , the input image 802 may be an image captured from one or more cameras inside the vehicle. The input image 802 may depict an occupant of the vehicle, including the driver, who may be depicted in the center. The input image 802 may depict the driver wearing the seatbelt incorrectly (e.g., the driver wearing the seatbelt behind the back). The seatbelt locator 804 may analyze the input image 802 to identify and model the seatbelt in the input image 802. The seatbelt locator 804 may visualize the modeled seatbelt via the modeled input image 806.

[0127] See Figure 8 , the modeled input image 806 may depict the input image 802 with the indicated seatbelt and seatbelt status (e.g., orientation / position). The modeled input image 806 may include a first visual indication of the driver's seatbelt with a corresponding label indicating the status of the driver's seatbelt. See Figure 8 , the modeled input image 806 may depict a box indicating the driver's seatbelt and a label "Seatbelt: Off" indicating that the driver's seatbelt is worn but incorrectly applied.

[0128] Figure 9 An example 900 of seatbelt positioning according to at least one embodiment is shown. The input image 902, the seatbelt locator 904, and the modeled input image 906 may be according to those described in conjunction with Figures 1 to 6Those described. The seatbelt locator 904 can be a collection of one or more hardware and / or software computing resources having executable instructions that, when executed, model one or more restraint devices from one or more images. The seatbelt locator 904 can obtain or otherwise receive an input image 902, perform an initial classification of regions of the input image 902 that represent restraint devices (e.g., seatbelts), apply a set of constraints to the regions of the input image 902 to refine the initial classification to obtain a refined classification, and generate a model of the restraint device based at least in part on the refined classification. The seatbelt locator 904 can identify and model a seatbelt depicted in an image (e.g., input image 902).

[0129] Referring Figure 9 , the input image 902 can be an image captured from one or more cameras inside the vehicle. The input image 902 can depict an occupant of the vehicle, including a driver who may be depicted in the center. The input image 902 can depict the driver wearing the seatbelt incorrectly (e.g., the driver wearing the seatbelt under the arm). The seatbelt locator 904 can analyze the input image 902 to identify and model the seatbelt in the input image 902. The seatbelt locator 904 can visualize the modeled seatbelt via a modeled input image 906.

[0130] See Figure 9 , the modeled input image 906 can depict the input image 902 with an indicated seatbelt and seatbelt status (e.g., orientation / position). The modeled input image 906 can include a first visual indication of the driver's seatbelt with a corresponding label indicating the status of the driver's seatbelt. See Figure 9 , the modeled input image 906 can depict a box indicating the driver's seatbelt and a label "Seatbelt: Off" indicating that the driver's seatbelt is worn but incorrectly applied.

[0131] Figure 10 Illustrates an example 1000 of seatbelt positioning according to at least one embodiment. The modeled input image 1002 can be according to those described in connection with Figures 1 to 6 Those described. The seatbelt locator can be a collection of one or more hardware and / or software computing resources having executable instructions that, when executed, model one or more restraint devices from one or more images. The seatbelt locator can obtain or otherwise receive an input image, perform an initial classification of regions of the input image that represent restraint devices (e.g., seatbelts), apply a set of constraints to the regions of the input image to refine the initial classification to obtain a refined classification, generate a model of the restraint device based at least in part on the refined classification, and visualize the model as a modeled input image 1002.

[0132] The input image 1002 for modeling can be based on input images that can be captured from one or more cameras. The input images can be captured from a camera with a standard FOV. The seat belt locator can identify and model one or more seat belts from images of any FOV. The seat belt locator can identify and model one or more seat belts from images that can depict portions of the seat belts. See Figure 10 , the input image 1002 for modeling can depict a standard FOV image having a seat belt identified by a shaded area and a corresponding label "Seat Belt: On" indicating that the seat belt is worn and properly applied.

[0133] Figure 11 An example 1100 of seat belt positioning according to at least one embodiment is shown. The input image 1102 for modeling can be according to those described in conjunction with Figures 1 to 6 . The seat belt locator can be a collection of one or more hardware and / or software computing resources having executable instructions that, when executed, model one or more restraint devices from one or more images. The seat belt locator can obtain or otherwise receive the input image, perform an initial classification of regions of the input image that represent restraint devices (e.g., seat belts), apply a set of constraints to the regions of the input image to refine the initial classification to obtain a refined classification, generate a model of the restraint device at least in part based on the refined classification, and visualize the model as the input image 1102 for modeling.

[0134] The input image 1102 for modeling can be based on input images that can be captured from one or more cameras. The input image can depict a driver who has pushed the seat belt away from the driver's body. The seat belt locator can identify and model one or more seat belts from images that depict one or more seat belts in any suitable position (e.g., pushed away, extended, pulled in). See Figure 11 , the input image 1102 for modeling can depict an image having a seat belt identified by a dashed curve and a corresponding label "Seat Belt: Off" indicating that the seat belt is worn but not properly applied.

[0135] It should be noted that while Figures 7 to 11Illustrates an example of a seat belt locator that identifies, models, visualizes, and marks one or more seat belts. However, the seat belt locator can identify, model, visualize, and mark one or more seat belts in any suitable manner. The seat belt locator can visualize one or more seat belts in an image by indicating the one or more seat belts via a contour, a box, a curve, a bounding box, a shaded area, a patterned area, or other indication. The seat belt locator can mark the one or more seat belts in the image with different indications (e.g., symbols, characters, labels, or other indications). Indications that a seat belt is worn and correctly applied can be represented in different ways, such as displaying text on a display screen in a vehicle indicating: "Seat belt: On", "Seat belt: Applied", "Seat belt: Correctly applied", etc. Indications that a seat belt is worn but incorrectly applied can be represented in different ways, such as "Seat belt: Off", "Seat belt: Incorrect", "Seat belt: Incorrectly applied", etc. Indications that a seat belt is not worn can be represented in different ways, such as "Seat belt: Off", "Seat belt: Not applied", "Seat belt not applied", etc. The indications can also include audio indications (e.g., computer voice indications), visual indications (e.g., indicator lights, switches), physical indications (e.g., indicating vibrations), etc.

[0136] Now refer to Figure 12 , each block of the method 1200 described herein includes a computational process that can be performed using any combination of hardware, firmware, and / or software. For example, different functions can be performed by a processor executing instructions stored in a memory. These methods can also be embodied as computer-usable instructions stored on a computer storage medium. These methods can be provided by a stand-alone application, service, or hosted service (independently or in combination with another hosted service) or a plug-in of another product, to name a few. Additionally, by way of example, the method 1200 is described with respect to Figure 1 a system for seat belt localization. However, these methods can alternatively or additionally be performed by any one system or any combination of systems, including but not limited to those described herein.

[0137] Figure 12 is a flowchart showing a method 1200 for generating a model for an input image according to some embodiments of the present disclosure. At block 1202, the method 1200 includes: obtaining an input image. The input image can represent a subject (such as a driver, a passenger, or other occupant) and a restraint device (such as a seat belt, a safety harness, or other restraint device) applied to the subject. The input image can be an image captured from one or more image capture devices (such as a camera or other device) that can be located inside a vehicle.

[0138] At block 1204, method 1200 includes: determining pixel candidates. A system that executes at least a portion of method 1200 may perform an initial classification on a region of an image that represents a restraint device to. The initial classification may result in the determination of pixel candidates. For any given pixel of an input image, the system may use the pixel and its neighboring pixel information to predict whether the pixel belongs to a seat belt or a background pixel. The system may determine a region of the input image that depicts or otherwise represents a restraint device (e.g., a seat belt) to determine pixel candidates.

[0139] At block 1206, method 1200 includes: removing false positives to determine eligible pixel candidates. A system that executes at least a portion of method 1200 may apply a set of constraints to a region of the image to refine the initial classification to obtain a refined classification. The refined classification may result in the determination of eligible pixel candidates. The set of constraints may include characteristics such as parameters of a camera that captured the input image (e.g., camera lens dimensions, focus, principal point, distortion parameters), standard seat belt dimension ranges (e.g., width and / or length ranges), standard seat belt characteristics and / or parameters (e.g., standard seat belt color, material), vehicle parameters (e.g., vehicle layout, vehicle components), and the like. The system may apply the set of constraints to the pixel candidates such that pixels in the pixel candidates that do not comply with or otherwise do not conform to the set of constraints may be removed to determine eligible pixel candidates.

[0140] At block 1208, method 1200 includes: building a model based on the eligible pixel candidates. A system that executes at least a portion of method 1200 may generate a model of the restraint device at least in part based on the refined classification. The system may generate the model based on the pixels of the eligible pixel candidates corresponding to the seat belt. At block 1210, method 1200 includes: mapping the model onto the input image. The system may visualize the model onto the input image as the modeled input image. The modeled input image may include an indication of the position / orientation / location of the seat belt and the status of the seat belt indicating whether it is worn, not worn, applied, or incorrectly applied.

[0141] Example autonomous vehicle

[0142] Figure 13AFIG. is an illustration of an example autonomous vehicle 1300 in accordance with some embodiments of the present disclosure. The autonomous vehicle 1300 (alternatively referred to herein as "vehicle 1300") may include, but is not limited to, passenger vehicles such as cars, trucks, buses, ambulances, shuttles, electric or motorized bicycles, motorcycles, fire trucks, police cars, ambulances, boats, construction vehicles, underwater vehicles, drones, and / or other types of vehicles (e.g., driverless and / or accommodating one or more passengers). Autonomous vehicles are generally described according to levels of automation defined by the "Classification and Definitions of Terms Related to Driving Automation Systems for On-Road Motor Vehicles" of the National Highway Traffic Safety Administration (NHTSA) under the U.S. Department of Transportation and the Society of Automotive Engineers (SAE) (Standard No. J3016 - 201806 released on June 15, 2018, Standard No. J3016 - 201609 released on September 30, 2016, and prior and future versions of this standard). Vehicle 1300 may be capable of having functionality according to one or more of levels 3 - 5 of the autonomous driving level. For example, depending on the embodiment, vehicle 1300 may have conditional automation (level 3), high automation (level 4), and / or full automation (level 5).

[0143] Vehicle 1300 may include components such as a chassis, a body, wheels (e.g., 2, 4, 6, 8, 18, etc.), tires, axles, and other components of the vehicle. Vehicle 1300 may include a propulsion system 1350, such as an internal combustion engine, a hybrid power plant, a fully electric motor, and / or another type of propulsion system. Propulsion system 1350 may be connected to the driveline of vehicle 1300, which may include a transmission, to effect the propulsion of vehicle 1300. Propulsion system 1350 may be controlled in response to signals received from throttle / accelerator 1352.

[0144] A steering system 1354, which may include a steering wheel, may be used to steer vehicle 1300 (e.g., along a desired path or route) when propulsion system 1350 is operating (e.g., when the vehicle is in motion). Steering system 1354 may receive signals from a steering actuator 1356. For fully autonomous (level 5) functionality, the steering wheel may be optional.

[0145] A brake sensor system 1346 may be used to operate vehicle brakes in response to signals received from a brake actuator 1348 and / or a brake sensor.

[0146] May include one or more CPUs, a system-on-chip (SoC) 1304( Figure 13C) and / or one or more controllers 1336 of one or more GPUs may provide signals (e.g., representing commands) to one or more components and / or systems of vehicle 1300. For example, one or more controllers may send signals to operate vehicle brakes via one or more brake actuators 1348, operate steering system 1354 via one or more steering actuators 1356, and operate propulsion system 1350 via one or more throttles / accelerators 1352. One or more controllers 1336 may include one or more on-board (e.g., integrated) computing devices (e.g., supercomputers) that process sensor signals and output operation commands (e.g., signals representing commands) to enable autonomous driving and / or assist a human driver in driving vehicle 1300. One or more controllers 1336 may include a first controller 1336 for autonomous driving functions, a second controller 1336 for functional safety functions, a third controller 1336 for artificial intelligence functions (e.g., computer vision), a fourth controller 1336 for infotainment functions, a fifth controller 1336 for redundancy in emergency situations, and / or other controllers. In some examples, a single controller 1336 may handle two or more of the above functions, two or more controllers 1336 may handle a single function, and / or any combination thereof.

[0147] One or more controllers 1336 may provide signals for controlling one or more components and / or systems of vehicle 1300 in response to sensor data (e.g., sensor inputs) received from one or more sensors. Sensor data may be received from, for example and without limitation, global navigation satellite system sensors 1358 (e.g., global positioning system sensors), RADAR sensors 1360, ultrasonic sensors 1362, LIDAR sensors 1364, inertial measurement unit (IMU) sensors 1366 (e.g., accelerometers, gyroscopes, magnetic compasses, magnetometers, etc.), microphones 1396, stereo cameras 1368, wide-angle cameras 1370 (e.g., fisheye cameras), infrared cameras 1372, surround cameras 1374 (e.g., 360-degree cameras), remote and / or mid-range cameras 1398, speed sensors 1344 (e.g., for measuring the rate of vehicle 1300), vibration sensors 1342, steering sensors 1340, brake sensors (e.g., as part of brake sensor system 1346), and / or other sensor types.

[0148] One or more of the controllers 1336 may receive input (e.g., represented by input data) from the instrument cluster 1332 of the vehicle 1300 and provide output (e.g., represented by output data, display data, etc.) via a human machine interface (HMI) display 1334, an audible annunciator, a speaker, and / or via other components of the vehicle 1300. These outputs may include information such as vehicle speed, velocity, time, map data (e.g., Figure 13C The HMI display 1334 may include information such as the HD map 1322 of the vehicle 1300 , location data (e.g., the location of the vehicle 1300 on the map), directions, locations of other vehicles (e.g., an occupancy grid), information about objects and states of objects as sensed by the controller 1336 , and the like. For example, the HMI display 1334 may display information about the presence of one or more objects (e.g., street signs, warning signs, traffic light changes, etc.) and / or information about driving maneuvers that the vehicle has made, is making, or will make (e.g., changing lanes now, leaving 34B in two miles, etc.).

[0149] The vehicle 1300 further includes a network interface 1324 that can communicate via one or more networks using one or more wireless antennas 1326 and / or a modem. For example, the network interface 1324 may be capable of communicating via LTE, WCDMA, UMTS, GSM, CDMA2000, etc. The one or more wireless antennas 1326 may also enable communication between objects in the environment (e.g., vehicles, mobile devices, etc.) using one or more local area networks such as Bluetooth, Bluetooth LE, Z-wave, ZigBee, etc. and / or one or more low power wide area networks (LPWANs) such as LoRaWAN, SigFox, etc.

[0150] Figure 13B For use according to some embodiments of the present disclosure Figure 13A 1300. The cameras and respective fields of view are an example embodiment and are not intended to be limiting. For example, additional and / or replaceable cameras may be included and / or the cameras may be located at different locations on the vehicle 1300.

[0151] The camera types for the camera can include, but are not limited to, digital cameras that can be adapted to be used with components and / or systems of the vehicle 1300. The camera can operate under Automotive Safety Integrity Level (ASIL) B and / or under another ASIL. The camera type can have any image capture rate, such as 60 frames per second (fps), 920 fps, 240 fps, etc., depending on the embodiment. The camera may be capable of using a rolling shutter, a global shutter, another type of shutter, or a combination thereof. In some examples, the color filter array can include a red clear (RCCC) color filter array, a red clear blue (RCCB) color filter array, a red blue green clear (RBGC) color filter array, a Foveon X3 color filter array, a Bayer sensor (RGGB) color filter array, a monochrome sensor color filter array, and / or another type of color filter array. In some embodiments, clear pixel cameras, such as cameras having RCCC, RCCB, and / or RBGC color filter arrays, can be used in efforts to improve light sensitivity.

[0152] In some examples, one or more of the cameras can be used to perform Advanced Driver Assistance System (ADAS) functions (e.g., as part of a redundant or fail-safe design). For example, a multi-functional monocular camera can be installed to provide functions including lane departure warning, traffic sign assistance, and smart headlight control. One or more of the cameras (e.g., all cameras) can simultaneously record and provide image data (e.g., video).

[0153] One or more of the cameras can be mounted in a mounting component, such as a custom-designed (3-D printed) component, to cut off stray light and reflections from inside the vehicle (e.g., reflections from the dashboard reflected in the windshield mirror) that may interfere with the image data capture ability of the camera. Regarding the wing mirror mounting component, the wing mirror component can be custom 3-D printed such that the camera mounting plate matches the shape of the wing mirror. In some examples, one or more cameras can be integrated into the wing mirror. For side view cameras, one or more cameras can also be integrated into the four pillars at each corner of the cab.

[0154] A camera having a field of view that includes an environmental portion in front of the vehicle 1300 (e.g., a front camera) can be used for surround view to help identify the forward path and obstacles, and to assist in providing information crucial for generating an occupancy grid and / or determining a preferred vehicle path with the help of one or more controllers 1336 and / or a control SoC. The front camera can be used to perform many of the same ADAS functions as LIDAR, including emergency braking, pedestrian detection, and collision avoidance. The front camera can also be used for ADAS functions and systems, including lane departure warning (“LDW”), adaptive cruise control (“ACC”), and / or other functions such as traffic sign recognition.

[0155] A variety of cameras can be used in a front-facing configuration, including, for example, a monocular camera platform including a CMOS (Complementary Metal Oxide Semiconductor) color imager. Another example could be a wide-angle camera 1370, which can be used to sense objects (such as pedestrians, intersection traffic, or bicycles) entering the field of view from the periphery. Although Figure 13B only one wide-angle camera is illustrated in the figure, any number of wide-angle cameras 1370 can be present on the vehicle 1300. Additionally, a long-range camera 1398 (such as a long-range stereo camera pair) can be used for depth-based object detection, especially for objects for which a neural network has not been trained. The long-range camera 1398 can also be used for object detection and classification and basic object tracking.

[0156] One or more stereo cameras 1368 can also be included in the front-facing configuration. The stereo camera 1368 can include an integrated control unit that includes a scalable processing unit that can provide a multi-core microprocessor with an integrated CAN or Ethernet interface and programmable logic (e.g., an FPGA) on a single chip. Such a unit can be used to generate a 3-D map of the vehicle environment, including distance estimates for all points in the image. An alternative stereo camera 1368 can include a compact stereo vision sensor that can include two camera lenses (one on the left and one on the right) and an image processing chip that can measure the distance from the vehicle to a target object and use the generated information (such as metadata) to activate autonomous emergency braking and lane departure warning functions. Other types of stereo cameras 1368 can be used in addition to or in place of those described herein.

[0157] Cameras having a field of view of an environmental portion including the side of the vehicle 1300 (such as side-view cameras) can be used for surround view, providing information used to create and update an occupancy grid and generate side-impact collision warnings. For example, surround cameras 1374 (such as the four surround cameras 1374 shown in Figure 13B the figure) can be placed around the vehicle 1300. The surround cameras 1374 can include wide-angle cameras 1370, fish-eye cameras, 360-degree cameras, and / or the like. For example, four fish-eye cameras can be placed in the front, rear, and sides of the vehicle. In an alternative arrangement, the vehicle can use three surround cameras 1374 (such as on the left, right, and rear) and can utilize one or more other cameras (such as a forward-facing camera) as a fourth surround camera.

[0158] A camera (e.g., a rear-view camera) having a field of view that includes an environmental portion behind vehicle 1300 can be used to assist with parking, surround view, rear collision warning, and creating and updating an occupancy grid. A variety of cameras can be used, including but not limited to cameras that are also suitable as front cameras as described herein (e.g., long-range and / or mid-range cameras 1398, stereo cameras 1368, infrared cameras 1372, etc.).

[0159] Figure 13C For an example autonomous vehicle 1300 according to some embodiments of the present disclosure Figure 13A is a block diagram of an example system architecture. It should be understood that this and other arrangements described herein are presented only as examples. Other arrangements and elements (e.g., machines, interfaces, functions, orders, function groupings, etc.) can be used in addition to or instead of those shown, and some elements can be omitted entirely. Further, many of the elements described herein are functional entities that can be implemented as discrete or distributed components or in combination with other components, and in any suitable combination and location. The various functions described herein as being performed by entities can be implemented by hardware, firmware, and / or software. For example, the various functions can be implemented by a processor executing instructions stored in a memory.

[0160] Figure 13C Each of the components, features, and systems in vehicle 1300 is illustrated as being connected via bus 1302. Bus 1302 can include a Controller Area Network (CAN) data interface (alternatively referred to herein as the "CAN bus"). CAN can be a network within vehicle 1300 used to assist in controlling various features and functions of vehicle 1300, such as driving brakes, acceleration, braking, steering, windshield wipers, etc. The CAN bus can be configured to have dozens or even hundreds of nodes, each with its own unique identifier (e.g., CAN ID). The CAN bus can be read to find steering wheel angle, ground speed, engine revolutions per minute (RPM), button positions, and / or other vehicle status indicators. The CAN bus can be ASIL B compliant.

[0161] Although the bus 1302 is described herein as a CAN bus, this is not intended to be limiting. For example, in addition to or alternatively to a CAN bus, FlexRay and / or Ethernet can be used. Further, although a single line is used to represent the bus 1302, this is not intended to be limiting. For example, any number of buses 1302 can be present, which can include one or more CAN buses, one or more FlexRay buses, one or more Ethernet buses, and / or one or more other types of buses using different protocols. In some examples, two or more buses 1302 can be used to perform different functions, and / or can be used for redundancy. For example, a first bus 1302 can be used for a collision avoidance function, and a second bus 1302 can be used for drive control. In any example, each bus 1302 can communicate with any component of the vehicle 1300, and two or more buses 1302 can communicate with the same component. In some examples, each SoC 1304, each controller 1336, and / or each computer within the vehicle can have access to the same input data (e.g., input from sensors of the vehicle 1300), and can be connected to a common bus such as a CAN bus.

[0162] The vehicle 1300 can include one or more controllers 1336, such as those described herein with respect to Figure 13A The controllers 1336 can be used for a variety of functions. The controllers 1336 can be coupled to any other different components and systems of the vehicle 1300, and can be used for control of the vehicle 1300, artificial intelligence of the vehicle 1300, infotainment for the vehicle 1300, and / or the like.

[0163] The vehicle 1300 can include one or more system-on-chips (SoC) 1304. The SoC 1304 can include a CPU 1306, a GPU 1308, a processor 1310, a cache 1312, an accelerator 1314, a data store 1316, and / or other components and features not shown. In a variety of platforms and systems, the SoC 1304 can be used to control the vehicle 1300. For example, one or more SoC 1304 can be combined with an HD map 1322 in a system (e.g., a system of the vehicle 1300), and the HD map can obtain map refreshes and / or updates via a network interface 1324 from one or more servers (e.g., Figure 13D one or more servers 1378) of.

[0164] The CPU 1306 may include a CPU cluster or a CPU complex (alternatively referred to herein as a "CCPLEX"). The CPU 1306 may include multiple cores and / or L2 caches. For example, in some embodiments, the CPU 1306 may include eight cores in a coherent multi-processor configuration. In some embodiments, the CPU 1306 may include four dual-core clusters, each of which has a dedicated L2 cache (e.g., a 2MB L2 cache). The CPU 1306 (e.g., CCPLEX) may be configured to support simultaneous cluster operation such that any combination of the clusters of the CPU 1306 can be active at any given time.

[0165] The CPU 1306 may implement power management capabilities including one or more of the following features: Each hardware block may automatically perform clock gating when idle to save dynamic power; Each core clock may be gated when the core is not actively executing instructions due to the execution of WFI / WFE instructions; Each core may be independently power gated; When all cores are clock gated or power gated, each core cluster may be independently clock gated; and / or When all cores are power gated, each core cluster may be independently power gated. The CPU 1306 may further implement enhanced algorithms for managing power states, where allowed power states and desired wake-up times are specified, and the hardware / microcode determines the best power state for the cores, clusters, and CCPLEX to enter. The processing cores may support a simplified power state entry sequence in software, and this work is offloaded to the microcode.

[0166] The GPU 1308 may include an integrated GPU (alternatively referred to herein as an "iGPU"). The GPU 1308 may be programmable and efficient for parallel workloads. In some examples, the GPU 1308 may use an enhanced tensor instruction set. The GPU 1308 may include one or more streaming microprocessors, each of which may include an L1 cache (e.g., an L1 cache with at least 96KB of storage capacity), and two or more of these streaming microprocessors may share an L2 cache (e.g., an L2 cache with 512KB of storage capacity). In some embodiments, the GPU 1308 may include at least eight streaming microprocessors. The GPU 1308 may use a computer-based application programming interface (API). Additionally, the GPU 1308 may use one or more parallel computing platforms and / or programming models (e.g., NVIDIA's CUDA).

[0167] In automotive and embedded use cases, the GPU 1308 can be power optimized for best performance. For example, the GPU 1308 can be fabricated on fin field-effect transistors (FinFETs). However, this is not intended to be restrictive, and the GPU 1308 can be fabricated using other semiconductor manufacturing processes. Each streaming microprocessor can incorporate several mixed-precision processing cores divided into multiple blocks. By way of example and not limitation, 64 PF32 cores and 32 PF64 cores can be divided into four processing blocks. In such an example, each processing block can be allocated 16 FP32 cores, 8 FP64 cores, 16 INT32 cores, two mixed-precision NVIDIA tensor cores for deep learning matrix arithmetic, an L0 instruction cache, a warp scheduler, a dispatch unit, and / or a 64KB register file. Additionally, the streaming microprocessor can include separate parallel integer and floating-point data paths to enable efficient execution of workloads by leveraging a mix of compute and addressing computations. The streaming microprocessor can include separate thread scheduling capabilities to allow for finer-grained synchronization and cooperation between parallel threads. The streaming microprocessor can include a combined L1 data cache and shared memory unit to improve performance while simplifying programming.

[0168] The GPU 1308 can include, in some examples, a high-bandwidth memory (HBM) that provides a peak memory bandwidth of approximately 900 GB / s and / or a 16 GB HBM2 memory subsystem. In some examples, in addition to or alternatively to HBM memory, synchronous graphics random access memory (SGRAM), such as fifth-generation graphics double data rate synchronous random access memory (GDDR5), can be used.

[0169] The GPU 1308 can include unified memory technology that includes access counters to allow memory pages to be more precisely migrated to the processors that most frequently access them, thereby improving the efficiency of the memory ranges shared between processors. In some examples, address translation service (ATS) support can be used to allow the GPU 1308 to directly access the CPU 1306 page tables. In such an example, when the GPU 1308 memory management unit (MMU) experiences a miss, an address translation request can be transmitted to the CPU 1306. In response, the CPU 1306 can look up the virtual-physical mapping for the address in its page table and transmit the translation back to the GPU 1308. In this way, the unified memory technology can allow for a single unified virtual address space for the memories of both the CPU 1306 and the GPU 1308, thus simplifying GPU 1308 programming and porting applications to the GPU 1308.

[0170] In addition, GPU 1308 may include an access counter that can track how frequently the GPU 1308 accesses the memory of other processors. The access counter can help ensure that memory pages are moved to the physical memory of the processor that most frequently accesses those pages.

[0171] SoC 1304 may include any number of caches 1312, including those described herein. For example, cache 1312 may include an L3 cache (e.g., which is connected to both CPU 1306 and GPU 1308) that is available to both CPU 1306 and GPU 1308. Cache 1312 may include a write-back cache that can track the state of lines, for example, by using a cache coherence protocol (e.g., MEI, MESI, MSI, etc.). Depending on the embodiment, the L3 cache may include 4MB or more, but smaller cache sizes may also be used.

[0172] SoC 1304 may include an arithmetic logic unit (ALU) that can be used to perform processing for any of the various tasks or operations regarding vehicle 1300 - such as processing a DNN. In addition, SoC 1304 may include a floating point unit (FPU) - or other math co-processor or digital co-processor type - for performing mathematical operations within the system. For example, SoC 104 may include one or more FPUs integrated within the execution units of CPU 1306 and / or GPU 1308.

[0173] SoC 1304 may include one or more accelerators 1314 (e.g., hardware accelerators, software accelerators, or a combination thereof). For example, SoC 1304 may include a hardware acceleration cluster that may include optimized hardware accelerators and / or large on-chip memories. This large on-chip memory (e.g., 4MB SRAM) may enable the hardware acceleration cluster to accelerate neural networks and other computations. The hardware acceleration cluster may be used to supplement GPU 1308 and offload some of the tasks of GPU 1308 (e.g., freeing up more cycles of GPU 1308 for performing other tasks). As an example, accelerator 1314 may be used for targeted workloads that are stable enough to be easily accelerated (e.g., perception, convolutional neural network (CNN), etc.). When used herein, the term "CNN" may include all types of CNNs, including region-based or region convolutional neural networks (RCNN) and fast RCNN (e.g., for object detection).

[0174] The accelerator 1314 (e.g., a hardware acceleration cluster) may include a Deep Learning Accelerator (DLA). The DLA may include one or more Tensor Processing Units (TPUs) that can be configured to provide an additional one trillion operations per second for deep learning applications and inference. The TPU may be an accelerator configured to perform image processing functions (e.g., for CNN, RCNN, etc.) and optimized for performing image processing functions. The DLA may be further optimized for a specific set of neural network types and floating-point operations, as well as inference. The design of the DLA may provide higher performance per millimeter than a general-purpose GPU and far exceed the performance of a CPU. The TPU may perform several functions, including single-instance convolution functions, supporting INT8, INT16, and FP16 data types for both features and weights, respectively, and post-processor functions.

[0175] The DLA may execute neural networks, especially CNNs, quickly and efficiently for any of a variety of functions on processed or unprocessed data, such as, and not limited to: CNNs for object recognition and detection using data from camera sensors; CNNs for distance estimation using data from camera sensors; CNNs for emergency vehicle detection and identification and detection using data from microphones; CNNs for face recognition and vehicle owner recognition using data from camera sensors; and / or CNNs for security and / or safety-related events.

[0176] The DLA may perform any function of the GPU 1308, and by using an inference accelerator, for example, a designer may configure the DLA or the GPU 1308 for any function. For example, a designer may focus the processing and floating-point operations of a CNN on the DLA and leave other functions to the GPU 1308 and / or other accelerators 1314.

[0177] The accelerator 1314 (e.g., a hardware acceleration cluster) may include a Programmable Vision Accelerator (PVA), which may alternatively be referred to herein as a computer vision accelerator. The PVA may be designed and configured to accelerate computer vision algorithms for Advanced Driver Assistance Systems (ADAS), autonomous driving, and / or augmented reality (AR) and / or virtual reality (VR) applications. The PVA may provide a balance between performance and flexibility. For example, each PVA may include, for example, and not limited to, any number of Reduced Instruction Set Computer (RISC) cores, Direct Memory Access (DMA), and / or any number of vector processors.

[0178] The RISC cores can interact with an image sensor (such as the image sensor of any camera described herein), an image signal processor, and / or the like. Each of these RISC cores can include any number of memories. Depending on the embodiment, the RISC cores can use any of several protocols. In some examples, the RISC cores can execute a real-time operating system (RTOS). The RISC cores can be implemented using one or more integrated circuit devices, application-specific integrated circuits (ASICs), and / or storage devices. For example, the RISC cores can include an instruction cache and / or tightly coupled RAM.

[0179] The DMA can enable components of the PVA to access system memory independently of the CPU 1306. The DMA can support any number of features used to optimize the PVA, including but not limited to supporting multi-dimensional addressing and / or circular addressing. In some examples, the DMA can support addressing up to six or more dimensions, which can include block width, block height, block depth, horizontal block step, vertical block step, and / or depth step.

[0180] The vector processor can be a programmable processor that can be designed to efficiently and flexibly execute programming for computer vision algorithms and provide signal processing capabilities. In some examples, the PVA can include a PVA core and two vector processing subsystem partitions. The PVA core can include a processor subsystem, one or more DMA engines (such as two DMA engines), and / or other peripherals. The vector processing subsystem can operate as the main processing engine of the PVA and can include a vector processing unit (VPU), an instruction cache, and / or vector memory (such as VMEM). The VPU core can include a digital signal processor, such as, for example, a single instruction multiple data (SIMD), very long instruction word (VLIW) digital signal processor. The combination of SIMD and VLIW can enhance throughput and rate.

[0181] Each of the vector processors may include an instruction cache and may be coupled to dedicated memory. As a result, in some examples, each of the vector processors may be configured to execute independently of the other vector processors. In other examples, the vector processors included in a particular PVA may be configured to employ data parallelization. For example, in some embodiments, multiple vector processors included in a single PVA may execute the same computer vision algorithm, but on different regions of an image. In other examples, the vector processors included in a particular PVA may execute different computer vision algorithms simultaneously on the same image, or even execute different algorithms on sequential images or portions of an image. Among other things, any number of PVAs may be included in a hardware acceleration cluster, and any number of vector processors may be included in each of these PVAs. In addition, the PVA may include additional error correction code (ECC) memory to enhance overall system security.

[0182] Accelerator 1314 (e.g., a hardware acceleration cluster) may include an on-chip computer vision network and SRAM to provide high-bandwidth, low-latency SRAM for Accelerator 1314. In some examples, the on-chip memory may include at least 4MB SRAM consisting of, for example and without limitation, eight field-configurable memory blocks, which may be accessed by both the PVA and the DLA. Each pair of memory blocks may include an Advanced Peripheral Bus (APB) interface, configuration circuitry, a controller, and a multiplexer. Any type of memory may be used. The PVA and the DLA may access the memory via a backbone that provides high-speed memory access to the PVA and the DLA. The backbone may include (e.g., using APB) an on-chip computer vision network that interconnects the PVA and the DLA to the memory.

[0183] The on-chip computer vision network may include an interface that determines that both the PVA and the DLA provide ready and valid signals before transmitting any control signals / address / data. Such an interface may provide separate phases and separate channels for transmitting control signals / address / data, as well as burst communication for continuous data transmission. This type of interface may conform to the ISO 26262 or IEC 61508 standards, but other standards and protocols may also be used.

[0184] In some examples, SoC 1304 can include a real-time ray tracing hardware accelerator, such as that described in U.S. Patent Application No. 16 / 101,232, filed on August 10, 2018. The real-time ray tracing hardware accelerator can be used to quickly and efficiently determine the position and extent of objects (e.g., within a world model) in order to generate a real-time visualization simulation for RADAR signal interpretation, for sound propagation synthesis and / or analysis, for SONAR system simulation, for general wave propagation simulation, for comparison with LIDAR data for positioning and / or other functional purposes, and / or for other uses. In some embodiments, one or more tree traversal units (TTUs) can be used to perform one or more ray tracing-related operations.

[0185] Accelerator 1314 (e.g., a hardware accelerator cluster) has a wide range of autonomous driving applications. The PVA can be a programmable vision accelerator that can be used in critical processing stages in ADAS and autonomous vehicles. The capabilities of the PVA are a good match for algorithm domains that require predictable processing, low power, and low latency. In other words, the PVA performs well in semi-dense or dense regular computations, even on small data sets that require predictable runtimes with low latency and low power. Thus, in the context of a platform for autonomous vehicles, the PVA is designed to run classical computer vision algorithms because they are effective in object detection and integer math operations.

[0186] For example, according to one embodiment of the technology, the PVA is used to perform computer stereo vision. In some examples, an algorithm based on semi-global matching can be used, but this is not intended to be limiting. Many applications for level 3 - 5 autonomous driving require instant motion estimation / stereo matching (e.g., structure from motion, pedestrian recognition, lane detection, etc.). The PVA can perform computer stereo vision functions on inputs from two monocular cameras.

[0187] In some examples, the PVA can be used to perform dense optical flow. For example, the PVA can be used to process raw RADAR data (e.g., using a 4D fast Fourier transform) before transmitting the next RADAR pulse to provide processed RADAR signals. In other examples, the PVA is used for time-of-flight depth processing, which, for example, processes raw time-of-flight data to provide processed time-of-flight data.

[0188] DLA can be used to run any type of network to enhance control and driving safety, including, for example, a neural network that outputs a confidence metric for each object detection. Such confidence values can be interpreted as probabilities or as providing a relative "weight" of each detection compared to other detections. The confidence value enables the system to make further decisions regarding which detections should be considered true positive detections rather than false positive detections. For example, the system can set a threshold for the confidence and consider only detections that exceed the threshold as true positive detections. In an automatic emergency braking (AEB) system, false positive detections can cause the vehicle to automatically perform emergency braking, which is clearly undesirable. Therefore, only the most confident detections should be considered as triggers for AEB. DLA can run a neural network for regressing confidence values. The neural network can take as its input at least some subset of parameters, such as bounding box dimensions, a ground plane estimate obtained (e.g., from another subsystem), outputs of an inertial measurement unit (IMU) sensor 1366 related to the vehicle 1300 orientation, distance, a 3D position estimate of an object obtained from a neural network and / or other sensors (such as a LIDAR sensor 1364 or a RADAR sensor 1360), etc.

[0189] SoC 1304 can include one or more data stores 1316 (e.g., memories). The data store 1316 can be an on-chip memory of the SoC 1304, which can store neural networks to be executed on the GPU and / or DLA. In some examples, for redundancy and safety, the data store 1316 can be large enough in capacity to store multiple instances of the neural network. The data store 1316 can include an L2 or L3 cache 1312. References to the data store 1316 can include references to memories associated with the PVA, DLA, and / or other accelerators 1314 as described herein.

[0190] The SoC 1304 may include one or more processors 1310 (e.g., embedded processors). The processor 1310 may include a boot and power management processor, which may be a dedicated processor and subsystem for handling boot power and management functions as well as security implementation related. The boot and power management processor may be part of the SoC 1304 boot sequence and may provide runtime power management services. The boot power and management processor may provide clock and voltage programming, assist in system low power state transitions, manage the SoC 1304 thermal and temperature sensors, and / or manage the SoC 1304 power state. Each temperature sensor may be implemented as a ring oscillator whose output frequency is proportional to temperature, and the SoC 1304 may use the ring oscillator to detect the temperature of the CPU 1306, GPU 1308, and / or accelerator 1314. If it is determined that the temperature exceeds a threshold, then the boot and power management processor may enter a temperature fault routine and place the SoC 1304 in a lower power state and / or place the vehicle 1300 in a driver safety stop mode (e.g., safely stop the vehicle 1300).

[0191] The processor 1310 may further include a set of embedded processors that can be used as an audio processing engine. The audio processing engine may be an audio subsystem that allows for full hardware support for multi-channel audio over multiple interfaces and a wide range of flexible audio I / O interfaces. In some examples, the audio processing engine is a dedicated processor core with a digital signal processor with dedicated RAM.

[0192] The processor 1310 may further include an always-on processor engine that may provide the necessary hardware features to support low power sensor management and wake-up use cases. The always-on processor engine may include a processor core, tightly coupled RAM, support peripherals (e.g., timers and interrupt controllers), various I / O controller peripherals, and routing logic.

[0193] The processor 1310 may further include a security cluster engine that includes a dedicated processor subsystem for handling security management of automotive applications. The security cluster engine may include two or more processor cores, tightly coupled RAM, support peripherals (e.g., timers, interrupt controllers, etc.), and / or routing logic. In a secure mode, the two or more cores may operate in a lockstep mode and act as a single core with comparison logic to detect any differences between their operations.

[0194] The processor 1310 may further include a real-time camera engine that may include a dedicated processor subsystem for handling real-time camera management.

[0195] The processor 1310 may further include a high dynamic range signal processor, which may include an image signal processor, which is a hardware engine that is part of the camera processing pipeline.

[0196] The processor 1310 may include a video image compositor that may be a processing block (e.g., implemented on a microprocessor) that implements the video post-processing functions required for a video playback application to generate the final image for the player window. The video image compositor may perform lens distortion correction on the wide-angle camera 1370, the surround camera 1374, and / or the in-cab monitoring camera sensor. The in-cab monitoring camera sensor is preferably monitored by a neural network running on another instance of the advanced SoC, configured to identify in-cab events and respond accordingly. The in-cab system may perform lip reading to activate mobile phone services and make calls, dictate emails, change the vehicle destination, activate or change the vehicle's infotainment system and settings, or provide voice-activated web surfing. Certain functions are only available to the driver when the vehicle is operating in autonomous mode and are disabled otherwise.

[0197] The video image compositor may include enhanced temporal noise reduction for spatial and temporal noise reduction. For example, in the case of motion in the video, the noise reduction appropriately weights the spatial information, reducing the weight of the information provided by neighboring frames. In the case where the image or a portion of the image does not include motion, the temporal noise reduction performed by the video image compositor may use information from a previous image to reduce the noise in the current image.

[0198] The video image compositor may also be configured to perform stereo correction on input stereo lens frames. When the operating system desktop is in use and the GPU 1308 does not need to continuously render new surfaces, the video image compositor may further be used for user interface composition. Even when the GPU 1308 is powered on and actively performing 3D rendering, the video image compositor may be used to relieve the burden on the GPU 1308 to improve performance and responsiveness.

[0199] The SoC 1304 may further include a Mobile Industry Processor Interface (MIPI) camera serial interface, a high-speed interface, and / or a video input block for receiving video and inputs from cameras and may be used for camera and related pixel input functions. The SoC 1304 may further include an input / output controller that may be software-controlled and may be used to receive I / O signals not committed to a specific role.

[0200] SoC 1304 may further include a wide range of peripheral device interfaces to enable communication with peripheral devices, audio codecs, power management, and / or other devices. SoC 1304 can be used to process data from cameras (connected via Gigabit Multimedia Serial Link and Ethernet), sensors (such as LIDAR sensor 1364, RADAR sensor 1360, etc. that can be connected via Ethernet), data from bus 1302 (such as the speed of vehicle 1300, steering wheel position, etc.), and data from GNSS sensor 1358 (connected via Ethernet or CAN bus). SoC 1304 may further include dedicated high-performance large-capacity storage controllers, which may include their own DMA engines and can be used to free the CPU 1306 from routine data management tasks.

[0201] SoC 1304 can be an end-to-end platform with a flexible architecture that spans levels 3 - 5 of automation, thus providing an integrated functional safety architecture for a platform that leverages and efficiently uses computer vision and ADAS technologies to achieve diversity and redundancy, along with deep learning tools to provide a flexible and reliable driving software stack. SoC 1304 can be faster, more reliable, and even more energy-efficient and space-efficient than conventional systems. For example, when combined with CPU 1306, GPU 1308, and data storage 1316, accelerator 1314 can provide a fast and efficient platform for level 3 - 5 autonomous vehicles.

[0202] Thus, this technology provides capabilities and functions that cannot be achieved by conventional systems. For example, computer vision algorithms can be executed on CPUs that can be configured using high-level programming languages such as the C programming language to perform various processing algorithms across a variety of visual data. However, CPUs often cannot meet the performance requirements of many computer vision applications, such as those related to execution time and power consumption. In particular, many CPUs cannot execute complex object detection algorithms in real time, which is a requirement for in-vehicle ADAS applications and for practical level 3 - 5 autonomous vehicles.

[0203] In contrast to conventional systems, the technology described herein allows multiple neural networks to be executed simultaneously and / or sequentially by providing a CPU complex, a GPU complex, and a hardware acceleration cluster, and combining the results to achieve level 3 - 5 autonomous driving functions. For example, a CNN executed on a DLA or a dGPU (such as GPU 1320) can include text and word recognition, allowing a supercomputer to read and understand traffic signs, including signs for which the neural network has not been specifically trained. The DLA can further include a neural network capable of recognizing, interpreting, and providing semantic understanding of the signs and passing that semantic understanding to a path planning module running on the CPU complex.

[0204] As another example, as required for level 3, 4, or 5 driving, multiple neural networks can run simultaneously. For example, a warning sign consisting of "Caution: Flashing lights indicate icy conditions" together with the electric lights can be interpreted independently or jointly by several neural networks. The sign itself can be recognized as a traffic sign by a first neural network deployed (e.g., a trained neural network), and the text "Flashing lights indicate icy conditions" can be interpreted by a second deployed neural network that informs the vehicle's path planning software (preferably executed on the CPU complex) that there is an icy condition when the flashing lights are detected. The flashing lights can be recognized by operating a third deployed neural network on multiple frames, and this neural network informs the vehicle's path planning software of the presence (or absence) of the flashing lights. All three neural networks can run simultaneously, for example, within the DLA and / or on the GPU 1308.

[0205] In some examples, the CNNs for face recognition and vehicle owner recognition can use data from the camera sensors to identify the presence of an authorized driver and / or vehicle owner of the vehicle 1300. A processing engine that is always on the sensor can be used to unlock the vehicle and turn on the lights when the vehicle owner approaches the driver's door, and in a security mode, to disable the vehicle when the vehicle owner leaves the vehicle. In this way, the SoC 1304 provides security against theft and / or carjacking.

[0206] In another example, the CNN for emergency vehicle detection and recognition can use data from the microphone 1396 to detect and recognize an emergency vehicle siren. In contrast to conventional systems that use a general classifier to detect the siren and manually extract features, the SoC 1304 uses the CNN to classify environmental and urban sounds as well as visual data. In a preferred embodiment, the CNN running on the DLA is trained to recognize the relative closing rate of an emergency vehicle (e.g., by using the Doppler effect). The CNN can also be trained to recognize emergency vehicles specific to the local area in which the vehicle is operating as recognized by the GNSS sensor 1358. Thus, for example, when operating in Europe, the CNN will seek to detect European sirens, and when in the United States, the CNN will seek to identify only North American sirens. Once an emergency vehicle is detected, with the assistance of the ultrasonic sensor 1362, a control program can be used to execute an emergency vehicle safety routine to slow down the vehicle, pull over to the side of the road, stop the vehicle, and / or idle the vehicle until the emergency vehicle passes.

[0207] The vehicle may include a CPU 1318 (e.g., a discrete CPU or dCPU) that may be coupled to the SoC 1304 via a high-speed interconnect (e.g., PCIe). The CPU 1318 may include, for example, an X86 processor. The CPU 1318 may be used to perform any of a variety of functions, including, for example, arbitrating potentially inconsistent results between ADAS sensors and the SoC 1304, and / or monitoring the status and health of the controller 1336 and / or the infotainment SoC 1330.

[0208] The vehicle 1300 may include a GPU 1320 (e.g., a discrete GPU or dGPU) that may be coupled to the SoC 1304 via a high-speed interconnect (e.g., NVIDIA's NVLINK). The GPU 1320 may provide additional artificial intelligence capabilities, for example, by executing redundant and / or different neural networks, and may be used to train and / or update neural networks based on inputs (e.g., sensor data) from sensors of the vehicle 1300.

[0209] The vehicle 1300 may further include a network interface 1324, which may include one or more wireless antennas 1326 (e.g., one or more wireless antennas for different communication protocols, such as cellular antennas, Bluetooth antennas, etc.). The network interface 1324 may be used to enable wireless connections to the cloud (e.g., to the server 1378 and / or other network devices), to other vehicles, and / or to computing devices (e.g., the passenger's client device) via the Internet. To communicate with other vehicles, a direct link may be established between the two vehicles, and / or an indirect link may be established (e.g., across a network and via the Internet). The direct link may be provided using a vehicle-to-vehicle communication link. The vehicle-to-vehicle communication link may provide the vehicle 1300 with information about vehicles approaching the vehicle 1300 (e.g., vehicles in front of, to the side of, and / or behind the vehicle 1300). This function may be part of the cooperative adaptive cruise control function of the vehicle 1300.

[0210] The network interface 1324 may include an SoC that provides modulation and demodulation functions and enables the controller 1336 to communicate over a wireless network. The network interface 1324 may include a radio frequency front end for upconverting from baseband to radio frequency and downconverting from radio frequency to baseband. The frequency conversion may be performed by known processes, and / or may be performed using a super-heterodyne process. In some examples, the radio frequency front end functions may be provided by a separate chip. The network interface may include wireless capabilities for communicating via LTE, WCDMA, UMTS, GSM, CDMA2000, Bluetooth, Bluetooth LE, Wi-Fi, Z-Wave, ZigBee, LoRaWAN, and / or other wireless protocols.

[0211] Vehicle 1300 may further include a data store 1328 that may include off-chip (e.g., outside of SoC 1304) storage devices. The data store 1328 may include one or more storage elements, including RAM, SRAM, DRAM, VRAM, flash memory, hard drives, and / or other components and / or devices that can store at least one bit of data.

[0212] Vehicle 1300 may further include a GNSS sensor 1358 (e.g., GPS and / or assisted GPS sensors) for assisting mapping, perception, occupancy grid generation, and / or path planning functions. Any number of GNSS sensors 1358 may be used, including, for example and without limitation, a GPS using a USB connector with an Ethernet to serial (RS-232) bridge.

[0213] Vehicle 1300 may further include a RADAR sensor 1360. The RADAR sensor 1360 may be used by the vehicle 1300 for remote vehicle detection even in darkness and / or adverse weather conditions. The RADAR functional safety level may be ASIL B. The RADAR sensor 1360 may use CAN and / or bus 1302 (e.g., to transmit data generated by the RADAR sensor 1360) for control as well as access to object tracking data and, in some examples, access Ethernet to access raw data. A variety of RADAR sensor types may be used. For example and without limitation, the RADAR sensor 1360 may be suitable for front, rear, and side RADAR use. In some examples, a pulsed Doppler RADAR sensor is used.

[0214] The RADAR sensor 1360 may include different configurations, such as long-range with a narrow field of view, short-range with a wide field of view, short-range side coverage, etc. In some examples, long-range RADAR may be used for adaptive cruise control functions. The long-range RADAR system may provide a wide field of view (e.g., within 250m) achieved through two or more independent scans. The RADAR sensor 1360 may help distinguish between static and moving objects and may be used by the ADAS system for emergency braking assistance and forward collision warning. The long-range RADAR sensor may include a single station multi-mode RADAR with multiple (e.g., six or more) fixed RADAR antennas and high-speed CAN and FlexRay interfaces. In an example with six antennas, the central four antennas may create a focused beam pattern that is designed to record the surroundings of the vehicle 1300 at a higher rate with minimal traffic interference from adjacent lanes. The other two antennas may extend the field of view, making it possible to quickly detect vehicles entering or leaving the lane of the vehicle 1300.

[0215] As an example, a mid-range RADAR system can include a range of up to 960 m (front) or 80 m (rear) and a field of view of up to 42 degrees (front) or 950 degrees (rear). A short-range RADAR system can include, but is not limited to, RADAR sensors designed to be mounted at both ends of the rear bumper. When mounted at both ends of the rear bumper, such a RADAR sensor system can create two beams that continuously monitor the rear and the blind spots alongside the vehicle.

[0216] The short-range RADAR system can be used in an ADAS system for blind spot detection and / or lane change assistance.

[0217] Vehicle 1300 can further include ultrasonic sensors 1362. Ultrasonic sensors 1362 that can be placed in the front, rear, and / or sides of vehicle 1300 can be used for parking assistance and / or creating and updating an occupancy grid. A variety of ultrasonic sensors 1362 can be used, and different ultrasonic sensors 1362 can be used for different detection ranges (e.g., 2.5 m, 4 m). Ultrasonic sensors 1362 can operate at ASIL B for functional safety levels.

[0218] Vehicle 1300 can include a LIDAR sensor 1364. The LIDAR sensor 1364 can be used for object and pedestrian detection, emergency braking, collision avoidance, and / or other functions. The LIDAR sensor 1364 can be at ASIL B for functional safety levels. In some examples, vehicle 1300 can include multiple LIDAR sensors 1364 (e.g., two, four, six, etc.) that can use Ethernet (e.g., to provide data to a gigabit Ethernet switch).

[0219] In some examples, the LIDAR sensor 1364 may be able to provide a list of objects and their distances for a 360-degree field of view. Commercially available LIDAR sensors 1364 can have, for example, an advertised range of approximately 100 m, an accuracy of 2 cm - 3 cm, and support for a 100 Mbps Ethernet connection. In some examples, one or more non-protruding LIDAR sensors 1364 can be used. In such examples, the LIDAR sensor 1364 can be implemented as a small device that can be embedded in the front, rear, sides, and / or corners of vehicle 1300. In such examples, the LIDAR sensor 1364 can provide a field of view of up to 120 degrees horizontally and 35 degrees vertically, with a range of 200 m, even for low-reflectivity objects. The front-mounted LIDAR sensor 1364 can be configured for a horizontal field of view between 45 degrees and 135 degrees.

[0220] In some examples, LIDAR technologies such as 3D flash LIDAR can also be used. 3D flash LIDAR uses the flash of a laser as the emission source to illuminate the vehicle's surroundings up to about 200m. The flash LIDAR unit includes a receiver that records the laser pulse transit time and reflected light on each pixel, which in turn corresponds to the range from the vehicle to the object. Flash LIDAR can allow for the generation of highly accurate and distortion-free images of the surroundings using each laser flash. In some examples, four flash LIDAR sensors can be deployed, one on each side of the vehicle 1300. Available 3D flash LIDAR systems include solid-state 3D staring array LIDAR cameras (e.g., non-scanning LIDAR devices) that have no moving parts other than a fan. The flash LIDAR device can use Class I (eye-safe) laser pulses of 5 nanoseconds per frame and can capture the reflected laser in the form of 3D range point clouds and co-registered intensity data. By using flash LIDAR, and because flash LIDAR is a solid-state device with no moving parts, the LIDAR sensor 1364 can be less susceptible to motion blur, vibration, and / or shock.

[0221] The vehicle can further include an IMU sensor 1366. In some examples, the IMU sensor 1366 can be located at the center of the rear axle of the vehicle 1300. The IMU sensor 1366 can include, for example and without limitation, accelerometers, magnetometers, gyroscopes, magnetic compasses, and / or other sensor types. In some examples, such as in six-axis applications, the IMU sensor 1366 can include accelerometers and gyroscopes, while in nine-axis applications, the IMU sensor 1366 can include accelerometers, gyroscopes, and magnetometers.

[0222] In some embodiments, the IMU sensor 1366 can be implemented as a miniature high-performance GPS-aided inertial navigation system (GPS / INS) that combines microelectromechanical systems (MEMS) inertial sensors, a high-sensitivity GPS receiver, and advanced Kalman filtering algorithms to provide estimates of position, velocity, and attitude. Thus, in some examples, the IMU sensor 1366 can enable the vehicle 1300 to estimate the heading without input from a magnetic sensor by directly observing the change in velocity from the GPS to the IMU sensor 1366 and correlating it. In some examples, the IMU sensor 1366 and the GNSS sensor 1358 can be combined into a single integrated unit.

[0223] The vehicle can include a microphone 1396 placed in and / or around the vehicle 1300. Among other things, the microphone 1396 can be used for emergency vehicle detection and identification.

[0224] The vehicle may further include any number of camera types, including a stereo camera 1368, a wide-angle camera 1370, an infrared camera 1372, a surround camera 1374, a long-range and / or mid-range camera 1398, and / or other camera types. These cameras can be used to capture image data around the entire periphery of the vehicle 1300. The camera types used depend on the embodiment and the requirements of the vehicle 1300, and any combination of camera types can be used to provide the necessary coverage around the vehicle 1300. Additionally, the number of cameras can vary according to the embodiment. For example, the vehicle may include six cameras, seven cameras, ten cameras, twelve cameras, and / or another number of cameras. As an example and without limitation, these cameras can support Gigabit Multimedia Serial Link (GMSL) and / or Gigabit Ethernet. Each of the cameras is described in more detail herein with respect to Figure 13A and Figure 13B is described in more detail.

[0225] The vehicle 1300 may further include a vibration sensor 1342. The vibration sensor 1342 can measure the vibration of components of the vehicle such as an axle. For example, a change in vibration can indicate a change in the road surface. In another example, when two or more vibration sensors 1342 are used, the difference between the vibrations can be used to determine the friction or slip of the road surface (e.g., when there is a vibration difference between a powered drive axle and a free-rotating axle).

[0226] The vehicle 1300 may include an ADAS system 1338. In some examples, the ADAS system 1338 may include a SoC. The ADAS system 1338 may include autonomous / adaptive / auto cruise control (ACC), cooperative adaptive cruise control (CACC), forward collision warning (FCW), automatic emergency braking (AEB), lane departure warning (LDW), lane keeping assist (LKA), blind spot warning (BSW), rear cross-traffic warning (RCTW), collision warning system (CWS), lane centering (LC), and / or other features and functions.

[0227] The ACC system can use RADAR sensors 1360, LIDAR sensors 1364, and / or cameras. The ACC system can include longitudinal ACC and / or lateral ACC. Longitudinal ACC monitors and controls the distance to the vehicle immediately in front of the vehicle 1300 and automatically adjusts the vehicle speed to maintain a safe distance from the vehicle ahead. Lateral ACC performs distance keeping and, when necessary, advises the vehicle 1300 to change lanes. Lateral ACC is related to other ADAS applications such as LC and CWS.

[0228] The CACC uses information from other vehicles, which can be received indirectly from other vehicles via the network interface 1324 and / or the wireless antenna 1326 via a wireless link or through a network connection (e.g., via the Internet). The direct link can be provided by a vehicle-to-vehicle (V2V) communication link, while the indirect link can be an infrastructure-to-vehicle (I2V) communication link. Generally, the V2V communication concept provides information about the immediately preceding vehicle (e.g., the vehicle immediately in front of vehicle 1300 and in the same lane as it), while the I2V communication concept provides information about traffic further ahead. The CACC system can include either or both of the I2V and V2V information sources. Given the information of the vehicle in front of vehicle 1300, the CACC can be more reliable, and it has the potential to improve the smoothness of traffic flow and reduce road congestion.

[0229] The FCW system is designed to alert the driver to a hazard so that the driver can take corrective action. The FCW system uses a front camera and / or RADAR sensor 1360 coupled to a dedicated processor, DSP, FPGA, and / or ASIC, which is electrically coupled to driver feedback such as a display, speaker, and / or vibrating component. The FCW system can provide warnings in the form of, for example, audible, visual warnings, vibrations, and / or rapid braking pulses.

[0230] The AEB system detects an impending forward collision with another vehicle or other object and can automatically apply the brakes if the driver does not take corrective action within a specified time or distance parameter. The AEB system can use a front camera and / or RADAR sensor 1360 coupled to a dedicated processor, DSP, FPGA, and / or ASIC. When the AEB system detects a hazard, it typically first alerts the driver to take corrective action to avoid the collision, and if the driver does not take corrective action, then the AEB system can automatically apply the brakes in an effort to prevent or at least mitigate the impact of the predicted collision. The AEB system can include technologies such as dynamic brake support and / or collision imminent braking.

[0231] The LDW system provides visual, audible, and / or tactile warnings such as steering wheel or seat vibrations to alert the driver when vehicle 1300 crosses a lane marking. When the driver indicates an intentional lane departure by activating the turn signal, the LDW system is not activated. The LDW system can use a front-side-facing camera coupled to a dedicated processor, DSP, FPGA, and / or ASIC, which is electrically coupled to driver feedback such as a display, speaker, and / or vibrating component.

[0232] The LKA system is a variant of the LDW system. If vehicle 1300 starts to leave the lane, then the LKA system provides steering input or braking to correct the vehicle 1300.

[0233] The BSW system detects and warns the driver of vehicles in the vehicle's blind spot. The BSW system can provide visual, audible, and / or tactile alerts to indicate that merging or changing lanes is unsafe. The system can provide additional warnings when the driver uses the turn signal. The BSW system can use a rear-facing camera and / or RADAR sensor 1360 coupled to a dedicated processor, DSP, FPGA, and / or ASIC, which is electrically coupled to driver feedback such as a display, speaker, and / or vibrating component.

[0234] The RCTW system can provide visual, audible, and / or tactile notifications when an object is detected outside the rear camera range while the vehicle 1300 is in reverse. Some RCTW systems include AEB to ensure that the vehicle brakes are applied to avoid a crash. The RCTW system can use one or more rear RADAR sensors 1360 coupled to a dedicated processor, DSP, FPGA, and / or ASIC, which is electrically coupled to driver feedback such as a display, speaker, and / or vibrating component.

[0235] Conventional ADAS systems may be prone to false positive results, which can be annoying and distracting to the driver, but are typically not catastrophic because the ADAS system alerts the driver and allows the driver to decide whether a safe condition truly exists and act accordingly. However, in an autonomous vehicle 1300, in the case of conflicting results, the vehicle 1300 itself must decide whether to heed the results from the main computer or an auxiliary computer (e.g., the first controller 1336 or the second controller 1336). For example, in some embodiments, the ADAS system 1338 can be a backup and / or auxiliary computer for providing perception information to a backup computer rationality module. The backup computer rationality monitor can run redundant and diverse software on hardware components to detect faults in perception and dynamic driving tasks. The output from the ADAS system 1338 can be provided to the supervisory MCU. If the outputs from the main computer and the auxiliary computer conflict, then the supervisory MCU must determine how to reconcile the conflict to ensure safe operation.

[0236] In some examples, the host computer may be configured to provide a confidence score to the supervisory MCU indicating the host computer's confidence in the selected result. If the confidence score exceeds a threshold, then the supervisory MCU may follow the direction of the host computer regardless of whether the secondary computer provides conflicting or inconsistent results. In cases where the confidence score does not meet the threshold and where the host computer and the secondary computer indicate different results (e.g., conflict), the supervisory MCU may arbitrate between these computers to determine an appropriate result.

[0237] The supervisory MCU may be configured to run a neural network that is trained and configured to determine conditions under which the secondary computer provides a false alarm based on outputs from the host computer and the secondary computer. Thus, the neural network in the supervisory MCU can learn when the output of the secondary computer can be trusted and when it cannot. For example, when the secondary computer is a RADAR-based FCW system, the neural network in the supervisory MCU can learn when the FCW system is identifying a metallic object that is not in fact dangerous, such as a drain grate or manhole cover that triggers an alarm. Similarly, when the secondary computer is a camera-based LDW system, the neural network in the supervisory MCU can learn to disregard the LDW when a cyclist or pedestrian is present and lane departure is actually the safest strategy. In embodiments including a neural network running on the supervisory MCU, the supervisory MCU may include at least one of a DLA or a GPU suitable for running the neural network with associated memory. In a preferred embodiment, the supervisory MCU may include components of SoC 1304 and / or be included as a component of SoC 1304.

[0238] In other examples, the ADAS system 1338 may include a secondary computer that performs ADAS functions using traditional computer vision rules. In this way, the secondary computer may use classical computer vision rules (if-then), and the presence of a neural network in the supervisory MCU can improve reliability, safety, and performance. For example, diverse implementations and intentional non-identity make the overall system more fault-tolerant, especially for failures caused by software (or software-hardware interface) functions. For example, if there is a software vulnerability or error in the software running on the host computer and the non-identical software code running on the secondary computer provides the same overall result, then the supervisory MCU can be more confident that the overall result is correct and that the vulnerability in the software or hardware used by the host computer does not cause a substantial error.

[0239] In some examples, the output of the ADAS system 1338 can be fed to the perception block of the host computer and / or the dynamic driving task block of the host computer. For example, if the ADAS system 1338 indicates a forward collision warning due to an object being immediately in front, the perception block can use this information when identifying the object. In other examples, the auxiliary computer can have its own neural network, which is trained and thus reduces the risk of false positives as described herein.

[0240] The vehicle 1300 can further include an infotainment SoC 1330 (e.g., an in-vehicle infotainment system (IVI)). Although illustrated and described as an SoC, the infotainment system can not be an SoC and can include two or more discrete components. The infotainment SoC 1330 can include a combination of hardware and software that can be used to provide audio (e.g., music, personal digital assistant, navigation instructions, news, radio, etc.), video (e.g., TV, movies, streaming, etc.), telephone (e.g., hands-free calling), network connectivity (e.g., LTE, Wi-Fi, etc.), and / or information services (e.g., navigation system, rear parking assistance, radio data system, vehicle-related information such as fuel level, total distance covered, brake fuel level, oil level, door open / close, air filter information, etc.) to the vehicle 1300. For example, the infotainment SoC 1330 can include a radio, a disc player, a navigation system, a video player, USB and Bluetooth connectivity, an in-vehicle computer, in-vehicle entertainment, Wi-Fi, steering wheel audio controls, hands-free voice controls, a head-up display (HUD), an HMI display 1334, a telematics device, a control panel (e.g., for controlling various components, features, and / or systems, and / or interacting therewith), and / or other components. The infotainment SoC 1330 can further be used to provide information (e.g., visual and / or auditory) to the vehicle's user, such as information from the ADAS system 1338, autonomous driving information such as planned vehicle maneuvers, trajectories, surrounding environment information (e.g., intersection information, vehicle information, road information, etc.), and / or other information.

[0241] The infotainment SoC 1330 can include GPU functionality. The infotainment SoC 1330 can communicate with other devices, systems, and / or components of the vehicle 1300 via a bus 1302 (e.g., a CAN bus, Ethernet, etc.). In some examples, the infotainment SoC 1330 can be coupled to a supervisory MCU such that in the event of a failure of the main controller 1336 (e.g., the main and / or backup computer of the vehicle 1300), the GPU of the infotainment system can perform some self-driving functions. In such examples, the infotainment SoC 1330 can place the vehicle 1300 in the driver safe parking mode as described herein.

[0242] Vehicle 1300 may further include an instrument cluster 1332 (such as a digital instrument panel, an electronic instrument cluster, a digital instrument panel, etc.). The instrument cluster 1332 may include a controller and / or a supercomputer (such as a discrete controller or supercomputer). The instrument cluster 1332 may include a set of instruments, such as a speedometer, fuel level, oil pressure, tachometer, odometer, turn indicator, shift position indicator, seat belt warning light, parking brake warning light, engine malfunction light, airbag (SRS) system information, lighting controls, safety system controls, navigation information, etc. In some examples, the information may be displayed and / or shared between the infotainment SoC 1330 and the instrument cluster 1332. In other words, the instrument cluster 1332 may be included as part of the infotainment SoC 1330, or vice versa.

[0243] Figure 13D A system schematic for communication between a cloud-based server and Figure 13A an example autonomous vehicle 1300 in accordance with some embodiments of the present disclosure. The system 1376 may include a server 1378, a network 1390, and vehicles including the vehicle 1300. The server 1378 may include multiple GPUs 1384(A)-1384(H) (collectively referred to herein as GPUs 1384), PCIe switches 1382(A)-1382(H) (collectively referred to herein as PCIe switches 1382), and / or CPUs 1380(A)-1380(B) (collectively referred to herein as CPUs 1380). The GPUs 1384, CPUs 1380, and PCIe switches may be interconnected by high-speed interconnects such as, for example, and without limitation, an NVLink interface 1388 developed by NVIDIA and / or a PCIe connection 1386. In some examples, the GPUs 1384 are connected via NVLink and / or an NVSwitch SoC, and the GPUs 1384 and the PCIe switches 1382 are connected via a PCIe interconnect. Although eight GPUs 1384, two CPUs 1380, and two PCIe switches are illustrated, this is not intended to be limiting. Depending on the embodiment, each of the servers 1378 may include any number of GPUs 1384, CPUs 1380, and / or PCIe switches. For example, each of the servers 1378 may include eight, sixteen, thirty-two, and / or more GPUs 1384.

[0244] Server 1378 can receive image data from a vehicle via network 1390, the image data representing an image showing an unexpected or changed road condition such as a recently started road work. Server 1378 can transmit neural network 1392, updated neural network 1392, and / or map information 1394, including information about traffic and road conditions, to the vehicle via network 1390. Updates to map information 1394 can include updates to HD map 1322, such as information about construction sites, potholes, curves, floods, or other obstacles. In some examples, neural network 1392, updated neural network 1392, and / or map information 1394 can be represented and / or generated based on data received from new training and / or from any number of vehicles in the environment and / or experience of training performed at a data center (e.g., using server 1378 and / or other servers).

[0245] Server 1378 can be used to train a machine learning model (e.g., a neural network) based on training data. The training data can be generated by vehicles and / or can be generated in a simulation (e.g., using a game engine). In some examples, the training data is labeled (e.g., in cases where the neural network benefits from supervised learning) and / or undergoes other preprocessing, while in other examples, the training data is not labeled and / or preprocessed (e.g., in cases where the neural network does not require supervised learning). Training can be performed according to any one or more categories of machine learning techniques, including but not limited to: categories such as supervised training, semi-supervised training, unsupervised training, self-learning, reinforcement learning, joint learning, transfer learning, feature learning (including principal component and clustering analysis), multilinear subspace learning, manifold learning, representation learning (including alternative dictionary learning), rule-based machine learning, anomaly detection, and any variations or combinations thereof. Once the machine learning model is trained, the machine learning model can be used by the vehicle (e.g., transmitted to the vehicle via network 1390), and / or the machine learning model can be used by server 1378 to remotely monitor the vehicle.

[0246] In some examples, server 1378 can receive data from a vehicle and apply the data to a latest real-time neural network for real-time intelligent inference. Server 1378 can include a deep learning supercomputer powered by GPU 1384 and / or a dedicated AI computer, such as DGX and DGX Station machines developed by NVIDIA. However, in some examples, server 1378 can include a deep learning infrastructure of a data center powered only by a CPU.

[0247] The deep learning infrastructure of server 1378 may be capable of fast real-time inference and can use this ability to evaluate and verify the health of the processors, software, and / or associated hardware in vehicle 1300. For example, the deep learning infrastructure can receive periodic updates from vehicle 1300, such as an image sequence and / or objects located in the image sequence that vehicle 1300 has identified (e.g., via computer vision and / or other machine learning object classification techniques). The deep learning infrastructure can run its own neural network to identify the objects and compare them with the objects identified by vehicle 1300. If the results do not match and the infrastructure concludes that the AI in vehicle 1300 has failed, then server 1378 can transmit a signal to vehicle 1300, instructing the fail-safe computer in vehicle 1300 to take control, notify the passengers, and complete a safe parking operation.

[0248] For inference, server 1378 may include GPU 1384 and one or more programmable inference accelerators (e.g., NVIDIA's TensorRT). The combination of a GPU-powered server and inference acceleration can enable real-time response. In other examples, such as when performance is less critical, CPU, FPGA, and other processor-powered servers can be used for inference.

[0249] Example computing device

[0250] Figure 14 A block diagram of an example computing device 1400 suitable for implementing some embodiments of the present disclosure. Computing device 1400 may include an interconnection system 1402 that directly or indirectly couples the following devices: a memory 1404, one or more central processing units (CPUs) 1406, one or more graphics processing units (GPUs) 1408, a communication interface 1410, I / O ports 1412, input / output components 1414, a power supply 1416, one or more presentation components 1418 (e.g., a display), and one or more logic units 1420.

[0251] Although Figure 14 the various boxes are shown as being connected via an interconnection system 1402 with lines, this is not intended to be restrictive and is for clarity only. For example, in some embodiments, a presentation component 1418, such as a display device, may be considered an I / O component 1414 (e.g., if the display is a touchscreen). As another example, CPU 1406 and / or GPU 1408 may include memory (e.g., memory 1404 may represent a storage device in addition to the memory of GPU 1408, CPU 1406, and / or other components). In other words, Figure 14The computing device is merely illustrative. No distinction is made between categories such as "workstation", "server", "laptop", "desktop", "tablet", "client device", "mobile device", "handheld device", "gaming console", "electronic control unit (ECU)", "virtual reality system", "augmented reality system", and / or other device or system types, as all of these are considered within the scope of Figure 14 the computing device.

[0252] The interconnect system 1402 may represent one or more links or buses, such as an address bus, a data bus, a control bus, or a combination thereof. The interconnect system 1402 may include one or more types of buses or links, such as an Industry Standard Architecture (ISA) bus, an Extended Industry Standard Architecture (EISA) bus, a Video Electronics Standards Association (VESA) bus, a Peripheral Component Interconnect (PCI) bus, a Peripheral Component Interconnect Express (PCIe) bus, and / or another type of bus or link. In some embodiments, there are direct connections between components. For example, the CPU 1406 may be directly connected to the memory 1404. Additionally, the CPU 1406 may be directly connected to the GPU 1408. In cases where there are direct or point-to-point connections between components, the interconnect system 1402 may include a PCIe link to effect the connection. In these examples, a PCI bus need not be included in the computing device 1400.

[0253] The memory 1404 may include any of a variety of computer-readable media. Computer-readable media can be any available media that can be accessed by the computing device 1400. Computer-readable media can include volatile and non-volatile media as well as removable and non-removable media. By way of example and not limitation, computer-readable media can include computer storage media and communication media.

[0254] Computer storage media can include volatile and non-volatile media and / or removable and non-removable media, implemented in any method or technology for storing information such as computer-readable instructions, data structures, program modules, and / or other data types. For example, the memory 1404 may store computer-readable instructions (e.g., which represent programs and / or program elements, such as an operating system). Computer storage media can include, but is not limited to, RAM, ROM, EEPROM, flash memory, or other storage technologies, CD-ROM, digital versatile disk (DVD), or other optical disk storage, magnetic cassettes, magnetic tape, magnetic disk storage, or other magnetic storage devices, or any other medium that can be used to store the desired information and that can be accessed by the computing device 1400. As used herein, computer storage media does not include signals per se.

[0255] A computer storage medium can include computer-readable instructions, data structures, program modules, and / or other data types in a modulated data signal such as a carrier wave or other transmission mechanism, and includes any information conveyance medium. The term "modulated data signal" can refer to a signal that has one or more of its characteristics set or changed in such a manner as to encode information in the signal. By way of example, and not limitation, a computer storage medium can include wired media such as a wired network or direct wired connection, and wireless media such as sound, RF, infrared, and other wireless media. Any combination of the foregoing should also be included within the scope of computer-readable media.

[0256] The CPU 1406 can be configured to execute at least some of the computer-readable instructions to control one or more components of the computing device 1400 to perform one or more of the methods and / or processes described herein. Each of the CPUs 1406 can include one or more cores (e.g., one, two, four, eight, twenty-eight, seventy-two, etc.) capable of simultaneously processing a large number of software threads. The CPU 1406 can include any type of processor and can include different types of processors depending on the type of computing device 1400 being implemented (e.g., a processor with fewer cores for a mobile device and a processor with more cores for a server). For example, depending on the type of computing device 1400, the processor can be an Advanced RISC Machine (ARM) processor implemented using Reduced Instruction Set Computing (RISC) or an x86 processor implemented using Complex Instruction Set Computing (CISC). In addition to one or more microprocessors or complementary coprocessors such as a math coprocessor, the computing device 1400 can also include one or more CPUs 1406.

[0257] In addition to or instead of the CPU 1406, one or more GPUs 1408 may be configured to execute at least some of the computer-readable instructions to control one or more components of the computing device 1400 to perform one or more of the methods and / or processes described herein. One or more GPUs 1408 may be integrated GPUs (e.g., having one or more CPUs 1406) and / or one or more GPUs 1408 may be discrete GPUs. In an embodiment, one or more GPUs 1408 may be coprocessors of one or more CPUs 1406. The computing device 1400 may use the GPUs 1408 to render graphics (e.g., 3D graphics) or perform general computing. For example, one or more GPUs 1408 may be used for general-purpose computing on GPUs (GPGPU). One or more GPUs 1408 may include hundreds or thousands of cores capable of processing hundreds or thousands of software threads simultaneously. The GPUs 1408 may generate pixel data for an output image in response to a rendering command (e.g., a rendering command received from the CPU 1406 via a host interface). The GPUs 1408 may include graphics memory, such as display memory, for storing pixel data or any other suitable data, such as GPGPU data. The display memory may be included as part of the memory 1404. One or more GPUs 1408 may include two or more GPUs operating in parallel (e.g., via a link). The link may directly connect the GPUs (e.g., using NVLINK) or may connect the GPUs through a switch (e.g., using NVSwitch). When combined, each GPU 1408 may generate pixel data or GPGPU data for different portions of the output or different outputs (e.g., the first GPU for the first image and the second GPU for the second image). Each GPU may include its own memory or may share memory with other GPUs.

[0258] In addition to and / or in lieu of the CPU 1406 and / or GPU 1408, the logic unit 1420 may be configured to execute at least some of the computer-readable instructions to control one or more components of the computing device 1400 to perform one or more of the methods and / or processes described herein. In an embodiment, the CPU 1406, GPU 1408, and / or logic unit 1420 may perform any combination of methods, processes, and / or portions thereof, separately or jointly. One or more logic units 1420 may be part of and / or integrated in one or more of the CPU 1406 and / or GPU 1408, and / or one or more logic units 1420 may be discrete components or otherwise outside of the CPU 1406 and / or GPU 1408. In an embodiment, one or more logic units 1420 may be a coprocessor of one or more CPUs 1406 and / or one or more GPUs 1408.

[0259] Examples of the logic unit 1420 include one or more processing cores and / or their components, such as tensor cores (TC), tensor processing units (TPU), pixel vision cores (PVC), visual processing units (VPU), graphics processing clusters (GPC), texture processing clusters (TPC), streaming multiprocessors (SM), tree traversal units (TTU), artificial intelligence accelerators (AIA), deep learning accelerators (DLA), arithmetic logic units (ALU), application specific integrated circuits (ASIC), floating point units (FPU), I / O elements, peripheral component interconnect (PCI) or peripheral component interconnect express (PCIe) elements, etc.

[0260] The communication interface 1410 may include one or more receivers, transmitters, and / or transceivers that enable the computing device 1400 to communicate with other computing devices via an electronic communication network, including wired and / or wireless communication. The communication interface 1410 may include components and functions that enable communication over any of several different networks, such as wireless networks (e.g., Wi-Fi, Z-Wave, Bluetooth, Bluetooth LE, ZigBee, etc.), wired networks (e.g., communication via Ethernet or InfiniBand), low power wide area networks (e.g., LoRaWAN, SigFox, etc.), and / or the Internet.

[0261] The I / O port 1412 can enable the computing device 1400 to be logically coupled to other devices including I / O components 1414, presentation components 1418, and / or other components, some of which may be built into (e.g., integrated into) the computing device 1400. Exemplary I / O components 1414 include microphones, mice, keyboards, joysticks, game pads, game controllers, dish satellite antennas, scanners, printers, wireless devices, and so on. The I / O components 1414 can provide a natural user interface (NUI) that processes user-generated air gestures, voice, or other physiological inputs. In some instances, the input can be transmitted to appropriate network elements for further processing. The NUI can implement any combination of speech recognition, stylus recognition, face recognition, biometric recognition, on-screen and near-screen gesture recognition, air gestures, head and eye tracking, and touch recognition associated with the display of the computing device 1400 (described in more detail below). The computing device 1400 can include depth cameras such as stereo camera systems, infrared camera systems, RGB camera systems, touch screen technologies, and combinations thereof for gesture detection and recognition. Additionally, the computing device 1400 can include an accelerometer or gyroscope that enables motion detection (e.g., as part of an inertial measurement unit (IMU)). In some examples, the output of the accelerometer or gyroscope can be used by the computing device 1400 to render immersive augmented reality or virtual reality.

[0262] The power supply 1416 can include a hard-wired power supply, a battery power supply, or a combination thereof. The power supply 1416 can power the computing device 1400 so that the components of the computing device 1400 can operate.

[0263] The presentation component 1418 can include a display (e.g., a monitor, a touch screen, a television screen, a head-up display (HUD), other display types, or a combination thereof), speakers, and / or other presentation components. The presentation component 1418 can receive data from other components (e.g., the GPU 1408, the CPU 1406, etc.) and output the data (e.g., as images, videos, sounds, etc.).

[0264] This disclosure can be described in the general context of machine-usable instructions or computer code, including computer-executable instructions such as program modules, executed by a computer or other machines such as personal digital assistants or other handheld devices. Generally, program modules, including routines, programs, objects, components, data structures, etc., refer to code that performs specific tasks or implements specific abstract data types. This disclosure can be practiced in a variety of system configurations, including handheld devices, consumer electronics, general-purpose computers, more specialized computing devices, and so on. This disclosure can also be practiced in a distributed computing environment where tasks are performed by remote processing devices linked through a communication network.

[0265] As used herein, the recitation of "and / or" with respect to two or more elements shall be construed to mean only one element or a combination of elements. For example, "element A, element B, and / or element C" may include only element A, only element B, only element C, element A and element B, element A and element C, element B and element C, or elements A, B, and C. Further, "at least one of element A or element B" may include at least one of element A, at least one of element B, or at least one of element A and at least one of element B. Still further, "at least one of element A and element B" may include at least one of element A, at least one of element B, or at least one of element A and at least one of element B.

[0266] The subject matter of the present disclosure is described in detail herein to meet statutory requirements. However, the description itself is not intended to limit the scope of the present disclosure. On the contrary, the inventors have contemplated that the claimed subject matter may also be embodied in other ways, including steps different from or combinations of steps similar to those described herein in connection with other current or future technologies. Moreover, although the terms "step" and / or "block" may be used herein to imply different elements of a method employed, these terms should not be construed as implying any particular order among or between the various steps disclosed herein unless the order of the steps is expressly recited.

Claims

1. A method, comprising: Receiving an image representing a subject and a restraint device corresponding to the subject; Performing an initial classification of a region of the image representing the restraint device, at least in part based on a plurality of pixels including at least one edge of the restraint device; Applying a set of constraints to pixel candidates in the region of the image to refine the pixel candidates to obtain refined pixel candidates; And Generating a model indicating the application of the restraint device, at least in part based on the refined pixel candidates.

2. The method according to claim 1, wherein the initial classification is performed by using pixels along a specified direction and a plurality of adjacent pixels to determine a set of pixels that are part of the restraint device.

3. The method according to claim 2, wherein determining the set of pixels that are part of the restraint device is at least in part based on intensity levels of the pixels along the specified direction and the plurality of adjacent pixels.

4. The method according to claim 1, wherein the constraint device includes one or more attributes, and the one or more attributes at least include: The width size of the restraint device and one or more anchors of the restraint device in a vehicle.

5. The method according to claim 1, wherein the image is captured by a camera comprising one or more configurations, the one or more configurations comprising at least: The calibration of the camera, the camera pose, and the type of camera lens.

6. The method according to claim 1, further comprising: Activating a signal to indicate that the restraint device is in an inappropriate position relative to the subject.

7. The method according to claim 1, further comprising: Using the model to determine the position of the restraint device relative to the subject, wherein the model approximates the shape of the restraint device.

8. A system, comprising: One or more processors; And A memory storing computer-executable instructions that can be executed by the one or more processors to cause the system to: Perform an initial classification of a region of an image representing the restraint device applied to a subject, at least in part based on a plurality of pixels including at least one edge of the restraint device; Apply a set of constraints to pixel candidates in the region of the image to refine the pixel candidates to obtain refined pixel candidates; And Generate a model indicating the application of the restraint device, at least in part based on the refined pixel candidates.

9. The system according to claim 8, wherein the model includes a polynomial curve having coefficients calculated to match the shape of the restraint device.

10. The system according to claim 8, wherein: The system includes a parallel processing unit PPU; The image includes one or more regions; and Classifying the one or more regions in parallel using one or more threads of the PPU.

11. The system according to claim 8, wherein the model indicates whether the restraint device is correctly applied to the subject.

12. The system according to claim 8, wherein the set of constraints includes at least a width range.

13. The system according to claim 8, wherein the region of the image corresponds to a grouping of one or more pixels of the image.

14. The system according to claim 8, wherein the image is obtained from one or more networks of one or more systems in a vehicle.

15. The system according to claim 14, wherein the model of the restraint device is provided to one or more systems of the vehicle via the one or more networks.

16. A vehicle, comprising: A propulsion system; An image capture device capable of capturing an image of at least one passenger of the vehicle; And A computer system comprising instructions executable by the computer system to at least: Access an image representing a subject and a restraint device applied to the subject; Perform an initial classification of a region of the image representing the restraint device, at least in part based on a plurality of pixels including at least one edge of the restraint device; Apply a set of constraints to pixel candidates in the region of the image to refine the pixel candidates to obtain refined pixel candidates; Generate a model indicating the application of the restraint device, at least in part based on the refined pixel candidates; And Control the function of at least one subsystem of the vehicle, at least in part based on the model.

17. The vehicle according to claim 16, wherein the restraint device is a seat belt of the vehicle.

18. The vehicle according to claim 16, wherein one subsystem of the vehicle is a warning system indicating whether the restraint device is worn and correctly applied.

19. The vehicle according to claim 16, wherein one subsystem of the vehicle sends a signal to control the propulsion of the vehicle via the propulsion system.

20. The vehicle according to claim 16, wherein one subsystem of the vehicle transmits an indication of whether the restraint device is worn and correctly applied to one or more remote systems.

Citation Information

Patent Citations

  • Method for programmable timeouts of tree traversal mechanisms in hardware

    US10885698B2

  • Neural network based determination of gaze direction using spatial models

    US11657263B2

  • Machine learning-based seatbelt detection and usage recognition using fiducial marking

    US12005855B2

  • Neural network based facial analysis using facial landmarks and associated confidence values

    US20210182625A1

  • Automobile safety belt detection method and automobile safety belt detection device

    CN104417490A