Deep learning-based 3D driver distraction detection system and method thereof

The deep learning-based 3D driver distraction detection system addresses the challenge of accurately detecting distances between a driver's head and distracting items by estimating 3D shapes and reconstructing surfaces, effectively meeting the standards for driver attention monitoring.

WO2025107112A1PCT designated stage expired Publication Date: 2025-05-30HARMAN INT IND INC +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2023/132639
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-11-20
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

Current driver distraction detection methods cannot accurately detect the distance between a driver's head and distracting items like phones or cigarettes, failing to meet the requirements of the Chinese standard GB/T 41797-2022.

Method used

A deep learning-based 3D driver distraction detection system that estimates the 3D shape of human heads and other items, predicts their distances, and reconstructs 3D surfaces to detect driver distractions effectively.

Benefits of technology

The system provides accurate 3D distance predictions between a driver's head and distracting items, enhancing the detection of driver distractions and meeting the requirements of GB/T 41797-2022.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2023132639_30052025_PF_FP_ABST
    Figure CN2023132639_30052025_PF_FP_ABST
Patent Text Reader

Abstract

The provided system and method for deep learning-based 3D detecting driver distraction comprises the steps of generating the feature maps and generating the 3D key points proposals. In the model training phase, plant of annotated single frame images (with labeled 3D key points coordinates (x, y, z) of each object, e.g., head, phone, etc. ). Then the surface of human head, phone, cigarette, etc. could be reconstructed and the distance between the human head and other objects such as phone, cigarette could be calculated. Based on those estimated distances, our OMS / DMS system can detect driver distractive behaviors such as phone calls, smoking and so on.
Need to check novelty before this filing date? Find Prior Art

Description

DEEP LEARNING-BASED 3D DRIVER DISTRACTION DETECTION SYSTEM AND METHOD THEREOFTECHNICAL FIELD

[0001] The inventive subject matter generally relates to image processing in driver assistance. More particularly, the inventive subject matter relates to a deep learning-based 3D driver distraction detection system and method.BACKGROUND

[0002] As a new C-NCAP requirement, GB / T 41797-2022, which is the Chinese standard of performance requirements and test methods for driver attention monitoring system which defines a phone call activity to be drive holding the phone which has less than 5cm distance to the human face, is going to come into practice. In this regard, the original equipment manufacturers (OEMs) in the automotive industry shall need to provide corresponding services to fulfill this requirement. This standard also specifies the testing procedure that requires the distance between the driver’s head and the phone.

[0003] However, the current classification method in the market for driver manual distractive detection cannot provide accurate distance detection, nor can it provide an effective solution that meets the requirement of GB / T 41797-2022.

[0004] In addition, some clients have also raised requirements for predicting the distance between the driver's head and other items, such as cigarettes, etc., but our current solution using classifiers cannot explicitly detect such smoking activities based on the distance detected such as between the driver’s head and cigarette, either.

[0005] Therefore, it is desirable to design an effective new detection solution that can estimate the 3D shape of human head and the other items, and predict the distance therebetween to meet these requirements.

[0006] SUMMARY OF THE INVENTIVE SUBJECT MATTER

[0007] The solution according to the inventive subject matter is described with respect to the claimed system for deep learning-based 3D detecting driver distraction as well as with respect to the claimed methods for deep learning-based 3D detecting driver distraction. This deep learning-based 3D detecting solution may estimate the 3D shape of the human head and the other items, such as phone or cigarettes, and predict the distance therebetween to meet the requirements of detecting driver distraction.

[0008] In one aspect, a system for deep learning-based 3D detecting driver distraction is provided. The system for deep learning-based 3D detecting driver distraction comprises a detector module, which comprises an encoder configured to extract features of an input image frame and cabin layout information, a decoder configured to generate at least one feature map and a detection head configured to classify each of at least two objects in the image frame and regress 3D key point coordinates of key points of said each of the at least two objects. The system for deep learning-based 3D detecting driver distraction further comprises a surface reconstruction module configured to reconstruct at least two 3D surfaces for the at least two objects, respectively, and a calculation module configured to calculate a 3D distance between the at least two 3D surfaces.

[0009] In another aspect, a method for deep learning-based 3D detecting driver distraction is provided. The method for deep learning-based 3D detecting driver distraction comprises the following steps of extracting, via an encoder of a detector module, features of an input image frame and cabin layout information, generating, via a decoder of the detector module, at least one feature map, and classifying, via a detection head of the detector module, each of at least two objects in the image frame and regressing 3D key point coordinates of key points of said each of the at least two objects. The method for deep learning-based 3D detecting driver distraction further comprises the steps of reconstructing, via a surface reconstruction module, at least two 3D surfaces for the at least two objects, respectively, and calculating, via a calculation module, a 3D distance between the at least two 3D surfaces.

[0010] In yet another aspect, a computer-readable medium storing instruction is provided that, when executed by one or more processors, may perform the steps of the method for deep learning-based 3D detecting driver distraction.BRIEF DESCRIPTION OF THE DRAWINGS

[0011] The present inventive subject matter may be better understood from reading the following description of non-limiting embodiments, with reference to the attached drawings. In the FIG. s, like reference numeral designates corresponding parts, wherein below:

[0012] FIG. 1 illustrates an exemplary structure diagram of the system for deep learning-based 3D detecting driver distraction, in accordance with the one or more embodiments of the inventive subject matter;

[0013] FIG. 2 illustrates an exemplary diagram of the detector module of the system for deep learning-based 3D detecting driver distraction, in accordance with the one or more embodiments of the inventive subject matter;

[0014] FIG. 3 illustrates an exemplary flowchart of the method for deep learning-based 3D detecting driver distraction, in accordance with the one or more embodiments of the inventive subject matter; and

[0015] FIG. 4 illustrates an exemplary process diagram of the method for deep learning-based 3D detecting driver distraction, in accordance with the one or more embodiments of the inventive subject matter.DETAILED DESCRIPTION

[0016] The detailed description of the one or more embodiments of the inventive subject matter is disclosed hereinafter; however, it is understood that the disclosed embodiments are merely exemplary of the inventive subject matter that may be embodied in various and alternative forms. The figures are not necessarily to scale; some features may be exaggerated or minimized to show details of particular components. Therefore, specific structural and function details disclosed herein are not to be interpreted as limiting, but merely as a representative basis for teaching one skilled in the art to variously employ the present inventive subject matter.

[0017] The 3D key point detection solution provided in the inventive subject matter may estimate the 3D shape of human head and other items, such as phone or cigarette, in a cabin of vehicle, and predict the distance between the head and the item to meet the requirement of GB / T 41797-2022.

[0018] In one aspect, the inventive subject matter provides a system for deep learning-based 3D detecting driver distraction.

[0019] FIG. 1 illustrates an exemplary structure diagram 100 of the system for deep learning-based 3D detecting driver distraction, in accordance with the one or more embodiments of the inventive subject matter.

[0020] As shown in Figure 1, the provided system comprises a detector module 110, a reconstruction module 120 and a calculation module 130. The detector module 110 may be of a form of neural network model. In the model training phase, plenty of annotated single frame images (with labeled 3D key point coordinates (x, y, z) of each object, e.g., head, phone, cigarette, etc. ) are used to train the neural network model in the detector module 110. The trained detector module 110 may output predicted 3D key points of human faces, phones or cigarettes. These identified 3D key points then may be input into the reconstruction module 120, where the 3D curved surfaces of the human head, the phone or the cigarette, etc. may be reconstructed, and thus the distance between the human head and the other object, such as the phone or the cigarette, can be calculated in the calculation module 130. Based on those estimated distances, the OMS / DMS system on the automobile may detect driver distractive behaviors, such as phone calls, smoking, and so on.

[0021] FIG. 2 illustrates an exemplary diagram 200 of the detector module of the system for deep learning-based 3D detecting driver distraction, in accordance with the one or more embodiments of the inventive subject matter. The pipeline of the detector module (which can be the detector module 110 shown in FIG. 1) includes three parts: an encoder 210 to extract features of images and cabin layout information 230, a decoder 240 to generate at least one feature map 250 for 3D object detection, and a detection head to classify each object and regress the 3D key point coordinates of their key points into feature vectors 260, 270, correspondingly.

[0022] As shown in the main components of the detector module of Fig. 2, firstly, the input for the encoder 210 includes the cabin layout information 230 for the cabin of the vehicle, which may comprise the dimensional information of the cabin, such as the length, width, and height inside the vehicle's cabin, the layout of seat positions, and the camera positioning in the cabin. Driver’s behaviors can be captured by the camera arranged, for example, for the Driver Monitor System (DMS) in the cabin. That is, the DMS camera may capture image frames 220 inside the cabin, which, along with the cabin layout information 230, can be input into the encoder 210 of the  detector module.

[0023] Due to the fact that the detector module is of a neural network model, which may have already been trained, it is possible for the detector module to classify and regress the points in the input image frame 230, according to their objectiveness probability score and classification probability score.

[0024] As mentioned above, the encoder 210 may extract features of images from the input image frame 220. Then, the decoder 240 may generate and output the at least one feature map 250 for 3D key points detection. In the feature map, each of the points in the input image frame may be described with an objectiveness probability score, obj, a classification probability score, cls, and its 3D coordinate, (x, y, z) . In an example, for a point in the input image frame 230, its probability of belonging to a certain object may be firstly determined. For the point with its objectiveness probability score higher than a threshold, its probability of belonging to any certain classification may be then determined. The point may either belong to the human head, or belong to mobile phones or cigarettes, and so on, depending on its classification probability score. Next, for those points in the input image frame with the objectiveness probability scores and the classification probability scores both higher than set thresholds, correspondingly, the coordinates (x, y, z) of these 3D key points will be collected, along with its objectiveness probability score and classification probability score, to generate feature maps. An input image frame may generate at least one feature map for 3D object detection.

[0025] The value of the objectiveness probability score and the value of the classification probability score both can be from 0 to 1, with a set threshold of 0.25, respectively, for example. For a point in the input image frame, if its values of the objectiveness probability score and of the classification probability score are higher than 0.25, respectively, the 3D key point coordinate of this key point will be regressed to the corresponding feature vector.

[0026] Exemplary feature vectors for 3D points detection can be as shown on the right of Fig. 2, which are extracted from the at least one feature map. The upper vector 260 is formed by those points extracted from the at least one feature map belonging to the human head classification. In the example, 3D key point coordinates, (x, y, z) , of those points with the value of the objectiveness probability score, obj, higher than 0.25 and the value of the classification probability score, cls, for human-head being higher than 0.25 may be regressed into the upper vector 260 for the human head. Similarly, the lower vector 270 is formed by those points extracted from the at least one  feature map belonging to the phone (or cigarette) classification. In the example, 3D key point coordinates, (x, y, z) , of those points with the value of the objectiveness probability score higher than 0.25 and the value of the phone (or cigarette) classification probability score higher than 0.25 may be regressed into the lower vector 270 for the phone (or cigarette) . The 3D key point coordinates, (x, y, z) , may be filled in the slots following the corresponding header of the corresponding feature vector, respectively. Alternatively, the value for the objectiveness probability score may be different from that for the classification probability score. It can be conceived that, the existence, classifications and 3D key point coordinates of other objects may also be contained in the at least one feature map, and their feature vectors can be extracted, either.

[0027] All the regression values (x, y, z) in each positive feature vector may construct the 3D key points of the human head, the phone, etc. in the expected order. In subsequent processing, the 3D key points regressed to the human head may be used to reconstruct a 3D curved surface of the partial human head, and the face, mouth, eye, etc. can be reconstructed on the surface. Similarly, the 3D key points regressed to the phone (or cigarette) may be used to reconstruct a 3D surface of the phone (or cigarette) . Thereby, after getting the 3D key point coordinates (x, y, z) output from the detector module 110, the 3D surface of each of the objects may get reconstructed. The 3D surface of the human mouth and other objects such as a phone, a cigarette can be reconstructed, respectively, in the reconstruction module 120. There are tons of potential robust algorithms to choose from for the surface reconstruction process, which is a mature area. For example, Power Crust, Poisson, and other surface reconstruction algorithms would be the proper ones to use. And eventually, the required distance from the driver’s head to the phone (or cigarette) for prediction can be calculated in the calculation module 130.

[0028] In another aspect, the inventive subject matter provides a method for deep learning-based 3D detecting driver distraction.

[0029] FIG. 3 illustrates an exemplary flowchart of the method for deep learning-based 3D detecting driver distraction, in accordance with the one or more embodiments of the inventive subject matter.

[0030] Firstly, the detector module takes a form of the neural network model, such as the detector module 110 shown in FIG. 1. The neural network training process includes feeding the neural network with paired cabin layout information and the input image frames. The cabin layout information can be obtained from factory information set for the vehicle type, and different  vehicle types can have different cabin layout information, such as the cabin sizes or dimensions including the length, width, height, etc. of the interior space, the positioning of the interior facilities, such as the front and rear seats, and the positioning of the DMS camera, in the cabin. The cabin layout information can be fixed for the same type of vehicles. Since different types may have different cabin layout or size parameters, the neural network model will eventually predict the 3D key points based on the information of the layout and size of the car cabin. Therefore, the neural network model shall be trained according to the vehicle type. The detector module needs to be trained only once for each of various known vehicle models, without the need to repeatedly train the neural network model of the detector module when the system to be used in a new known vehicle type. Only those newly designed unknown types of vehicle need to undergo the training steps. Therefore, in step S310, the type of the vehicle should be firstly determined.

[0031] For the training process for vehicle types unknown by the neural network model, in step S320, plenty of annotated single frame images (with labeled 3D key points coordinates (x, y, z) of each object, e.g., head, phone, confidential, etc. ) are used to train the neural network model of the detector module. The neural network model outputs the predicted 3D key points of the human face, the phone or the cigarette. Then, a loss function can be used to supervise the model to learn how to predict the 3D key points which is close to the ground truth, and the neural network model will finally predict the 3D key points based on the layout information of the type of car cabin.

[0032] For those types of vehicles already known by the neural network model, in step S330, the detector module may get the input of the cabin layout information and the input image frame captured from the DMS camera. The encoder of the detector module may extract features of the input image frame. Different types of the vehicle cabins may have a variety of arrangements with various cabin layout information, that is why the cabin layout information needs to be as the input parameter to the system.

[0033] Based on this information, in step S340, the decoder of the detector module may predict the 3D key points of the human face and the phone to generate at least one feature map. To inference 3D key points in a new known type of car cabin, the cabin layout parameter may be changed, according to, but the one only input image frame captured by the camera would not need to be modified.

[0034] Next, in step S350, the decoder head of the detector module may classify each object and regress 3D key point coordinates of their key points.

[0035] After getting the 3D key points output from the detector module, in step S360, the 3D surface could get reconstructed. All the (x, y, z) values in each positive feature vector construct the 3D key points of the human head, the phone, etc. in the expected order. From those 3D key points of the human head, we can get the 3D key points of the mouth, the face and so on. The surface reconstruction process can be performed by, for example, Power Crust, Poisson, and / or other surface reconstruction algorithms.

[0036] Therefore, in step S360, the 3D surface of the human mouth and other objects such as a phone, a cigarette, may be reconstruct, respectively, and then the required distance for prediction can be calculated. As an example, a driver distraction alarm may be generated when the calculated 3D distance is shorter than a threshold, such as 5cm set for GB / T 41797-2022.

[0037] FIG. 4 illustrates an exemplary process diagram 400 of the method for deep learning-based 3D detecting driver distraction, in accordance with the one or more embodiments of the inventive subject matter.

[0038] On the left side of FIG. 4, an input image frame 410 captured by the DMS camera in the cabin has been input into the detector module 420 in the system for deep learning-based 3D detecting driver distraction as provided according to the present inventiveness subject matter. As shown schematically on the left side of Figure 4, the head of a driver inside the carriage and his phone are captured in this frame of image. In addition, detector module 420 may also obtained cabinet layout information, for example, from the factory setting information, as described earlier. It can be conceived that the images of the interior facilities in the cabin captured in the input image frame 410 can be corresponding to the obtain cabinet layout information of this cabin. As shown in FIG. 4, at least the coordinates of the driver's seat in the input image frame 410 can be corresponding to the cabin layout information, referring to the camera coordinate frame. After processing with the detector module 420 in the form of the neural network model, 16 key points along the facial contour 422 on the head, 14 key points at the mouth 424 on the head, 8 key points on the nose 426, 6 key points on each of his two eyes 428, and 7 key points at each of his two eyebrows 430 are detected from the input image frame 410. Therefore, the neural network model can regress these 64 key points on the driver’s head, and the 3D key point coordinates, (x0, y0, z0) … (x63, y63, z63) , of these 64 key points can be extracted into the head vector 260 as  shown in FIG. 2. The head vector 260 can reconstruct a 3D curved surface 440 of the driver’s head, including his face, mouth, nose, eyes, and eyebrows of the driver, as can be seen in FIG. 4.

[0039] In addition, key points such as mobile phones may also be detected from image frame 410. For example, a mobile phone can be viewed as a hexahedron, and the 8 vertices 450 of the hexahedron can be taken as the 3D key points of the phone. Therefore, the neural network model can classify these 8 key points on the phone, and the 3D key point coordinates (x0, y0, z0) … (x7, y7, z7) of these 8 key points can be regressed into the phone vector 270 as shown in Figure 2. The phone vector can reconstruct the 3D surface 460 of the phone, as shown in FIG. 4.

[0040] Therefore, the 3D distance, d, 470 between the driver’s head and the phone can be ultimately calculated for detecting driver distractions. As shown by the instance in FIG. 4, the 3D distance can be detected in real-time, or at certain time intervals. When the detected distance is shorter than the set minimum distance of 5cm according to GB / T 41797-2022, a driver distraction alarm will be generated to worn the driver to focus on driving.

[0041] The inventive subject matter is designed to solve the problem in ADAS space 2 manual distraction detection. The invention provides a new 3D key points detection head to estimate the 3D contours of driver head and phone. It can provide a quantified prediction of the distance between the head and phone, which is required by GB / T 41797-2022. There is no such solution in this industry may address the detection of the 3D distance. Meanwhile, the inventive subject matter can be based on a 3D object detection pipeline which is a lightweight model with regards to 3D object segmentation. And this method can also be extended to detect smoking behavior by the distance between the cigarette and the driver face. In general, this inventive subject matter may improve the precision level and interpretability of our manual distraction detection model. Also, this method can also be extended to 3D perception tasks in L2 / L3 / L4 automatic driving scene, such as obstacle detection, pedestrian detection and so on. Compared with segmentation-based methods, the inventive subject matter is based on 3D key points detection, which may be much lighter and faster. Compared with other 3D object detection methods which only detect 3D bounding boxes, the inventive subject matter can detect the detail 3D shape of object, which is much more precise and accurate. Therefore, the inventive subject matter may realize the above advantages in ADAS / Automatic driving business for our nowadays life.

[0042] Any combination of one or more computer-readable media may be used to perform the  method provided in one and more embodiments of the present inventive subject matter. The computer-readable medium may be a computer-readable signal medium or a computer-readable storage medium. The computer-readable storage medium may be, for example, but not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples (non-exhaustive list) of the computer-readable storage medium may include, for example: an electrical connection with one or more wires, portable computer floppy disks, hard disks, random access memory (RAM) , read-read-only memory (ROM) , erasable programmable read only memory (EPROM or flash memory) , optical fibers, portable compact disc read only memory (CD-ROM) , optical storage devices, magnetic storage devices, or any suitable combinations of the foregoing. In the context of the disclosure, the computer-readable storage medium may be any tangible medium that can include or store a program for use by or in connection with an instruction execution system, apparatus, or device.

[0043] As used in the disclosure, an element or step listed in the singular form and preceded by the word "one / a" should be understood as not excluding a plurality of said elements or steps, unless such exception is specifically stated. Furthermore, references to "embodiments" or "examples" of the disclosure are not intended to be construed as exclusive, also including the existence of other embodiments of the recited features. The terms "first" , "second" , "third" , etc. are used only for identification and are not intended to emphasize a numerical requirement or positioning order of their objects.

[0044] References in the present inventive subject matter to the method for reproducing a constant sound field include the following content:

[0045] Item 1: In one or more embodiments, the present inventive subject matter provides a system for deep learning-based 3D detecting driver distraction, comprising:

[0046] a detector module comprising:

[0047] an encoder configured to extract features of an input image frame and cabin layout information;

[0048] a decoder configured to generate at least one feature map; and

[0049] a detection head configured to classify each of at least two objects in the image frame and regress 3D key point coordinates of key points of said each of the at least two objects,

[0050] a surface reconstruction module configured to reconstruct at least two 3D surfaces for the  at least two objects, respectively, and

[0051] a calculation module configured to calculate a 3D distance between the at least two 3D surfaces.

[0052] Item 2. The system of item 1, wherein the detector module may be of a form of neural network model, which may be trained with plenty of paired cabin layout information and image frames of cabins of known various types of vehicles.

[0053] Item 3. The system of item 1 or 2, wherein the cabin layout information may comprise cabin dimensions, layout of seats, and camera positioning in the cabin.

[0054] Item 4. The system of any of items 1-3, wherein the input image frame can be captured from the camera arranged in the cabin.

[0055] Item 5. The system of any of items 1-4, wherein the at least one feature map comprises at least two feature vectors each with an objectiveness probability score, a classification probability score, and at least one regressed 3D key point coordinate of the key points of said each of the at least two objects.

[0056] Item 6. The system of any of items 1-5, wherein the at least two objects may comprise a human head and another item, and wherein said another item may comprise a phone or a cigarette.

[0057] Item 7. The system of any of items 1-6, wherein the number of the 3D key points from one of the at least two objects can be different than the number of the 3D key points from the other of the at least two objects.

[0058] Item 8. The system of any of items 1-7, wherein the number of the 3D key points from the human head can be more than the number of the 3D key points from said another item.

[0059] Item 9. The system of any of items 1-8, wherein the surface reconstruction module configured to reconstruct at least two 3D surfaces for the at least two objects, respectively, may use algorithms comprising Power Crust or Poisson.

[0060] Item 10. The system of any of items 1-9, wherein a driver distraction alarm may be generated when the calculated 3D distance is shorter than a threshold.

[0061] Item 11: In one or more embodiments, the present inventive subject matter further provides a method for deep learning-based 3D detecting driver distraction, comprising steps of:

[0062] extracting, via an encoder of a detector module, features of an input image frame and cabin layout information;

[0063] generating, via a decoder of the detector module, at least one feature map;

[0064] classifying, via a detection head of the detector module, each of at least two objects in the image frame and regressing 3D key point coordinates of key points of said each of the at least two objects;

[0065] reconstructing, via a surface reconstruction module, at least two 3D surfaces for the at least two objects, respectively, and

[0066] calculating, via a calculation module, a 3D distance between the at least two 3D surfaces.

[0067] Item 12. The method apparatus of item 11, wherein the detector module may be of a form of neural network model, which may be trained with plenty of paired cabin layout information and image frames of cabins of known various types of vehicles.

[0068] Item 13. The method apparatus of item 11 or 12, wherein the cabin layout information may comprise cabin dimensions, layout of seats, and camera positioning in the cabin.

[0069] Item 14. The method apparatus of any of items 11-13, wherein the input image frame can be captured from the camera arranged in the cabin.

[0070] Item 15. The method apparatus of any of items 11-14, wherein the at least one feature map comprises at least two feature vectors each with an objectiveness probability score, a classification probability score, and at least one regressed 3D key point coordinate of the key points of said each of the at least two objects.

[0071] Item 16. The method apparatus of any of items 11-15, wherein the at least two objects may comprise a human head and another item, and wherein said another item may comprise a phone or a cigarette.

[0072] Item 17. The method apparatus of any of items 11-16, wherein the number of the 3D key points from one of the at least two objects can be different than the number of the 3D key points from the other of the at least two objects.

[0073] Item 18. The method apparatus of any of items 11-17, wherein the number of the 3D key points from the human head can be more than the number of the 3D key points from said another item.

[0074] Item 19. The method apparatus of any of items 11-18, wherein reconstructing the at least two 3D surfaces for the at least two objects, respectively, may use algorithms comprising Power Crust or Poisson.

[0075] Item 20. The method apparatus of any of items 11-19, further comprises generating a driver distraction alarm when the calculated 3D distance is shorter than a threshold.

[0076] Item 21. In one or more embodiments, the present inventive subject matter further provides a computer-readable medium storing instruction that, when executed by one or more processors, may perform the steps of the method for deep learning-based 3D detecting driver distraction, according to any one of Items 11-20.

Claims

1.A system for deep learning-based 3D detecting driver distraction, comprising:a detector module comprising:an encoder configured to extract features of an input image frame and cabin layout information;a decoder configured to generate at least one feature map; anda detection head configured to classify each of at least two objects in the image frame and regress 3D key point coordinates of key points of said each of the at least two objects,a surface reconstruction module configured to reconstruct at least two 3D surfaces for the at least two objects, respectively, anda calculation module configured to calculate a 3D distance between the at least two 3D surfaces.2.The system of claim 1, wherein the detector module may be of a form of neural network model, which may be trained with plenty of paired cabin layout information and image frames of cabins of known various types of vehicles.3.The system of claim 1, wherein the cabin layout information may comprise cabin sizes, positioning of seats, and positioning of a camera in the cabin.4.The system of claim 2, wherein the input image frame can be captured from the camera arranged in the cabin.5.The system of claim 1, wherein the at least one feature map comprises at least two feature vectors each with an objectiveness probability score, a classification probability score, and at least one regressed 3D key point coordinate of the key points of said each of the at least two objects.6.The system of claim 5, wherein the at least two objects may comprise a human head and another item, and wherein said another item may comprise a phone or a cigarette.7.The system of claim 5, wherein the number of the 3D key points from one of the at least two objects can be different than the number of the 3D key points from the other of the at least two objects.8.The system of claim 7, wherein the number of the 3D key points from the human head can be more than the number of the 3D key points from said another item.9.The system of claim 1, wherein the surface reconstruction module configured to reconstruct at least two 3D surfaces for the at least two objects, respectively, may use algorithms comprising Power Crust or Poisson.10.The system of claim 1, wherein a driver distraction alarm may be generated when the calculated 3D distance is shorter than a threshold.11.A method for deep learning-based 3D detecting driver distraction, comprising steps of:extracting, via an encoder of a detector module, features of an input image frame and cabin layout information;generating, via a decoder of the detector module, at least one feature map;classifying, via a detection head of the detector module, each of at least two objects in the image frame and regressing 3D key point coordinates of key points of said each of the at least two objects;reconstructing, via a surface reconstruction module, at least two 3D surfaces for the at least two objects, respectively, andcalculating, via a calculation module, a 3D distance between the at least two 3D surfaces.12.The method of claim 11, wherein the detector module may be of a form of neural network model, which may be trained with plenty of paired cabin layout information and image frames of cabins of known various types of vehicles.13.The method of claim 11, wherein the cabin layout information may comprise cabin sizes,  positioning of seats, and positioning of a camera in the cabin.14.The method of claim 12, wherein the input image frame can be captured from the camera arranged in the cabin.15.The method of claim 11, wherein the at least one feature map comprises at least two feature vectors each with an objectiveness probability score, a classification probability score, and at least one regressed 3D key point coordinate of the key points of said each of the at least two objects.16.The method of claim 15, wherein the at least two objects may comprise a human head and another item, and wherein said another item may comprise a phone or a cigarette.17.The method of claim 15, wherein the number of the 3D key points from one of the at least two objects can be different than the number of the 3D key points from the other of the at least two objects.18.The method of claim 17, wherein the number of the 3D key points from the human head can be more than the number of the 3D key points from said another item.19.The method of claim 11, wherein reconstructing the at least two 3D surfaces for the at least two objects, respectively, may use algorithms comprising Power Crust or Poisson.20.The method of claim 11, further comprises generating a driver distraction alarm when the calculated 3D distance is shorter than a threshold.21.A computer-readable medium storing instruction that, when executed by one or more processors, may perform the steps of the method according to any one of claims 11-20.

Citation Information

Patent Citations

  • Detection of driver behaviors using in-vehicle systems and methods

    US20160046298A1

  • Detection of driver behaviors using in-vehicle systems and methods

    US9714037B2