System and method for detecting movements and actions of occupants in a vehicle

The system addresses the inefficiencies of existing 2D image-based occupant detection systems by using learning engines to estimate 3D coordinates and movements, effectively monitoring occupant awareness and alerting for distracted driving.

DE102024133404A1Pending Publication Date: 2025-06-05MERCEDES BENZ GROUP AG

Patent Information

Application Number
DE102024133404
Authority / Receiving Office
DE · DE
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-05
Filing Date
2024-11-14
Publication Date
2025-06-05

AI Technical Summary

Technical Problem

Existing systems for detecting occupant movements and actions in vehicles from 2D images are less accurate and inefficient due to the lack of depth information in 2D images, which limits the effective estimation of 3D coordinates and movements.

Method used

A system and method that utilize a processing device equipped with a first learning engine for recognizing 2D joints and a second learning engine for estimating coarse depth, enabling the estimation of 3D coordinates and spatio-temporal pose of occupants from 2D images.

Benefits of technology

The system effectively detects movements and actions of occupants, monitors their awareness, and generates alerts for distracted drivers, providing an improved, efficient, and reliable solution for enhancing vehicle safety.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

The present disclosure relates to a system (100) for detecting the movement and action of occupants in a vehicle. The system (100) comprises a processing device (104) configured to detect images of occupants within the vehicle using an image capture unit (102). The processing device (104) further detects 2D joints associated with each occupant and accordingly determines coordinates of the detected 2D joints in the images to define a skeletal structure of the occupants, and further estimates the coarse depth of each pixel in the images. Furthermore, the processing device (104) estimates coarse 3D coordinates of the detected joints based on the determined coordinates and the estimated coarse depth to determine a spatio-temporal pose of the occupants and accordingly detect the movements and actions performed by the corresponding occupants.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure relates to the field of occupant perception detection systems. More specifically, the present disclosure provides a system and method for detecting occupant movements and actions in a vehicle. BACKGROUND

[0002] Driving a vehicle requires the driver's constant attention and alertness to ensure safe operation. Maintaining awareness of the surroundings and making informed decisions is critical to avoiding accidents and reducing the risk associated with driver distraction or fatigue. Therefore, driver attention monitoring has received considerable attention in recent years as a means of improving driver safety and gaining valuable insights into driver condition.

[0003] Known systems use various sensors and technologies, such as cameras, infrared sensors, and biometric measurements, to monitor the driver's movements and actions. This allows signs of driver distraction or fatigue to be detected and warnings or interventions to mitigate potential risks can be triggered. These solutions often assess the driver's level of attention and awareness while driving and initiate appropriate actions, such as issuing warnings or activating autonomous safety features.

[0004] Patent document number CN112070027B discloses a method for training a network and recognizing actions. The method includes the steps of updating model parameters of a pre-training model using a first sequence data set of a human body skeleton point sequence and a viewpoint label corresponding to each first sequence data value in the first sequence data set, followed by a further step of initializing model parameters of a human body action recognition model based on the updated model parameters of the pre-training model. The pre-training model and the human body action recognition model comprise feature extraction networks with the same structure.Furthermore, the method comprises the steps of updating the model parameters of the human motion recognition model using the second sequence data of the human skeleton point sequence and the motion category label corresponding to each second sequence data in the second sequence data set to obtain the trained human motion recognition model.

[0005] Another patent document, numbered CN112989947A, discloses a method and apparatus for estimating the three-dimensional coordinates of human keypoints. The estimated three-dimensional coordinates are used to determine and obtain thermodynamic diagrams containing the human keypoints and two-dimensional coordinates of each human keypoint from an image to be identified. The thermodynamic diagram of the keypoints of each human body and the image to be recognized are input into a trained depth recognition model to determine the depth information of the human body keypoint based on the thermodynamic diagram of the keypoints and the keypoint characteristics of the image to be recognized.Furthermore, the three-dimensional coordinates of each human body key point are determined according to the depth information of each human body key point and the corresponding two-dimensional coordinates.

[0006] Another patent document numbered CN114005105B discloses a method and apparatus for detecting driving behavior, as well as electronic equipment. The method includes the steps of acquiring a main driving window image, wherein the main driving window image includes a driver. The method further includes the steps of extracting the driver's human keypoints in the main driving window image using a specified CPN convolutional neural network to obtain human keypoint coordinates, and determining driving behavior feature data based on the human body keypoint coordinates, wherein the driving behavior feature data is used to characterize the driver's driving behavior.Furthermore, the method includes the steps of detecting whether the driver has abnormal driving behavior by a trained FD-ANN discrimination model according to the characteristic driving behavior data to obtain an abnormal driving behavior detection result.

[0007] Another patent document, numbered CN110543848B, discloses a method and apparatus for recognizing driver actions based on a three-dimensional convolutional neural network. The target model is obtained by training a constructed three-dimensional convolutional neural network. The three-dimensional convolutional neural network includes a plurality of sequentially connected combined layer structures, and each combined layer structure includes a convolutional layer and a pooling layer. The method involves analyzing videos recorded of train drivers during their journeys, extracting feature data, including optical flow features derived from pixel point changes, and feeding them into the target model, which calculates the probability that the driver's actions correspond to predefined actions.The feature extraction process includes the extraction of video images, primary colors, grayscale, gradients, and optical flow features from each image, which enables the identification of actions, with the optical flow features reflecting the pixel point changes between images and ultimately increasing the accuracy of identifying driver actions during train operation.

[0008] While the cited references focus on extracting features from captured images or videos to detect the coordinates of human keypoints and estimate driver actions, they do not consider the depth of keypoints in 2D images. As a result, the cited references may be less accurate, ineffective, and inefficient in detecting the movements and actions of vehicle occupants. Furthermore, the predictive models or CNN models of the cited references are not trained using the depth data of pixels associated with human joints in known images, limiting the existing solution to effectively estimating the 3D coordinates of the joints, the 3D position of the human skeleton, and human movements / actions from 2D images.

[0009] Predicting 3D poses from 2D images requires resource-intensive 3D CNN models, which are complicated by the need for extensive computational power and storage. The lack of annotated 3D data collected through costly lab setups with various sensors further complicates model training. Furthermore, this task is inherently difficult because it involves addressing an unsolved problem, making it difficult to accurately derive 3D poses from 2D images. Furthermore, accurately predicting the correct action of an occupant when they perform multiple actions simultaneously is a complex task. Finally, detecting overlapping activities of a seated occupant presents a unique challenge, further complicating the understanding and interpretation of human actions and poses.

[0010] Therefore, there is a need to overcome the drawbacks, shortcomings, and limitations of existing solutions by providing an improved, efficient, and reliable solution for detecting occupant movements and actions in a vehicle using 2D images. Furthermore, there is a need to monitor occupant alertness levels and alert occupants accordingly, taking proactive safety measures. SUBJECT OF THE PRESENT DISCLOSURE

[0011] A general object of the present disclosure is to detect the movements and actions of passengers and drivers in a vehicle and accordingly to monitor the driver's level of attention while driving.

[0012] One objective of the present disclosure is to detect the occupant's head, torso, neck and arm movements in order to assess his or her consciousness during the journey using 2D images.

[0013] One objective of the present disclosure is to estimate the 3D skeletal pose and the movements / actions of the occupants in the vehicle from 2D images.

[0014] Another object of the present disclosure is to provide an improved, efficient and reliable system and method for detecting the movement and action of occupants in a vehicle.

[0015] A further objective of this disclosure is to provide an improved, efficient and reliable system and method that monitors the state of consciousness of occupants and warns the occupants accordingly and takes proactive safety measures. SUMMARY

[0016] Aspects of the present disclosure relate to the field of occupant awareness detection systems. More specifically, the present disclosure provides a system and method for detecting the movement and action of occupants in a vehicle.

[0017] One aspect of the present disclosure relates to a system for detecting movements and actions of occupants in a vehicle. The system includes an image capture unit configured in the vehicle to capture one or more images of one or more occupants in the vehicle, and a processing device communicating with the image capture unit. The processing device includes one or more processors coupled to a memory that stores instructions executable by the processors. The processing device is configured to receive the captured images of the occupants from the image capture unit in real time, detect one or more 2D joints associated with each of the occupants in the received images using a first learning machine, and accordingly determine the coordinates of the detected 2D joints in the received images.to define a skeletal structure of the respective occupants, estimate the coarse depth of each pixel in the received images using a second learning machine, estimate coarse 3D coordinates of the detected joints based on the determined coordinates and the estimated coarse depth of the pixels associated with the detected joints, to determine a spatiotemporal pose of the one or more occupants, and monitor the spatiotemporal pose of the one or more occupants to detect the movements and actions performed by the respective occupants.

[0018] In one aspect, the processing device may be configured to monitor the attention of the occupants within the vehicle based on the detected movement and actions performed by the respective occupants and generate an alert when at least one driver among the one or more occupants is detected to be distracted or inattentive while driving the vehicle.

[0019] In one aspect, the one or more joints may include the elbow, wrist, shoulder, head, torso, and neck.

[0020] In one aspect, the first learning module may comprise a modified convolutional neural network (CCN) that enables the processing device to generate heat maps in the received images to predict probabilities for the presence of each of the joints in the pixels and, accordingly, to detect the one or more 2D joints.

[0021] In one aspect, the second learning engine may include a CNN-based encoding / decoding module and a transformer-based adaptive bin module that enable the processing device to estimate the coarse depth of each pixel in the received images.

[0022] In one aspect, the first learning machine may be trained by a mean square error loss function using a plurality of known images associated with occupants performing known actions and movements, and corresponding joint data and coordinates. Furthermore, the second learning machine may be trained by a different loss function using the plurality of known images associated with occupants and corresponding depth data acquired with an external depth sensor.

[0023] Another aspect of the present disclosure relates to a method for detecting movements and actions of occupants in a vehicle. The method comprises: capturing one or more images of one or more occupants in the vehicle by an image capture unit; detecting one or more 2D joints associated with each of the occupants in the captured images by a processing device using a first learning engine and correspondingly determining coordinates of the detected 2D joints in the received images to define a skeletal structure of the corresponding occupants; estimating a coarse depth of each pixel in the captured images by the processing device using a second learning engine;Estimating, by the processing device, coarse 3D coordinates of the detected joints based on the determined coordinates and the estimated coarse depth of the pixels associated with the detected joints to determine a spatiotemporal posture of the one or more occupants; and monitoring, by the processing device, the spatiotemporal posture of the one or more occupants to detect the movement and action performed by the corresponding occupants.

[0024] In one aspect, the method may comprise the steps of: monitoring, by the processing device, the awareness of the occupants within the vehicle based on the detected movement and the action performed by the respective occupants, and generating an alarm, by the processing device, when at least one driver among the one or more occupants is detected as being distracted or inattentive while driving the vehicle.

[0025] In one aspect, the method may comprise the steps of enabling the processing device, by the first learning engine, to generate heat maps in the images to predict probabilities for the presence of each of the joints in each of the pixels and, accordingly, to detect the one or more 2D joints.

[0026] In one aspect, the first learning engine may be trained using a mean squared error loss function, using a plurality of known images associated with occupants performing known actions and movements, as well as corresponding joint data and coordinates of the respective joints. Furthermore, the second learning engine may be trained using the plurality of known images associated with occupants and corresponding depth data acquired with an external depth sensor.

[0027] Various objects, features, aspects and advantages of the subject invention will become more apparent from the following detailed description of preferred embodiments together with the accompanying drawings in which like numerals represent like components. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] The accompanying drawings are included to provide a further understanding of the present disclosure and are incorporated in and constitute a part of this specification. The drawings illustrate exemplary embodiments of the present disclosure and, together with the description, serve to explain the principles of the present disclosure. Fig. 1 shows an exemplary block diagram of the proposed system for detecting the movement and action of occupants in a vehicle according to an embodiment of the present invention. Fig. 2 shows an exemplary architecture illustrating functional units of the processing device associated with the proposed system according to an embodiment of the present invention. Fig. 3 shows exemplary steps of the proposed method for detecting the movement and action of occupants in a vehicle according to an embodiment of the present disclosure. Fig. 4 shows an exemplary flowchart illustrating the overall operation of the proposed system in accordance with an embodiment of the present disclosure. Fig. 5 shows an exemplary skeleton structure created by the processing device according to an embodiment of the present disclosure. Fig. 6 shows an exemplary representation of the process for identifying and classifying actions using the 3D keypoints and the spatio-temporal pose according to an embodiment of the present disclosure. Fig. 7 shows an exemplary architecture of the second learning module including a CNN-based encoder-decoder module and a transformer-based adaptive bin module, in accordance with an embodiment of the present disclosure. DETAILED DESCRIPTION

[0029] The following is a detailed description of the embodiments of the disclosure illustrated in the accompanying drawings. The embodiments are detailed enough to clearly convey the disclosure. However, the necessary detail is not intended to limit the foreseeable variations of embodiments; on the contrary, it is intended to cover all modifications, equivalents, and alternatives that fall within the spirit and scope of the present disclosures as defined by the appended claims.

[0030] The embodiments discussed herein relate to the field of occupant awareness detection systems. More specifically, the present disclosure provides a system and method for detecting the movement and action of occupants in a vehicle.

[0031] With reference to Fig. 1, where an exemplary block diagram of the proposed system for detecting the movement and action of occupants in a vehicle (also referred to herein simply as the system) is disclosed, the system 100 may include an image capture unit 102 configured in the vehicle to capture one or more images or videos of one or more occupants in the vehicle. In an exemplary embodiment, the image capture unit 102 may include, but is not limited to, a still camera, a thermal camera, and / or an IR camera. The image capture unit 102 may be positioned in the interior of the vehicle in a predetermined direction and orientation such that the image capture unit 102 faces the occupants, in particular the driver of the vehicle. The system 100 may further include a processing device 104 that communicates with the image capture unit 102.In one embodiment, the processing device 104 may be, but is not limited to, an electronic control unit (ECU 106) of the vehicle. In another embodiment, the processing device 104 may also be a central server that communicates with the image acquisition unit 102.

[0032] Additionally, system 100 may include one or more warning devices 108 selected from the vehicle's speakers 108-1, a display and / or lights 108-2 present in the vehicle's interior, and haptic feedback devices 108-3 configured with the vehicle's steering wheel and / or seats. Warning devices 108 may be in communication with controller 106 or processing device 104 and configured to generate audible, visual, and haptic warnings.

[0033] In one embodiment, the processing device 104 may be configured to use a first learning machine (214 in Fig. 2) detects one or more 2D joints associated with each of the occupants in the images captured by the image acquisition unit 102. The joints may include, but are not limited to, the elbow, wrist, shoulder, head, torso, and neck. The processing device 104 may accordingly determine the coordinates of the detected 2D joints in the received images to define the skeletal structure of the corresponding occupants. Fig. 5 shows a skeletal structure 500 of the occupant with the recorded 2D joints. The detailed structure and functionality of the first learning machine were later developed in conjunction with Fig. 2 described.

[0034] The processing device 104 may be further configured to use a second learning machine (216 in Fig. 2) estimates the coarse depth of each pixel in the images acquired by the image acquisition unit 102. In one embodiment, the processing device 104 may determine the depth of each pixel in the acquired (IR) image and then select the depth value at the specific positions of the body joints. The depth values, together with the 2D keypoint positions from the pose estimation, may accordingly provide the 3D keypoints (body pose) for each occupant. The detailed structure and operation of the second learning engine were described later in connection with the Fig. 2 and Fig. 7 described.

[0035] Accordingly, the processing device 104 may analyze the determined coordinates and the estimated coarse depth of the pixels associated with the detected joints in the captured images to estimate coarse 3D coordinates of the detected joints in the images and further determine a spatiotemporal pose of the occupants. Finally, the processing device 104 may be configured to monitor the spatiotemporal pose of the occupants to detect the movements and actions performed by the corresponding occupants. This may enable the processing device 104 to determine and monitor the level of attention of the occupants and the driver while driving and to activate the warning devices 108 to generate a warning if at least the driver is distracted or inattentive while driving.

[0036] In one embodiment, processing device 104 may be configured to monitor the awareness of the vehicle's occupants in real time based on the detected movement and actions performed by the respective occupants. Accordingly, processing device 104 may be configured to actuate warning devices 108 to generate a warning when at least the driver of the vehicle is distracted or inattentive while driving.

[0037] In one embodiment, the processing device 104 may communicate with the image capture unit 102, the ECU 106, the speakers 108-1, the light / display 108-2, and the haptic feedback devices 108-3 via a network. Further, the network may be a wireless network, a wired network, or a combination thereof, which may be implemented as one of various types of networks, such as an intranet, local area network (LAN), wide area network (WAN), the Internet, and the like. Furthermore, the network may be either a dedicated network or a shared network. The shared network may represent an interconnection of various types of networks that may use a variety of protocols, such as Hypertext Transfer Protocol (HTTP), Transmission Control Protocol / Internet Protocol (TCP / IP), Wireless Application Protocol (WAP), and the like.

[0038] In one embodiment, system 100 may be implemented using any one or a combination of hardware components and software components, such as a cloud, a server, a computer system 100, a computing device, a network device, and the like. Furthermore, if the vehicle's ECU 106 has been successfully paired with the processing device 104 using methods described in detail in subsequent sections, the processing device 104 may interact with the image capture unit 102, the ECU 106, the speakers 108-1, the light / display 108-2, and the haptic feedback devices 108-3 via a website without requiring any mobile applications.

[0039] Referring to Fig. 2 shows the block diagram of Fig. 2 illustrates exemplary functional units of processing device 104, which may include one or more processors 202, memory 204, interface(s) 206, processing engine(s) 208, and database 210. The one or more processors 202 may be implemented as one or more microprocessors, microcomputers, microcontrollers, digital signal processors, central processing units, logic circuits, and / or any devices that manipulate data based on operating instructions. Among other capabilities, the one or more processors 202 are configured to retrieve and execute computer-readable instructions stored in a memory 204 of processing device 104. Memory 204 may store one or more computer-readable instructions or routines that may be retrieved and executed to create or release the data units via a network service.The memory 204 may comprise any non-volatile memory device, e.g., volatile memory such as RAM or non-volatile memory such as EPROM, flash memory, and the like.

[0040] In one embodiment, processing device 104 may also include one or more interfaces 206. Interface(s) 206 may include a variety of interfaces, such as interfaces for data input and output devices, referred to as I / O devices, storage devices, and the like. Interface(s) 206 may facilitate communication of processing device 104 with various devices connected to processing device 104. Interface(s) 206 may also provide a communication path for one or more components of processing device 104. Examples of such components include, but are not limited to, processing engine(s) 208 and database 210.

[0041] In one embodiment, the processing engine(s) 208 may be implemented as a combination of hardware and programming (e.g., programmable instructions) to implement one or more functions of the processing engine(s) 208. In the examples described herein, such combinations of hardware and programming may be implemented in a variety of ways. For example, the programming for the processing engine(s) 208 may consist of processor-executable instructions stored on a non-transitory, machine-readable storage medium, and the hardware for the processing engine(s) 208 may include a processing resource (e.g., one or more processors) to execute such instructions. In the present examples, the machine-readable storage medium may store instructions that, when executed by the processing resource, implement the processing engine(s) 208.In such examples, the processing device 104 may include the machine-readable storage medium storing the instructions and the processing resource for executing the instructions, or the machine-readable storage medium may be separate but accessible to the processing device 104 and the processing resource. In other examples, the processing engine(s) 208 may be implemented by electronic circuitry. The database 210 may include data either stored or generated as a result of functionality implemented by one of the components of the processing engine(s) 208.

[0042] In one embodiment, the processing engine(s) 208 may include a data acquisition unit 212, a first learning engine 214, a second learning engine 216, an action and motion estimation unit 218, an alerting unit 220, and other unit(s) 222. The other unit(s) 222 may implement functionality that complements applications or functions performed by the processing device 104 or the processing engine(s) 208.

[0043] According to one embodiment, the data acquisition unit 212 may cause the processing device 104 or the vehicle's ECU 106 to enable the image acquisition unit 102 to capture images or videos of the occupants within the vehicle. Furthermore, in one embodiment, during training of the learning machines 214, 216, the data acquisition unit 212 may cause the processing device 104 to receive a plurality of known images associated with humans (occupants) performing known actions and movements, along with the joint data and corresponding coordinates associated with the humans in the known images. The data acquisition unit 212 may also cause the processing device 104 to capture and receive, using an external depth sensor, the depth data associated with each pixel in the known images while training the second learning machines 214.The detailed structure and operation of the first learning machine and the second learning machine were described in the following sections.

[0044] According to one embodiment, the first learning engine 214 may include a modified convolutional neural network (CNN), which may cause the processing device 104 to detect the 2D joints associated with each of the occupants in the captured images and, accordingly, determine coordinates of the detected 2D joints (keypoints) in the images to define a skeletal structure of the corresponding occupants. In one embodiment, the modified CNN / first learning engine 214 may cause the processing device 104 to generate heat maps in the images to predict probabilities for the presence of each of the joints in the pixels and detect the 2D joints accordingly.

[0045] Furthermore, in one embodiment, the first learning machine 214 may be trained with a mean square error (MSE) loss function using a plurality (in lakhs) of known images associated with humans (occupants) performing known actions and movements, along with the joint data and corresponding coordinates associated with the humans in the known images. The mean square error loss function, MSE=1N∑i=1N(yi−yi)2, where yi is the predicted probability for a pixel.

[0046] The modified CNN can have an hourglass-shaped design to capture the data at any scale. The modified CNN can use skip connections to obtain spatial information at any resolution and forward it for higher sampling, which takes place further down the hourglass. The modified CNN can use a metric PCKh=∑j=1N∑i=1k(‖vi−vi'‖<δ)∑i=1k1 where N is the total number of images, v is the joint vector, δ is the threshold and k is the number of keypoints (joints).

[0047] According to one embodiment, the second learning engine 216 may cause the processing device 104 to estimate the coarse depth (monocle depth) of each pixel in the received images. The second learning engine 216 may cause the processing device 104 to determine the depth of each pixel in the acquired IR image and then select the depth value at the determined positions of the body joints. The depth values, together with the positions of the 2D keypoints from the pose estimation, may enable the processing device 104 to estimate the 3D keypoints (body pose) for each occupant.

[0048] In an embodiment relating to Fig. 7, the second learning engine 216 may include a CNN-based encoder-decoder module 702 along with a transformer-based adaptive bin module 704, which enables the processing device 104 to estimate the coarse depth of each pixel in the received images. The encoder / decoder module 702 may be used to extract depth-specific features from the images 706 captured by the image acquisition unit 102. Further, the transformer-based adaptive bin module 704 may use the depth-specific features (extracted by the encoder-decoder module) to calculate the accurate depth maps in the images.

[0049] In one embodiment, the second learning machine 216 may be trained by a loss function using the plurality (lakhs) of known images associated with humans (occupants) and corresponding depth data acquired with an external depth sensor. The loss function used may be L total = L pixel (Pixel-wise depth loss) + βL bins (Bin-center density loss), where Lpixel=α1T∑igi2−λT2(∑igi)2 and Lbins=chamfer(X,c(b))+chamfer(c(b),X), where, gi=log di∼−log di, d i is the value of the true depth, T is the number of valid GT, c(b) is the binding center, α = 10, and λ = 0.85

[0050] The second learning processor 216 may use adaptive binning to first determine the depth interval D = (d min , d max ) into N bins, where the bin widths b 718 are calculated adaptively for each image.

[0051] The deep repository c(b) 720 is calculated as follows: c(bi)=dmin+(dmax−dmin)(bi2+∑j=1i−1bj)

[0052] The second learning processor 216 may use the following metric: Average Relative Error (RLE): 1n∑pn|yp−γp|y where y p a pixel in the depth image y, y p ^ is a pixel in the predicted depth image □^ and n is the total number of pixels for each depth image.

[0053] According to one embodiment, the action and motion estimation unit 218 may cause the processing device 104 to estimate coarse 3D coordinates of the detected joints based on the determined coordinates and the estimated coarse depth of the pixels associated with the detected joints to determine a spatiotemporal pose of the one or more occupants. Furthermore, the action and motion estimation unit 218 may cause the processing device 104 to monitor the spatiotemporal pose of the occupants for a predefined frame or in real time to detect the movements and actions performed by the corresponding occupants.

[0054] Referring to Fig. 6, the action and motion estimation unit 218 may perform a classification of the actions 606 based on the 3D body pose points (body posture) stacked over a specific time window. A connected spatiotemporal graph may be constructed from the human pose points computed by the first learning engine 214 and the second learning engine 216. In one embodiment, a spatiotemporal graph convolutional network (ST-GCN) 602 may be trained to predict the occupant's intended action from the accumulated spatiotemporal graphs.

[0055] In one embodiment, a modified graph convolutional network modeled on the spatiotemporal architecture (ST-GCN) can be implemented. The graph nodes (body keypoints) can be divided into four subsets depending on which body part they belong to (hands, legs, head, torso). In an exemplary embodiment, a cross-entropy (CE) loss function can be used to train the modified graph convolutional network, where CE=−∑iCyilog(yi).

[0056] Furthermore, the action and motion estimation unit 218 and the warning unit 220 may cause the processing device 104 to monitor the awareness of the occupants in the vehicle based on the detected motion and the action performed by the respective occupants. Accordingly, the warning unit 220 may cause the processing device 104 to actuate the warning devices 108 (the vehicle's speakers, the display and / or lights in the vehicle's interior, and haptic feedback devices configured with the steering wheel and / or seats of the vehicle) to generate audible, visual, and / or haptic warnings.

[0057] With reference to the Fig. 3 and Fig.4 discloses the proposed method for detecting the movement and action of occupants in a vehicle. The method 300 may include the step 302 of capturing one or more images of one or more occupants in the vehicle by an image acquisition unit. The method 300 may further include the step 304 in which a processing device, using a first learning engine, detects one or more 2D joints 312 (keypoints) associated with each of the occupants in the captured images and, accordingly, determines coordinates of the detected 2D joints in the received images to define a skeletal structure of the corresponding occupants. Furthermore, the method 300 may include the step 306 in which the processing device, using a second learning engine, estimates the coarse depth 314 of each pixel in the captured images.

[0058] Further, the method 300 may include step 308, in which the processing device estimates coarse 3D coordinates of the detected joints based on the coordinates determined in step 304 and the coarse depth of the pixels associated with the detected joints, which are estimated in step 306, to determine a spatiotemporal pose of the one or more occupants. Furthermore, the method 300 may include step 310 of the processing device monitoring the spatiotemporal pose of the occupants to detect movements and actions performed by the corresponding occupants.

[0059] In one embodiment, method 300 may include the steps of monitoring, by the processing device, the awareness of the occupants within the vehicle based on the detected movement and actions performed by the respective occupants. If at least one driver among the occupants is detected as distracted or inattentive while driving the vehicle, method 300 may further include the steps of actuating vehicle speakers, the display, and / or lights provided within the vehicle's interior, and haptic feedback devices configured with the steering wheel and / or seats of the vehicle to generate an audible alarm, a visual alarm, and / or haptic alarms.

[0060] Thus, the present invention (system and method) provides an improved, efficient and reliable solution for detecting movements and actions of occupants in a vehicle using only 2D images, which enables the present invention to monitor the occupants' state of consciousness and to warn the occupants accordingly and to take proactive safety measures.

[0061] While the foregoing describes various embodiments of the invention, other and further embodiments of the invention may be devised without departing from the basic scope of the invention. The scope of the invention is determined by the following claims. The invention is not limited to the described embodiments, variations, or examples, which are included to enable a person of ordinary skill in the art to make and use the invention when combined with information and knowledge available to the person of ordinary skill in the art. BENEFITS OF THIS DISCLOSURE

[0062] The present disclosure detects the movements and actions of passengers and drivers in a vehicle and accordingly monitors the driver's state of consciousness while driving the vehicle.

[0063] This disclosure captures the occupant's head, torso, neck, and arm movements to assess their consciousness while driving using 2D images.

[0064] This disclosure estimates the 3D skeletal posture and the movements / actions of the occupants in the vehicle from 2D images.

[0065] The present disclosure provides an improved, efficient and reliable system and method for detecting the movement and action of occupants in a vehicle.

[0066] This disclosure provides an improved, efficient, and reliable system and method that monitors occupant consciousness and alerts occupants accordingly and takes proactive safety measures. QUOTES CONTAINED IN THE DESCRIPTION

[0000] This list of documents submitted by the applicant was generated automatically and is included solely for the convenience of the reader. This list is not part of the German patent or utility model application. The DPMA assumes no liability for any errors or omissions. Cited patent literature

[0000] CN 112070027B

[0004] CN 112989947A

[0005] CN 114005105

[0006] CN 110543848

[0007]

Claims

[1] A system (100) for detecting movements and actions of occupants in a vehicle, the system (100) comprising: an image capture unit (102) configured in the vehicle to capture one or more images of one or more occupants in the vehicle; and a processing device (104) in communication with the image capture unit (102), the processing device (104) comprising one or more processors (202) coupled to a memory (204) storing instructions executable by the processors (202), and the processing device (104) causing: to receive the captured images of the occupants from the image acquisition unit (102) in real time; using a first learning machine (214), to detect one or more 2D joints associated with each of the occupants in the received images and, accordingly, to determine coordinates of the detected 2D joints in the received images to define a skeletal structure of the corresponding occupants; using a second learning machine (216) to estimate the coarse depth of each pixel in the received images; estimate coarse 3D coordinates of the detected joints based on the determined coordinates and the estimated coarse depth of the pixels associated with the detected joints to determine a spatio-temporal position of the one or more occupants; and to monitor the spatio-temporal pose of one or more occupants in order to detect the movements and actions performed by the corresponding occupants. [2] The system (100) of claim 1, wherein the processing device (104) is configured to: to monitor the attention of the occupants within the vehicle based on the detected movement and the actions performed by the relevant occupants, and to generate a warning when at least one driver among the occupants is detected to be distracted or inattentive while driving the vehicle. [3] The system (100) of claim 1, wherein the one or more joints comprise elbow, wrist, shoulder, head, torso, and neck. [4] The system (100) of claim 1, wherein the first learning module (214) comprises a modified convolutional neural network (CCN) that enables the processing device (104) to generate heat maps in the received images to predict probabilities for the presence of each of the joints in the pixels and, accordingly, to detect the one or more 2D joints. [5] The system (100) of claim 4, wherein the second learning engine (216) comprises a CNN-based encoder-decoder module and a transformer-based adaptive bin module that enable the processing device (104) to estimate the coarse depth of each pixel in the received images. [6] The system (100) of claim 5, wherein the first learning machine (214) is trained by a mean square loss function using a plurality of known images associated with occupants performing known actions and movements and corresponding joint data and coordinates, and wherein the second learning machine (216) is trained by a different loss function using the plurality of known images associated with occupants and corresponding depth data acquired using an external depth sensor. [7] A method (300) for detecting movements and actions of occupants in a vehicle, the method (300) comprising: Capturing (302) one or more images of one or more occupants within the vehicle by an image capture unit (102); Detecting (304), by a processing device (104), using a first learning machine (214), one or more 2D joints associated with each of the occupants in the acquired images and determining coordinates of the detected 2D joints in the received images accordingly to define a skeletal structure of the corresponding occupants; Estimating (306) the coarse depth of each pixel in the captured images by the processing device (104) using a second learning machine (216); Estimating (308) coarse 3D coordinates of the detected joints by the processing device (104) based on the determined coordinates and the estimated coarse depth of the pixels associated with the detected joints to determine a spatio-temporal position of the one or more occupants; and Monitoring (310) the spatio-temporal pose of the one or more occupants by the processing device (104) to detect the movements and actions performed by the corresponding occupants. [8] The method (300) of claim 7, wherein the method (300) comprises the following steps: Monitoring the consciousness of the occupants within the vehicle by the processing device (104) based on the detected movement and the action performed by the respective occupants, and Generating a warning by the processing device (104) when at least one driver among the one or more occupants is detected as being distracted or inattentive while driving the vehicle. [9] The method (300) of claim 7, wherein the method (300) comprises the steps of enabling the processing device (104) through the first learning engine to generate heat maps in the images to predict probabilities of each of the joints being present in any of the pixels and, accordingly, detecting the one or more 2D joints. [10] The method (300) of claim 7, wherein the first learning machine (214) is trained by a mean square error loss function using a plurality of known images associated with occupants performing known actions and movements and corresponding joint data and coordinates of the respective joints, and wherein the second learning machine (216) is trained using the plurality of known images associated with occupants and corresponding depth data acquired using an external depth sensor.

Citation Information

Patent Citations

  • Driver action recognition method and device based on three-dimensional convolutional neural network

    CN110543848A

  • Network training, action recognition methods, devices, equipment and storage media

    CN112070027B

  • Method and device for estimating three-dimensional coordinates of human body key points

    CN112989947A

  • Driving behavior detection method and device and electronic equipment

    CN114005105A

Cited By

  • System for monitoring the posture of vehicle occupants

    DE202025103286U1