System and method for detecting human occupancy in vehicle

The system addresses occupancy detection challenges by using CNN and CRFs to process facial and body joints, ensuring accurate and efficient human occupancy detection in vehicles.

GB2636424APending Publication Date: 2025-06-18MERCEDES BENZ GROUP AG
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
GB2023019113
Authority / Receiving Office
GB · GB
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-12-14
Publication Date
2025-06-18

AI Technical Summary

Technical Problem

Existing occupancy detection systems in vehicles face challenges such as false positive/negative joint predictions due to occlusion, background variance, and variations in facial structures, leading to inefficient and unreliable human occupancy detection.

Method used

A system and method utilizing CNN and parallel CRFs to simultaneously process facial key-points and body joints, minimizing false positives and negatives by jointly determining and merging facial structure and body parts using an hourglass design Convolutional Neural Network (CNN) and Conditional Random Field (CRF) learning engine.

Benefits of technology

The system provides robust, intelligent, and reliable detection of human occupancy in vehicles, accurately determining the presence and position of occupants by integrating facial and body cues, thereby enhancing the reliability and efficiency of occupancy detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

Method of detecting human occupancy in a vehicle, comprising: acquiring an image 400; extracting localised temporal key points from the image 302, 404; inputting the key points into a Conditional Rand
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure relates to the field of image processing. In particular, the present disclosure provides a system and a method for detecting human occupancy in a vehicle. BACKGROUND

[0002] A vehicle is equipped with many sensors and devices for facilitating occupancy detection within the vehicle. Occupancy detection can also be carried out inside a vehicle using computer vision. Occupancy detection of driver / passengers is the fundamental problem for many in-car artificial intelligence (AI) use cases. Various such use cases can include interior automatically turning off lights of the vehicle if no one is sitting inside. Also, lights can be turned on automatically when person on other seat reach for a seat with no one occupied. Moreover, it also includes automated call for ambulance for number of occupants in case of emergencies. Further, deployment of airbags can also be optimized for maximum safety based on detection of the position and number of occupants.

[0003] Generally, occupancy detection is carried out only by focusing on one cue, for instance, human joints. However, it may lead to various challenges, for instance, it may cause failure of correct joint prediction due to occlusion. It may also lead to false positive / negative joint predictions because of background variance. Further, novel human poses during inference may not be correctly predicted.

[0004] Furthermore, facial structure detection also faces challenges due to variation in factors including age group, facial expressions, accessories, and cosmetics. Various techniques have been evolved to resolve aforementioned issues.

[0005] For instance, Patent document CN114333039A discloses a method, a device and a medium for portrait clustering. The method comprises the steps of extracting features of all snap shots collected by image collection equipment, and clustering the extracted face features and clustering human body features. When the confidence coefficient of the cluster features is larger than or equal to a first threshold value, generating a face clustering set and a human body clustering set; and when the preset requirements are met, clustering the snap shots to be clustered except the clustered snap shots to determine a new face clustering set and a new human body clustering set. Then, merging all the new clustering sets into the face clustering set or the human body clustering set, and finally performing fusion clustering on the merged face clustering set and the merged human body clustering set to serve as reference data for face recognition.

[0006] Patent document KR20210019645A discloses an apparatus and a method for monitoring a passenger in an autonomous vehicle. The apparatus and method are able to monitor the status of the passenger in a shuttle-type autonomous vehicle, and to, when a dangerous situation occurs or is expected, notify such fact to the outside. The apparatus comprises: a camera which photographs the inside of the autonomous vehicle in real time; a passenger and article processing unit which recognizes and tracks one or more of the passenger and articles from an image from the camera. Further, a posture and face direction estimation unit estimates the posture and face direction of the passenger based on the results of processing by the passenger and article processing unit; and a status determination unit determines whether the passenger is in danger or not based on output from the posture and face direction estimation unit. A notification unit which, when the passenger is determined to be in danger by the status determination unit, generates a corresponding notification message, and transmits the message to an external center.

[0007] While the cited references disclose various systems and methods for detecting occupancy in a vehicle, there is a possibility of providing more intelligent solution that enables efficient detection of human occupancy in a vehicle. OBJECTS OF THE PRESENT DISCLOSURE

[0008] A general object of the present disclosure is to provide a system and method for detecting human occupancy in a vehicle.

[0009] An object of the present disclosure is to provide a system and method that enables detection of human occupancy based on facial key-points or body joints, as well as a combination of both.

[0010] Another object of the present disclosure is to provide a system and method for detecting human occupancy using CNN and parallel CRFs for processing facial key-points and body joints, simultaneously.

[0011] Another object of the present disclosure is to provide a robust system and method that mitigates false positive results and false negative results while detecting the human occupancy.

[0012] Another object of the present disclosure is to provide a smart, intelligent, efficient, and reliable system and method for detecting human occupancy in the vehicle. SUMMARY

[0013] Aspects of the present disclosure relate to the field of image processing. In particular, the present disclosure provides a system and a method for detecting human occupancy in a vehicle.

[0014] An aspect of the present disclosure pertains to a method for detecting human occupancy in a vehicle. The method includes the steps of: acquiring, through an image acquisition unit, one or more multimedia frames of a Region of Interest (Rol) pertaining to a seating section inside the vehicle; receiving, at an occupancy detection device in communication with a Conditional Random Field (CRF) based learning engine, the one or more acquired multimedia frames from the image acquisition unit; and then extracting, at the occupancy detection device, one or more temporal attributes from the one or more received multimedia frames, and correspondingly localizing the extracted temporal attributes. Further, the method includes the steps of: determining, at the occupancy detection device, facial structure and body parts of the human-entity by mapping the localized temporal attributes; and further merging, at the occupancy detection device, the determined facial structure and body parts for facilitating detection of the human-entity in said section.

[0015] In an aspect, the method may include the step of jointly determining the facial structure and body parts of the human-entity by feeding the localized temporal attributes in the learning engine including at least two parallel branches having hourglass design Convolutional Neural Network (CNN), such that first temporal attributes may be fed in first branch of the learning engine, and other temporal attributes may be fed in respective branches of the learning engine.

[0016] In another aspect, the first temporal attributes may pertain to facial key-points, and the other temporal attributes may pertain to body joints.

[0017] In another aspect, the method may include the step of minimizing false positive and false negative detections by jointly determining the facial structure and body parts of the human-entity, and correspondingly merging them.

[0018] Another aspect of the present disclosure pertains to a human occupancy detection system implemented in a vehicle. The system includes an image acquisition unit and an occupancy detection device in communication with a Conditional Random Field (CRF) based learning engine. The image acquisition unit is configured to acquire one or more multimedia frames of a Region of Interest (Rol) pertaining to a seating section inside the vehicle. The occupancy detection device is coupled with the image acquisition unit, and the occupancy detection device includes one or more processors coupled to a memory storing instructions executable by the one or more processors. The occupancy detection device is configured to receive, from the image acquisition unit, the one or more acquired multimedia frames; and extract one or more temporal attributes from the one or more received multimedia frames, and correspondingly localize the extracted temporal attributes. Further, the occupancy detection device is configured to determine facial structure and body parts of the human-entity by mapping the localized temporal attributes; and then merge the determined facial structure and body parts for facilitating detection of the human-entity in said section.

[0019] In an aspect, the learning engine may include at least two parallel branches having hourglass design Convolutional Neural Network (CNN).

[0020] In another aspect, the learning engine may be configured to jointly determine the facial structure and body parts of the human-entity by analyzing the localized temporal attributes.

[0021] In one aspect, the learning engine may include one or more linear CRFs configured to form a human pose by aligning the determined facial structure and body parts of the humanentity.

[0022] In other aspect, the system may also be configured to minimize false positive and false negative detections by jointly determining the facial structure and body parts of the human-entity, and correspondingly merging them.

[0023] In another aspect, the temporal attributes may pertain to facial key-points and body joints.

[0024] Various objects, features, aspects and advantages of the inventive subject matter will become more apparent from the following detailed description of preferred embodiments, along with the accompanying drawing figures in which like numerals represent like components. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] The accompanying drawings are included to provide a further understanding of the present disclosure, and are incorporated in and constitute a part of this specification. The drawings illustrate exemplary embodiments of the present disclosure and, together with the description, serve to explain the principles of the present disclosure.

[0026] FIG. 1 illustrates an exemplary network architecture of the proposed system for detecting human occupancy in a vehicle, in order to elaborate its overall working, in accordance with an embodiment of the present disclosure.

[0027] FIG. 2 illustrates a block diagram representing functional units of estimating unit of the proposed system, in accordance with an embodiment of the present disclosure.

[0028] FIG. 3 illustrates a diagram depicting obtaining various facial key-points, in accordance with an exemplary embodiment of the present disclosure.

[0029] FIG. 4 illustrates a flow diagram depicting extraction of body joints along with the facial key-points and then merging by the proposed system in order to form a human pose, in accordance with an exemplary embodiment of the present disclosure.

[0030] FIG. 5A illustrates an exemplary representation of CNN for body joints and facial key-points localization, in accordance with an exemplary embodiment of the present disclosure.

[0031] FIG. 5B illustrates an exemplary representation of CRF for body structure likelihood representation, in accordance with an exemplary embodiment of the present disclosure.

[0032] FIG. 6 illustrates a chart representing step-wise functioning of the proposed system for detecting human occupancy in a vehicle, in accordance with an embodiment of the present disclosure.

[0033] FIG. 7A illustrates a first exemplary scenario explaining functioning of the proposed system, in accordance with an embodiment of the present disclosure.

[0034] FIG. 7B illustrates a second exemplary scenario explaining functioning of the proposed system, in accordance with an embodiment of the present disclosure.

[0035] FIG. 8 illustrates a flow diagram depicting steps involved in the proposed method for detecting human occupancy in a vehicle, in order to elaborate its overall working, in accordance with an embodiment of the present disclosure.

[0036] FIG. 9 illustrates an exemplary computer system in which or with which embodiments of the present invention can be utilized in accordance with embodiments of the present disclosure. DETAILED DESCRIPTION

[0037] The following is a detailed description of embodiments of the disclosure depicted in the accompanying drawings. The embodiments are in such details as to clearly communicate the disclosure. However, the amount of detail offered is not intended to limit the anticipated variations of embodiments; on the contrary, the intention is to cover all modifications, equivalents, and alternatives falling within the spirit and scope of the present disclosures as defined by the appended claims.

[0038] Embodiments explained herein relate to the field of image processing. In particular, the present disclosure provides a system and a method for detecting human occupancy in a vehicle.

[0039] Referring to FIG. 1, the proposed system 100 for detecting human occupancy in a vehicle (interchangeably, referred to as system 100, hereinafter) can be implemented within a vehicle, whereby the system 100 can facilitate detection of human occupancy within a seating section of said vehicle.

[0040] In an embodiment, the system 100 can include an image acquisition unit 102 that can be configured to acquire one or more multimedia frames of a Region of Interest (Rol) pertaining to the seating section inside the vehicle. In an exemplary embodiment, the image acquisition unit 102 can include a camera such that Field of View (FOV) of the camera may cover whole of the seating section of the vehicle. In one embodiment, the multimedia frames may pertain to continuous video associated with the seating section inside the vehicle. In other embodiment, the multimedia frames may pertain to one or more images associated with the seating section inside the vehicle. The camera may get actuated time-to-time upon sensing motion inside the vehicle, and may correspondingly capture instant images of the seating section.

[0041] In another embodiment, the system 100 can include an occupancy detection device 106 that can include one or more processors coupled to a memory storing instructions executable by the one or more processors. In one embodiment, the occupancy detection device 106 may be an integral part of Electronic Control Unit (ECU) of the vehicle. In other embodiment, the occupancy detection device 106 may act as an independent unit located at an strategic position within the vehicle.

[0042] The occupancy detection device 106 can be in communication with a Conditional Random Field (CRF) based learning engine 108. In an implementation, the learning engine 108 can include at least two parallel branches having hourglass design Convolutional Neural Network (CNN). Further, the occupancy detection device 106 can also be in communication with the image acquisition unit 102.

[0043] In an embodiment, the occupancy detection device 106 can be configured to receive the one or more acquired multimedia frames from the image acquisition unit 102. Further, the occupancy detection device 106 can be configured to extract one or more temporal attributes from the one or more received multimedia frames, and correspondingly localize the extracted temporal attributes. In an exemplary embodiment, the temporal attributes can pertain to facial key-points and body joints.

[0044] In another embodiment, the occupancy detection device 106 can be configured to determine facial structure and body parts of the human-entity by mapping the localized temporal attributes. For instance, the facial structure can be determined by localizing and mapping the facial key-points, and the body parts can be determined by localizing and mapping the body parts. In another embodiment, the learning engine 108 can include one or more linear CRFs configured to form a human pose by aligning the determined facial structure and body parts of the human-entity.

[0045] In one embodiment, the learning engine 108 can be configured to jointly determine the facial structure and body parts of the human-entity by feeding the localized temporal attributes. In other embodiment, the system 100 can also be configured to minimize false positive and false negative detections by jointly determining the facial structure and body parts of the human-entity, and correspondingly merging them for facilitating detection of the humanentity in said section.

[0046] According to an embodiment, the system 100 can include an interface device 104 that can be operatively coupled to the occupancy detection device 106, where the interface device 104 can include, but not limited to, display unit, GUI module, voice assistance module, and a user mobile computing device, such as smartphone, hi an embodiment, the system 100 can display the detected human occupancy through the interface device 104. In another embodiment, the system 100 can trigger alert signals, using the interface device 104, when no human is found in the vehicle or when number of humans / persons present in the vehicle are found to be more than seating capacity of the vehicle.

[0047] According to an embodiment, the occupancy detection device 106 can be in communication with the image acquisition unit 102, the interface device 104, and the learning engine 108, through a network 110. Further, the network 110 can be a wireless network, a wired network or a combination thereof that can be implemented as one of the different types of networks, such as Intranet, Local Area Network (LAN), Wide Area Network (WAN), Internet, and the like.

[0048] Further, the network 110 can either be a dedicated network or a shared network. The shared network can represent an association of different types of networks that can use variety of protocols, for example, Hypertext Transfer Protocol (HTTP), Transmission Control Protocol / Internet Protocol (TCP / IP), Wireless Application Protocol (WAP), and the like.

[0049] In an embodiment, the system 100 can be implemented using any or a combination of hardware components and software components such as a cloud, a server 112, a computing system, a computing device, a network device and the like. Further, the occupancy detection device 106 can interact with other components of the system 100, through a website or an application that can reside in the proposed system 100. In an implementation, the system 100 can be accessed by website or application that can be configured with any operating system, including but not limited to, Android™, iOS™, and the like. [00501 Referring to FIG. 2, showing a block diagram 200, the exemplary functional units of the occupancy detection device 106 can include one or more processor(s) 202. The one or more processor(s) 202 can be implemented as one or more microprocessors, microcomputers, microcontrollers, digital signal processors, central processing units, logic circuitries, and / or any devices that manipulate data based on operational instructions. Among other capabilities, the one or more processor(s) 202 are configured to fetch and execute computer-readable instructions stored in a memory 204 of the occupancy detection device 106. The memory 204 can store one or more computer-readable instructions or routines, which may be fetched and executed to create or share the data units over a network service. The memory 204 can include any non-transitory storage device including, for example, volatile memory such as RAM, or non-volatile memory such as EPROM, flash memory, and the like. [00511 In an embodiment, the occupancy detection device 106 can also include an interface(s) 206. The interface(s) 206 may include a variety of interfaces, for example, interfaces for data input and output devices, referred to as RO devices, storage devices, and the like. The interface(s) 206 may facilitate communication of the occupancy detection device 106 with various devices coupled to the occupancy detection device 106. The interface(s) 206 may also provide a communication pathway for one or more components of the occupancy detection device 106. Examples of such components include, but are not limited to, processing engine(s) 208 and database 210.

[0052] In an embodiment, the processing engine(s) 208 can be implemented as a combination of hardware and programming (for example, programmable instructions) to implement one or more functionalities of the processing engine(s) 208. In examples described herein, such combinations of hardware and programming may be implemented in several different ways. For example, the programming for the processing engine(s) 208 may be processor executable instructions stored on a non-transitory machine-readable storage medium and the hardware for the processing engine(s) 208 may include a processing resource (for example, one or more processors), to execute such instructions.

[0053] In the present examples, the machine-readable storage medium may store instructions that, when executed by the processing resource, implement the processing engine(s) 208. In such examples, the es occupancy detection device 106 can include the machine-readable storage medium storing the instructions and the processing resource to execute the instructions, or the machine-readable storage medium may be separate but accessible to the occupancy detection device 106 and the processing resource. In other examples, the processing engine(s) 208 may be implemented by electronic circuitry. The database 210 can include data that is either stored or generated as a result of functionalities implemented by any of the components of the processing engine(s) 208. In an embodiment, the processing engine(s) 208 can include an extracting unit 212, a merging unit 214, and other unit(s) 216. The other unit(s) 216 can implement functionalities that supplement applications or functions performed by the occupancy detection device 106 or the processing engine(s) 208.

[0054] According to an embodiment, the extracting unit 212 can facilitate extraction of one or more temporal attributes, pertaining to facial key-points and body joints, from the one or more multimedia frames, which are received from the image acquisition unit 102. In an exemplary embodiment, the extracted temporal attributes can further be localized by the other units 216 of the occupancy detection device 106.

[0055] In an implementation, as illustrated in FIG. 3, facial key-points 302-1, 302-2, 302-3... 302-N (also, collectively referred to as facial key-points 302, and individually referred to as facial key-point 302, herein) can be extracted from multimedia frame 300 by the extracting unit 212 for obtaining evidence of human occupancy in the vehicle. The extracted facial keypoints can be used to obtain likelihood of human facial structure present in the multimedia frame 300. In an instance, having high likelihood of human facial structure may provide consequential and adequate cues for the human occupancy. Further, the extracted facial keypoints can be merged with the extracted body joints, and thereby a robust human occupancy detection system can be designed.

[0056] In an exemplary embodiment, the extracting unit 212 can extract various facial key-points, such as, eye, mouth, nose, ears, and the like. These facial key-points can be analyzed for defining human face structure and hence said facial key-points can serve as important evidence of the human occupancy. Further, the extracting unit 212 can also extract various body joints such as, but not limited to, head, shoulder, neck, hip, knee and ankles, and can correspondingly determine related body parts. Furthermore, human body structure can be estimated by aligning said body parts.

[0057] According to another embodiment, the merging unit 214 can facilitate determination of facial structure of the human-entity by mapping the localized the facial keypoints. In other embodiment, the merging unit 214 can facilitate determination of body parts of the human-entity by mapping the localized body joints.

[0058] In an embodiment, the merging unit 214, along with the learning engine 108 including one or more linear CRFs, can form a human pose by aligning the determined facial structure and body parts of the human-entity. In another embodiment, the merging unit 214 can jointly determine the facial structure and body parts of the human-entity by feeding the localized temporal attributes into the learning engine 108.

[0059] In yet another embodiment, the merging unit 214 can also be configured to minimize false positive and false negative detections by jointly determining the facial structure and body parts of the human-entity, and correspondingly merging them for facilitating detection of the human-entity in said section.

[0060] For instance, as illustrated in FIG. 4, flow diagram 400 can include, at Step 1, extraction of body joints 404 along with the facial key-points 302 by the extracting unit 212, which is working in sync with CNN 402 of the learning engine 108. Further, the flow diagram 400 can include, at Step 2, feeding the extracted body joints 404 into the CRF 1, and feeding the extracted facial key-points 302 into the CRF 2. Furthermore, the flow diagram 400 can include, at Step 3, a process of merging of the body joints 404 along with the facial key-points 302, which are extracted at the Step 1, by the merging unit 214, which is working in sync with CRF 1 and CRF 2 of the learning engine 108 for merging cues, wherein the cues pertain to the body joints 404 and the facial key-points 302. In an example, by analyzing the merged cues, the system 100 can determine that left seat of the vehicle is empty and right seat of the vehicle is occupied by an adult.

[0061] Referring to FIGs. 5A and 5B, two parallel output CRF branches with Hourglass design CNN 402 as backbone may be responsible for jointly estimating the body joints and the facial key-points. In an exemplary embodiment, one CRF, say CRF 1, can be trained to obtain the likelihood of human body parts and corresponding structure (human pose estimation) from the extracted body joints. In another exemplary embodiment, one CRF, say CRF 2, can be trained to obtain the likelihood of human face structure (facial key-points estimation) from the extracted facial key-points.

[0062] In an exemplary embodiment, the facial key-points estimation can be defined as the localization of said key-points around vital areas of the face in various images or videos. It can include points around body parts, such as, mouth, nose, eyes, eyebrows, and the like. The facial key-points estimation can greatly improve robustness of the occupancy prediction. In another exemplary embodiment, the human pose estimation can be defined as the of localization of the body joints extracted from the images or videos.

[0063] In another exemplary embodiment, the CRFs could estimate facial and body structure by taking into consideration heat signatures of various body joints and facial keypoints. In an instance, as illustrated in the FIG. 5B, the CRF 1 can obtain heat signatures UV1, UV2, and UV3 associated with respective body joints, say joint 1, joint 2, and joint 3, and correspondingly the CRF 1 can estimate the likelihood of corresponding human body structure.

[0064] Referring to FIG. 6, as illustrated in flow chart 600, for detecting human occupancy, firstly the system 100 can obtain / extract temporal attributes (also, referred to as input cues / cues, herein) from an image or a video, for instance, body joints (also, referred to as human joints, herein) can be extracted at block 602, and facial key-points can be extracted at block 604.

[0065] Further, the system 100 can carry out structure likelihood estimation by processing the temporal attributes. In one exemplary embodiment, at block 606, the extracted body joints can be fed into CRF 1, which can be trained to obtain the likelihood of human body parts and corresponding structure by processing said body joints. In other exemplary embodiment, at block 608, the facial key-points can be fed into CRF 2, which can be trained to obtain the likelihood of human face structure by processing said facial key-points. In an implementation, usage two CRFs, i.e., CRF 1 and CRF 2, for analyzing two distinct types of temporal attributes can ensure minimization of false positives by learning human body and facial structures.

[0066] In an embodiment, the system 100 can detect occupancy based on individual temporal attributes / cues, for instance, the system 100 can detect occupancy separately for each of the body joints and facial key-points. In one embodiment, confidence value of the likelihood of human body parts and corresponding structure is compared with a first threshold value at block 610. In other embodiment, confidence value of the likelihood of human face structure is compared with a second threshold value at block 612.

[0067] In an embodiment, if the confidence value of the likelihood of human structure and the confidence value of the likelihood of human face structure are found to exceed corresponding threshold values, then both the temporal attributes can be merged at block 614. In an implementation, both the temporal attributes can be merged using OR operation, i.e, human occupancy can be detected with respect to any one or both of the temporal attributes. Further, based on the merged temporal attributes and result obtained using the OR operation, the system 100 can derive final occupancy decision at block 616.

[0068] In an exemplary embodiment, merging of two sources of cues can help the system 100 in minimizing false negatives produced by variation in surroundings or occlusions.

[0069] Referring to FIG. 7A, in first scenario 700, when body of a person / human 704 is occluded by an object 702 like briefcase or bag on this / her lap then confidence value of the extracted body joints cannot cross respective threshold value. In said scenario, integration of facial key-points showing high likelihood of forming human facial structure can prevent false negatives, which may have been obtained as outcome of the low confidence value of the extracted body joints. Table 1 can represent confidence values at CRF1 and CRF2, and corresponding predicted occupancy for the scenario 700. CRF1 confidence (Body Joints) CRF2 confidence (facial key points) Predicted Occupancy Without Facial Key points Low - Empty With Facial Key points Low High Occupied TABLE 1

[0070] Referring to FIG. 7B, in first scenario 710, when face of the person 704 holding newspaper 714 gets occluded due to wearing a mask 712, the confidence value of the predicted facial key-points joints cannot cross respective threshold value. In said scenario, integration of body joints showing high likelihood of forming body structure can prevent false positives, which may have been obtained as outcome of the low confidence value of the extracted facial key-points. Table 2 can represent confidence values at CRF1 and CRF2, and corresponding predicted occupancy for the scenario 720. CRF1 confidence (Body Joints) CRF2 confidence (facial key points) Predicted Occupancy Without Facial Key points Low - Empty With Facial Key points Low High Occupied TABLE 2

[0071] Referring to FIG. 8, where a flow diagram of the proposed method 800 for detecting human occupancy in a vehicle (also, referred to as method 800, herein) is shown, the method 800 can include step 802 of acquiring, through an image acquisition unit, one or more multimedia frames of a Region of Interest (Rol) pertaining to a seating section inside the vehicle.

[0072] In another embodiment, the method 800 can include step 804 of receiving, at an occupancy detection device in communication with a Conditional Random Field (CRF) based learning engine, the one or more acquired multimedia frames from the image acquisition unit. Further, the method 800 can include step 806 of extracting, at the occupancy detection device, one or more temporal attributes from the one or more multimedia frames, being received at the step 804, and correspondingly localizing the extracted temporal attributes.

[0073] In an embodiment, the method 800 can include step 808 of determining, at the occupancy detection device, facial structure and body parts of the human-entity by mapping the temporal attributes, being localized at the step 806. Further, the method 800 can include step 810 of merging, at the occupancy detection device, the facial structure and body parts, which are being determined at the step 808, for facilitating detection of the human-entity in said section.

[0074] In an embodiment, the method 800 can include the step of jointly determining the facial structure and body parts of the human-entity by feeding the localized temporal attributes in the learning engine including at least two parallel branches having hourglass design Convolutional Neural Network (CNN), such that first temporal attributes are fed in first branch of the learning engine, and other temporal attributes are fed in respective branches of the learning engine. In an exemplary embodiment, the first temporal attributes can pertain to facial key-points, and the other temporal attributes can pertain to body joints.

[0075] In another embodiment, the method 800 can include the step of minimizing false positive and false negative detections by jointly determining the facial structure and body parts of the human-entity, and correspondingly merging them.

[0076] In an embodiment, the proposed system and method can include a CNN based localization technique that predicts facial key-points along with human body joints. Moreover, having two source of cues / distinct types of temporal attributes for occupancy increases the reliability and robustness of the proposed system and method, as in case CNN-based deep learning technique fails for either facial joints or body joints due to variability in factors like surrounding, occlusions, and the like, the final occupancy detection using CRFs can still be accurate.

[0077] Further, even if there are few false positives, they cannot form human structure, for example, if shoulder, head, and neck are falsely detected when no one occupied the seat, they are most likely not detected in proportions as human body or facial structure by the CRF. In an embodiment, the proposed system and method can deal with both false positive and false negative scenarios very effectively. Therefore, any objects that is falsely detected as person occupied or any person falsely detected as unoccupied can be easily diagnosed and rectified by the proposed system and method.

[0078] In an embodiment, the proposed system and method can correctly detect number of occupants and their seating position in real time, which is very important for many intelligent interior and other such use cases. Hence, the proposed occupancy detection system and method can serve as a key factor in designing the intelligent interior.

[0079] Referring to FIG. 9, computer system 900 includes an external storage device 910, a bus 920, a main memory 930, a read only memory 940, a mass storage device 950, communication port 960, and a processor 970. A person skilled in the art will appreciate that computer system 900 may include more than one processor and communication ports. Examples of processor 970 include, but are not limited to, an Intel® Itanium® or Itanium 2 processor(s), or AMD® Opteron® or Athlon MP® processor(s), Motorola® lines of processors, FortiSOC™ system on a chip processors or other future processors. Processor 970 may include various modules associated with embodiments of the present invention. Communication port 960 can be any of an RS-232 port for use with a modem based dialup connection, a 10 / 100 Ethernet port, a Gigabit or 10 Gigabit port using copper or fiber, a serial port, a parallel port, or other existing or future ports. Communication port 960 may be chosen depending on a network, such a Local Area Network (LAN), Wide Area Network (WAN), or any network to which computer system connects.

[0080] In an embodiment, the memory 930 can be Random Access Memory (RAM), or any other dynamic storage device commonly known in the art. Read only memory 940 can be any static storage device(s) e.g., but not limited to, a Programmable Read Only Memory (PROM) chips for storing static information e.g., start-up or BIOS instructions for processor 970. Mass storage 950 may be any current or future mass storage solution, which can be used to store information and / or instructions. Exemplary mass storage solutions include, but are not limited to, Parallel Advanced Technology Attachment (PATA) or Serial Advanced Technology Attachment (SATA) hard disk drives or solid-state drives (internal or external, e.g., having Universal Serial Bus (USB) and / or Firewire interfaces), e.g. those available from Seagate (e.g., the Seagate Barracuda 7102 family) or Hitachi (e.g., the Hitachi Deskstar 7K1000), one or more optical discs, Redundant Array of Independent Disks (RAID) storage, e.g. an array of disks (e.g., SATA arrays), available from various vendors including Dot Hill Systems Corp., LaCie, Nexsan Technologies, Inc. and Enhance Technology, Inc.

[0081] In an embodiment, the bus 920 communicatively couples processor(s) 970 with the other memory, storage and communication blocks. Bus 920 can be, e.g. a Peripheral Component Interconnect (PCI) / PCI Extended (PCI-X) bus, Small Computer System Interface (SCSI), USB or the like, for connecting expansion cards, drives and other subsystems as well as other buses, such a front side bus (FSB), which connects processor 970 to software system.

[0082] In another embodiment, operator and administrative interfaces, e.g. a display, keyboard, and a cursor control device, may also be coupled to bus 920 to support direct operator interaction with computer system. Other operator and administrative interfaces can be provided through network connections connected through communication port 960. External storage device 910 can be any kind of external hard-drives, floppy drives, IOMEGA® Zip Drives, Compact Disc - Read Only Memory (CD-ROM), Compact Disc - Re-Writable (CD-RW), Digital Video Disk - Read Only Memory (DVD-ROM). Components described above are meant only to exemplify various possibilities. In no way should the aforementioned exemplary computer system limit the scope of the present disclosure.

[0083] Thus, it will be appreciated by those of ordinary skill in the art that the diagrams, schematics, illustrations, and the like represent conceptual views or processes illustrating systems and methods embodying this invention. The functions of the various elements shown in the figures may be provided through the use of dedicated hardware as well as hardware capable of executing associated software. Similarly, any switches shown in the figures are conceptual only. Their function may be carried out through the operation of program logic, through dedicated logic, through the interaction of program control and dedicated logic, or even manually, the particular technique being selectable by the entity implementing this invention. Those of ordinary skill in the art further understand that the exemplary hardware, software, processes, methods, and / or operating systems described herein are for illustrative purposes and, thus, are not intended to be limited to any particular named.

[0084] While the foregoing describes various embodiments of the invention, other and further embodiments of the invention may be devised without departing from the basic scope thereof. The scope of the invention is determined by the claims that follow. The invention is not limited to the described embodiments, versions or examples, which are included to enable a person having ordinary skill in the art to make and use the invention when combined with information and knowledge available to the person having ordinary skill in the art. ADVANTAGES OF THE PRESENT DISCLOSURE

[0085] The present disclosure provides a system and method that enables detection of human occupancy based on facial key-points and body joints, as well as combination of both.

[0086] The present disclosure provides a system and method for detecting human occupancy using CNN and parallel CRFs for processing facial key-points and body joints, simultaneously.

[0087] The present disclosure provides a robust system and method that mitigates false positive results and false negative results while detecting the human occupancy.

[0088] The present disclosure provides a smart, intelligent, efficient, and reliable system and method for detecting human occupancy in the vehicle.

Claims

1. A method (800) for detecting human occupancy in a vehicle, the method (800) comprising the steps of:acquiring (802), through an image acquisition unit, one or more multimedia frames of a Region of Interest (Rol) pertaining to a seating section inside the vehicle;receiving (804), at an occupancy detection device in communication with a Conditional Random Field (CRF) based learning engine, the one or more acquired multimedia frames from the image acquisition unit;extracting (806), at the occupancy detection device, one or more temporal attributes from the one or more received multimedia frames, and correspondingly localizing the extracted temporal attributes;determining (808), at the occupancy detection device, facial structure and body parts of the human-entity by mapping the localized temporal attributes; andmerging (810), at the occupancy detection device, the determined facial structure and body parts for facilitating detection of the human-entity in said section.

2. The method (800) as claimed in claim 1, wherein the method (800) comprises the step of jointly determining the facial structure and body pails of the human-entity by feeding the localized temporal attributes in the learning engine including at least two parallel branches having hourglass design Convolutional Neural Network (CNN), such that first temporal attributes are fed in first branch of the learning engine, and other temporal attributes are fed in respective branches of the learning engine.

3. The method (800) as claimed in claim 2, wherein the first temporal attributes pertain to facial key-points, and the other temporal attributes pertain to body joints.

4. The method (800) as claimed in claim 2, wherein the method (800) comprises the step of minimizing false positive and false negative detections by jointly determining the facial structure and body parts of the human-entity, and correspondingly merging them.

5. A human occupancy detection system (100) implemented in a vehicle, the system (100) comprising:an image acquisition unit (102) configured to acquire one or more multimedia frames of a Region of Interest (Rol) pertaining to a seating section inside the vehicle; andan occupancy detection device (106) in communication with a Conditional Random Field (CRF) based learning engine (108) and the image acquisition unit (102), the occupancy detection device (106) comprising one or more processors coupled to a memory storing instructions executable by the one or more processors, the occupancy detection device (106) configured to:receive, from the image acquisition unit (102), the one or more acquired multimedia frames;extract one or more temporal attributes from the one or more received multimedia frames, and correspondingly localize the extracted temporal attributes;determine facial structure and body parts of the human-entity by mapping the localized temporal attributes; andmerge the determined facial structure and body parts for facilitating detection of the human-entity in said section.

6. The system (100) as claimed in claim 5, wherein the learning engine (108) comprises at least two parallel branches having hourglass design Convolutional Neural Network (CNN).

7. The system (100) as claimed in claim 6, wherein the learning engine (106) is configured to jointly determine the facial structure and body parts of the human-entity by analyzing the localized temporal attributes.

8. The system (100) as claimed in claim 7, wherein the learning engine (108) comprises one or more linear CRFs configured to form a human pose by aligning the determined facial structure and body parts of the human-entity.

9. The system (100) as claimed in claim 7, wherein the system (100) is also configured to minimize false positive and false negative detections by jointly determining the facial structure and body parts of the human-entity, and correspondingly merging them.

10. The system (100) as claimed in claim 5, wherein the temporal attributes pertain to facial key-points and body joints.

Citation Information

Patent Citations

  • Equipment for controlling harmful animals using IoT

    KR1020220116610A

  • Portrait clustering method and device and medium

    CN114333039A

  • System, device, and methods for detecting and obtaining information on objects in a vehicle

    US20220114817A1

  • Systems, devices and methods for vehicle post-crash support

    WO2020136658A1