Systems, devices, and methods for measuring the mass of an object in a vehicle

By processing the 2D and 3D image sequences of the vehicle compartment, generating and analyzing the skeleton model of the occupants, the problems of inaccurate occupants' quality estimation and complex system in the prior art are solved, and an accurate, compact and low-cost occupants' quality measurement system is realized.

CN114144814BActive Publication Date: 2025-06-24GENTEX CORP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202080049934.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-07-09
Filing Date
2020-07-09
Publication Date
2025-06-24
Estimated Expiration
2040-07-09

AI Technical Summary

Technical Problem

The prior art has problems such as inaccuracy, high cost, large equipment size and complex alignment in estimating the quality of occupants in a vehicle, making it difficult to meet the requirements of precision, compactness and low cost.

Method used

The processor is used to process the 2D and 3D image sequences of the vehicle cabin, and the skeleton representation of the occupant is generated through the attitude detection algorithm. The scaling skeleton model is generated based on the depth value, and the skeleton model is analyzed to extract the characteristics of the occupant, and the quality of the occupant is estimated.

Benefits of technology

Accurate estimation of occupants in the vehicle is achieved, reducing the cost and volume of the system, simplifying the alignment process and improving the overall performance of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114144814B_ABST
    Figure CN114144814B_ABST
Patent Text Reader

Abstract

Methods and systems for estimating the mass of one or more occupants in a vehicle cabin are provided, including: obtaining a plurality of images of the one or more occupants, including a 2D (two-dimensional) image sequence and a 3D (three-dimensional) image sequence of the vehicle cabin captured by an image sensor; applying a pose detection algorithm to each of the obtained 2D image sequences to obtain one or more skeletal representations of the one or more occupants and combining one or more 3D images of the 3D image sequence with the one or more skeletal representations of the one or more occupants to obtain a skeletal model and analyzing the skeletal model to extract one or more features of each of the one or more occupants and also processing one or more of the extracted features of the skeletal model to estimate the mass of the one or more occupants.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-reference

[0002] This application claims the benefit of U.S. Provisional Application No. 62 / 871,787, filed Jul. 9, 2019, entitled "SYSTEMS, DEVICES AND METHODS FOR MEASURING THE MASS OF OBJECTS IN A VEHICLE" (Attorney Docket No. GR004 / USP), the entire disclosure of which is incorporated herein by reference. TECHNICAL FIELD

[0003] Some embodiments of the present invention relate to estimating the mass of one or more objects, and more particularly, but not exclusively, to measuring and determining the mass of an occupant in a vehicle and a system for controlling the vehicle based on the measured mass of the occupant, such as an airbag system of the vehicle.

[0004] Incorporated by reference

[0005] All publications, patents, and patent applications mentioned in this specification are incorporated herein by reference to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference. BACKGROUND OF THE INVENTION

[0006] Automobiles equipped with airbag systems are well known in the prior art. Thus, the introduction of airbag systems into automobiles has greatly improved the safety of vehicle occupants. In such airbag systems, an automobile collision is sensed and the airbag is rapidly inflated to ensure the safety of the occupants in an automobile collision. Now, such airbag systems have saved many lives. Unfortunately, if the mass and size of the occupant are very small, such as in the case where the occupant is a child, the airbag may also cause fatal injuries. To address this situation, the National Highway Traffic Safety Administration (NHTSA) has mandated that starting from the 2006 model year, all automobiles should be equipped with an automatic suppression system to detect the presence of a child or infant and suppress the airbag.

[0007] Accordingly, airbag and seatbelt technologies are now being developed to adjust the deployment of airbags based on the severity of a collision, the size and posture of the vehicle occupants, seatbelt usage, and the distance of the driver or occupant from the airbag. For example, an adaptive airbag system utilizes multi-stage airbags to adjust the pressure within the airbag. The greater the pressure within the airbag, the greater the force exerted by the airbag on the occupant when it contacts the occupant. These adjustments allow the system to deploy the airbag with a moderate force for most collisions and reserve the maximum force airbag only for the most severe collisions. An airbag control unit (ACU) communicates with one or more sensors of the vehicle to determine the position, mass, or relative size of the occupant. Information regarding the occupant and the severity of the collision is used by the ACU to determine whether the airbag should be inhibited or deployed, and if deployed, at what various output levels. For example, based on the measured mass of the received occupant (e.g., high, medium, low), the ACU can operate the airbag accordingly.

[0008] Prior art techniques for updating the ACU's previous mass estimates included utilizing mechanical solutions such as pressure pads or optical sensors embedded within the vehicle seat. For example, U.S. Patent No. 5,988,676, titled "Optical weight sensor for vehicular safety restraint systems," discloses an optical weight sensor configured to determine the weight of an occupant sitting on a vehicle seat and mounted between the seat frame and the mounting structure of the vehicle seat. Prior art mass estimation systems and airbag technologies may be less than ideal in at least some respects. Prior weight estimation systems were inaccurate, sometimes providing false mass estimates due to vehicle acceleration. Additionally, the technologies of prior solutions may be larger than would be ideal for use in a vehicle. Further, the cost of prior mass estimation technologies may be greater than would be ideal. Prior spectrometers may be somewhat bulky and may require more alignment than would be ideal in at least some cases.

[0009] In view of the foregoing, improved systems, devices, and methods for estimating and / or measuring and / or classifying the mass of an object within a vehicle interior would be beneficial. Ideally, such systems would be accurate, compact, integrated with other devices and systems such as those of the vehicle, rugged enough, and low cost. SUMMARY OF THE INVENTION

[0010] According to a first embodiment of the present invention, there is provided a method for estimating the mass of one or more occupants in a vehicle cabin, the method comprising: providing a processor configured to: obtain a plurality of images of the one or more occupants, wherein the plurality of images includes a 2D (two-dimensional) image sequence and a 3D (three-dimensional) image sequence of the vehicle cabin captured by an image sensor; apply a pose detection algorithm to each of the obtained 2D image sequences to obtain one or more skeletal representations of the one or more occupants; combine one or more 3D images of the 3D image sequence with the one or more skeletal representations of the one or more occupants to obtain at least one skeletal model for each of the one or more occupants, wherein the skeletal model includes information about the distances of one or more key points of the skeletal model from a viewpoint; analyze the one or more skeletal models to extract one or more features of each of the one or more occupants; process the one or more extracted features of the skeletal model to estimate the mass of each of the one or more occupants.

[0011] In an embodiment, the processor is configured to filter out one or more skeletal models based on predefined filtering criteria to obtain valid skeletal models.

[0012] In an embodiment, the predefined filtering criteria includes specific selection rules defining valid poses or orientations of the one or more occupants.

[0013] In an embodiment, the predefined filtering criteria is based on the measured confidence levels of one or more key points in the 2D skeletal representation.

[0014] In an embodiment, the confidence level is based on the measured probability heatmap of the one or more key points.

[0015] In an embodiment, the predefined filtering criteria is based on a high-density model.

[0016] In an embodiment, the processor is configured to generate one or more output signals, the one or more output signals including the estimated mass of each of the one or more occupants.

[0017] In an embodiment, the output signal is associated with the operation of one or more of the units of the vehicle.

[0018] In an embodiment, the units of the vehicle are selected from: airbags; electronic stability control (ESC) units; seat belts.

[0019] In an embodiment, the 2D image sequence is a visual image of the cabin.

[0020] In an embodiment, the 3D image sequence is one or more of the following: reflected light pattern images; stereoscopic images.

[0021] In an embodiment, the image sensor is selected from: a time-of-flight (ToF) camera; a stereo camera.

[0022] In an embodiment, the pose detection algorithm is configured to identify the pose and orientation of one or more occupants in the obtained 2D image.

[0023] In an embodiment, the pose detection algorithm is configured to: identify a plurality of key points of the one or more occupant body parts in at least one 2D image of the 2D image sequence; link pairs of the detected plurality of key points to generate a skeletal representation of the occupant in the 2D image.

[0024] In an embodiment, the key points are joints of the occupant's body.

[0025] In an embodiment, the pose detection algorithm is the OpenPose algorithm.

[0026] The method according to claim 1, wherein the one or more extracted features are one or more of: shoulder length; torso length; knee length; pelvic position; hip width of the occupant.

[0027] According to a second embodiment of the present invention, there is provided a method for estimating the mass of one or more occupants in a vehicle cabin, the method comprising: providing a processor configured to: obtain a plurality of images of the one or more occupants, wherein the plurality of images includes a 2D (two-dimensional) image sequence and a 3D (three-dimensional) image sequence of the vehicle cabin captured by an image sensor; apply a pose detection algorithm to each of the obtained 2D image sequences to obtain one or more skeletal representations of the one or more occupants; analyze one or more 3D images of the 3D image sequence to extract one or more depth values of the one or more occupants; apply the extracted depth values to the skeletal representation accordingly to obtain a scaled skeletal representation of the one or more occupants, wherein the scaled skeletal model includes information about the distance of the skeletal model from the viewpoint; analyze the scaled skeletal representation to extract one or more features of each of the one or more occupants; process the one or more extracted features to estimate the mass or body mass classification of each of the one or more occupants.

[0028] In an embodiment, the processor is configured to filter out one or more skeletal representations based on predefined filtering criteria to obtain a valid skeletal representation.

[0029] In an embodiment, the predefined filtering criteria includes specific selection rules defining valid poses or orientations of one or more occupants.

[0030] In an embodiment, the predefined filtering criteria are based on a measured confidence level of one or more key points in the 2D skeleton representation.

[0031] In an embodiment, the confidence level is based on a measured probability heatmap of the one or more key points.

[0032] In an embodiment, the predefined filtering criteria are based on a high-density model.

[0033] In an embodiment, the processor is configured to generate one or more output signals, the one or more output signals including the estimated mass or body mass classification of each of the one or more occupants.

[0034] In an embodiment, the output signal corresponds to an operation of one or more of the units of the vehicle.

[0035] According to a third embodiment of the present invention, there is provided a system for estimating the mass of one or more occupants in a vehicle cabin, the system comprising: a sensing device, the sensing device including: an illumination module, the illumination module including one or more illumination sources configured to illuminate the vehicle cabin; at least one imaging sensor configured to capture a 2D (two-dimensional) image sequence and a 3D (three-dimensional) image sequence of the vehicle cabin; and at least one processor configured to: apply a pose detection algorithm to each of the obtained 2D image sequences to obtain one or more skeleton representations of the one or more occupants; combine one or more 3D images of the 3D image sequence with the one or more skeleton representations of the one or more occupants to obtain at least one skeleton model for each of the one or more occupants, wherein the skeleton model includes information about the distances of one or more key points in the skeleton model from a viewpoint; analyze the one or more skeleton models to extract one or more features of each of the one or more occupants; process the one or more extracted features of the skeleton models to estimate the mass of each of the one or more occupants.

[0036] In an embodiment, the processor is configured to filter out one or more skeleton models based on predefined filtering criteria to obtain valid skeleton models.

[0037] In an embodiment, the predefined filtering criteria include specific selection rules defining valid poses or orientations of one or more occupants.

[0038] In an embodiment, the predefined filtering criteria are based on a measured confidence level of one or more key points in the 2D skeleton representation.

[0039] In an embodiment, the confidence level is based on a measured probability heatmap of one or more key points.

[0040] In an embodiment, the predefined filtering criteria are based on a high-density model.

[0041] In an embodiment, the sensing device is selected from: a ToF sensing device; a stereoscopic sensing device.

[0042] In an embodiment, the sensing device is a structured light pattern sensing device, and the at least one illumination source is configured to project modulated light onto the vehicle cabin in a predefined structured light pattern.

[0043] In an embodiment, the predefined structured light pattern is composed of a plurality of diffused light elements.

[0044] In an embodiment, the shape of the light element is one or more of the following: a point; a line; a stripe; or a combination thereof.

[0045] In an embodiment, the processor is configured to generate one or more output signals, the one or more output signals including the estimated mass or body mass classification of each of the one or more occupants.

[0046] In an embodiment, the output signal corresponds to the operation of one or more of the units of the vehicle.

[0047] In an embodiment, the units of the vehicle are selected from:

[0048] An airbag; an electronic stability control (ESC) unit; a seatbelt.

[0049] According to a third embodiment of the present invention, there is provided a non-transitory computer-readable storage medium storing computer program instructions, the computer program instructions causing the processor to perform the following steps when executed by a computer processor: obtaining a sequence of 2D (two-dimensional) images and a sequence of 3D (three-dimensional) images of one or more occupants, wherein the 3D images have a plurality of pattern features according to an illumination pattern; applying a pose detection algorithm to each of the obtained 2D image sequences to obtain one or more skeleton representations of the one or more occupants; combining one or more 3D images of the 3D image sequence with the one or more skeleton representations of the one or more occupants to obtain at least one skeleton model for each of the one or more occupants, wherein the skeleton model includes information about the distances of one or more key points in the skeleton model from a viewpoint; analyzing the one or more skeleton models to extract one or more features of each of the one or more occupants; processing the one or more extracted features of the skeleton models to estimate the mass or body mass classification of each of the one or more occupants. Description of the Drawings

[0050] A better understanding of the features and advantages of the present disclosure will be obtained by referring to the following detailed description of illustrative embodiments and the accompanying drawings, in which the principles of the present disclosure are used.

[0051] Figure 1A and Figure 1B Side views of a vehicle cabin before and after a car accident according to some embodiments of the present disclosure, in which an airbag is activated using a mass estimation system;

[0052] Figure 1C A schematic diagram of an imaging system configured to and capable of capturing images of a scene according to some embodiments of the present disclosure;

[0053] Figure 1D A schematic diagram of a sensing system according to some embodiments of the present disclosure, the sensing system being configured to and capable of capturing a structured light image of reflections of a vehicle cabin including one or more objects and analyzing the captured image to estimate the mass of one or more objects;

[0054] Figure 2A is a block diagram of a processor operating in the imaging system shown in Figure 1B according to some embodiments of the present disclosure;

[0055] Figure 2B is a flowchart showing steps of capturing one or more images of one or more objects and estimating the mass of the objects according to some embodiments of the present disclosure;

[0056] Figure 3A and Figure 3B show two images including a reflected light pattern according to some embodiments of the present disclosure;

[0057] Figure 4A and Figure 4B show a captured image including a skeleton with labeled representations according to some embodiments of the present disclosure;

[0058] Figure 4C-4G show a captured image including a skeleton with labeled representations according to some embodiments of the present disclosure;

[0059] Figure 4H-4K shows a data distribution of measured masses of one or more occupants in a vehicle varying with respective measured body characteristic features of the occupants according to some embodiments of the present disclosure;

[0060] Figure 5A and Figure 5B show examples of captured images of an interior passenger compartment of a vehicle filtered out based on predefined filtering criteria according to some embodiments of the present disclosure;

[0061] Figure 6 is a flowchart showing the generation of a skeletal model of each occupant according to some embodiments of the present disclosure;

[0062] Figure 7A is a schematic high - level flowchart of a method for measuring the mass of one or more occupants in a vehicle according to some embodiments of the present disclosure;

[0063] Figure 7B shows an image including a 3D layer and a skeletal layer combination of a vehicle interior passenger compartment according to some embodiments of the present disclosure;

[0064] Figure 7C is a schematic flowchart of a method for determining the mass of one or more occupants in a vehicle according to other embodiments of the present disclosure;

[0065] Figure 8 is a schematic flowchart of a method for measuring the mass of one or more occupants in a vehicle according to other embodiments of the present disclosure; and

[0066] Figure 9A-9C shows a graph of the mass prediction results of one or more occupants sitting in a vehicle compartment according to an embodiment. DETAILED DESCRIPTION

[0067] In the following description, various aspects of the present invention will be described. For purposes of explanation, specific details are set forth in order to provide a thorough understanding of the present invention. It will be apparent to those skilled in the art that there are other embodiments of the present invention that differ in details without affecting their basic nature. Accordingly, the present invention is not limited by what is shown in the drawings and described in the specification, but only as indicated in the appended claims, where the appropriate scope is determined only by the broadest interpretation of the said claims. The configurations disclosed herein can be combined in one or more of a variety of ways to provide, by analyzing one or more images of a riding object, an improved method, system, and apparatus for measuring the mass of one or more riding objects (e.g., a driver or a passenger) in a vehicle having an interior passenger compartment. One or more components of the configurations disclosed herein can be combined with each other in a variety of ways.

[0068] The systems and methods described herein include obtaining one or more images of a vehicle interior passenger compartment including one or more objects, e.g., one or more occupants (e.g., a vehicle driver or a passenger); and at least one processor configured to extract visual data and depth data from the obtained images, combine the visual data and the depth data, and analyze the combined data to estimate the mass of one or more objects in the vehicle.

[0069] According to other embodiments, a system and method including one or more imaging devices and one or more illumination sources as described herein can be used to capture one or more images of a passenger compartment inside a vehicle, the passenger compartment inside the vehicle including one or more objects, such as one or more occupants (e.g., a vehicle driver or passenger); and at least one processor configured to extract visual data and depth data from the captured images, combine the visual data and the depth data and analyze the combined data to estimate the mass of one or more objects in the vehicle.

[0070] Specifically, according to some embodiments, there is provided a method for measuring the mass of one or more seated objects (e.g., a driver or a passenger) in a vehicle having an interior passenger compartment, the method comprising using at least one processor: obtaining a plurality of images of the one or more occupants, wherein the plurality of images includes 2D (two-dimensional) images and 3D (three-dimensional) images, such as a sequence of 2D images and a sequence of 3D images of a vehicle cabin captured by an image sensor; applying a pose detection algorithm to each of the obtained 2D image sequences to obtain one or more skeletal representations of the one or more occupants; combining the one or more 3D images of the 3D image sequence with the one or more skeletal representations of the one or more occupants to obtain at least one skeletal model for each of the one or more occupants, wherein the skeletal model includes information about the distance of one or more key points of the skeletal model from the field of view; analyzing the one or more skeletal models to extract one or more features of each of the one or more occupants; processing the one or more extracted features of the skeletal model to estimate the mass or body mass classification of each of the one or more occupants.

[0071] According to some embodiments, the imaging device and the one or more illumination sources can be mounted and / or embedded in the vehicle, specifically mounted and / or embedded in the cabin of the vehicle (e.g., near the front mirror or dashboard of the vehicle and / or integrated into the overhead console).

[0072] According to another embodiment, there is provided an imaging system, the imaging system comprising: one or more illumination sources configured to project one or more light beams onto a vehicle cabin including one or more occupants in a predefined structured light pattern; an imaging device including a sensor configured to capture a plurality of images, the plurality of images including, for example, reflections of the structured light pattern from one or more occupants in the vehicle cabin; and one or more processors configured to: obtain a plurality of images of the one or more occupants, wherein the plurality of images includes one or more 2D (two-dimensional) images and 3D (three-dimensional) images, such as a sequence of 2D (two-dimensional) images and a sequence of 3D (three-dimensional) images of the vehicle cabin captured by an image sensor; apply a pose detection algorithm to each of the obtained 2D image sequences to obtain one or more skeleton representations of the one or more occupants; combine one or more 3D images of the 3D image sequence with the one or more skeleton representations of the one or more occupants to obtain at least one skeleton model for each of the one or more occupants, wherein the skeleton model includes information about the distance of one or more key points of the skeleton model from the field of view; analyze the one or more skeleton models to extract one or more features of each of the one or more occupants; and process the one or more extracted features of the skeleton model to estimate the mass or body mass classification of each of the one or more occupants.

[0073] According to some embodiments, the system and method are configured to generate one or more outputs, such as output signals, that may be associated with the operation of one or more devices, units, applications, or systems of the vehicle based on the measured mass. For example, the output signal may include information configured to optimize the unit performance of the vehicle once activated. In some cases, the units or systems of the vehicle may include the vehicle's airbags, seats, and / or electronic stability control (ESC) of the vehicle optimized according to the distribution and measured mass of the occupants.

[0074] Advantageously, the system and method according to an embodiment may include a sensing system including, for example, a single imaging device to capture one or more images of a scene and extract visual data, depth data, and other data such as a speckle pattern from the captured images to detect vibrations (e.g., micro-vibrations) in real time, for example. For example, according to an embodiment, an independent sensing system including, for example, a single imaging device and a single illumination source may be used to estimate the mass classification of the occupants of a vehicle. In some cases, the imaging system may include more than one imaging device and illumination source. In some cases, two or more imaging devices may be used.

[0075] As used herein, like characters refer to like elements.

[0076] Before presenting a detailed description of the present invention, it may be helpful to set forth definitions of certain terms that will be used hereinafter.

[0077] As used herein, the term "mass" encompasses the amount of matter contained in a body, as measured by its acceleration under a given force or by the force exerted on it by a gravitational field. However, as is commonly used, the present invention also refers to the "weight" of an object being measured, where "weight" encompasses the force exerted on the mass of a body by a gravitational field.

[0078] As used herein, the term "light" encompasses electromagnetic radiation having a wavelength in one or more of the ultraviolet, visible, or infrared portions of the electromagnetic spectrum.

[0079] As used herein, the term "structured light" is defined as the process of projecting a known pixel pattern onto a scene. The way these deform when incident on a surface allows a vision system to extract depth information and surface information of the objects in the scene.

[0080] As used in this application, the terms "pattern" and "pattern feature" refer to structured illumination as discussed hereinafter. The term "pattern" is used to denote the form and shape produced by any non-uniform illumination, particularly structured illumination employing multiple pattern features (e.g., lines, stripes, dots, geometric shapes, etc.) having uniform or different characteristics (e.g., shape, size, intensity, etc.). As a non-limiting example, a structured light illumination pattern can include multiple parallel lines as pattern features. In some cases, the pattern is known and calibrated.

[0081] As used herein, the term "modulated structured light pattern" is defined as the process of projecting modulated light onto a scene in a known pixel pattern.

[0082] As used herein, the term "depth map" is defined as an image containing information about the distance of the surfaces of the objects in a scene from a viewpoint. A depth map can be in the form of a grid connecting all points with z-axis data.

[0083] As used herein, the term "object" or "occupant" or "rider" is defined as any target being sensed, including any number of specific elements and / or backgrounds, and including a scene having specific elements. The disclosed systems and methods can be applied to an entire imaging target as an object and / or to specific elements within an imaging scene as objects. Non-limiting examples of "objects" can include one or more persons, such as vehicle passengers or drivers.

[0084] Now referring to the accompanying drawings, Figure 1AFIG. 0 is a side view of a vehicle 110 according to an embodiment, which shows a passenger compartment 105 that includes a vehicle 110 unit and a sensing system 100 configured to and capable of obtaining visual data (e.g., video images) and stereo data (e.g., depth maps), such as 2D (two-dimensional) images and 3D (three-dimensional) images of areas and objects within the vehicle and analyzing the visual data and stereo data, e.g., in real time or near real time, to obtain the mass (e.g., body mass classification) of an object (e.g., an occupant) in the vehicle.

[0085] Specifically, the sensing system 100 is configured to monitor areas and objects within the vehicle 110 to obtain video images and depth maps of the areas and objects and analyze the obtained video images and depth maps using one or more processors to estimate the mass of the object. According to an embodiment, non-limiting examples of such objects may be one or more of the vehicle's occupants (e.g., driver 111 or passenger 112).

[0086] According to some embodiments, the sensing system 100 may be installed, positioned, integrated, and / or embedded in the vehicle 110, particularly in the vehicle's compartment, such that the interior of the compartment and the objects present in the compartment may include, for example, one or more vehicle occupants (e.g., driver, passenger, pet, etc.), one or more objects associated with the compartment (e.g., doors, windows, headrests, armrests, etc.), and the like.

[0087] According to some embodiments, the system and method are configured to generate outputs, such as one or more output signals 106, 107, which may be associated with the operation of one or more of the vehicle's units to control one or more devices, applications, or systems of vehicle 110 based on the measured object mass. For example, output signals 106, 107 including the estimated mass of one or more occupants (e.g., driver 111 and passenger 112) as measured by the sensing system 110 may be transmitted to the ACU 108 and / or the vehicle computing system (VCS) 109, which are configured to activate one or more airbag systems in the event of an accident, such as the variable-intensity airbag system 111' for driver 111 and the variable-intensity airbag system 112' for passenger 112. According to an embodiment, the variable-intensity airbags 111', 112' may have different activation levels (e.g., strong / medium / weak), and the pressure within the variable-intensity airbags is accordingly activated to match the estimated mass classification of the vehicle occupants. In other words, signals may be sent to the ACU 108 or VCS 109, which activate one or more airbags according to the measured category of each occupant. Specifically, an adaptive airbag system may utilize multi-stage airbags to adjust the pressure within the airbags based on the received mass estimate. The greater the pressure within the airbag, the greater the force exerted by the airbag on the occupant when contacting the occupant. For example, as Figure 1B shown, in the case where the weight of driver 111 is about 100 kg and the weight of passenger 112 (a child) is less than 30 kg, during a collision, the airbags of each passenger and driver are deployed in real time according to the estimated mass of the passengers. For example, for the '100 kg' weight driver 111, it is deployed powerfully (high pressure), and for the 30 kg passenger, it is deployed less powerfully (medium or low pressure). Alternatively or in combination, an output including the mass estimation result of each occupant may be transmitted to control the seat belt pretensioner of the vehicle. For example, during a collision, pretension is applied to the seat belts (e.g., seat belts 111", 112") according to the mass estimate so that the passengers are optimally protected.

[0088] In other embodiments, the output including the mass estimation data may be used to optimize the electronic stability control (ESC) of the vehicle according to the occupant distribution in the vehicle; and / or activate or deactivate any of the vehicle's units to which the mass estimation may be related.

[0089] According to some embodiments, system 100 may include, for example, one or more sensors of different types, such as 2D imaging devices and / or 3D imaging devices and / or RF imaging devices and / or vibration sensors (micro-vibrations), etc., to capture sensed data of the vehicle cabin. Specifically, the 2D imaging device may capture images of the vehicle cabin from different angles, for example, and generate raw visual images of the cabin. In an embodiment, system 100 may include an imaging device and at least one processor, the imaging device being configured to capture 2D and 3D images of the vehicle cabin, the at least one processor being used to analyze the images to generate a depth map of the cabin. In another embodiment, system 100 may use one or more vibration sensors to detect vibrations (such as micro-vibrations) of one or more objects in the cabin, and / or analyze the captured 2D or 3D images to identify vibrations (such as micro-vibrations) of the objects.

[0090] According to another embodiment, system 100 may further include a face detector sensor and / or a face detection and / or face recognition software module for analyzing the captured 2D and / or 3D images.

[0091] In an embodiment, system 100 may include a computing unit including one or more processors or may communicate therewith, the one or more processors being configured to receive the sensed data captured by the sensors of system 100 and analyze the data according to one or more of computer vision algorithms and / or machine learning algorithms to estimate the mass of one or more occupants in the vehicle cabin, as will be described below herein.

[0092] Specifically, according to an embodiment, the one or more processors are configured to combine 2D data (e.g., captured 2D images) and 3D data (depth map) of the vehicle cabin to obtain a mass classification of one or more objects (such as vehicle occupants) in the vehicle cabin.

[0093] Advantageously, system 100 provides only minimal hardware, such as one or more sensors and imagers for capturing visual and depth images inside vehicle 110. In some cases, the interface connected to system 100 may supply the necessary power and transmit the acquired data to the computing and / or processing unit of the vehicle (such as VCS 109 and / or ACU 108), where all processing is carried out using its computing power. Thus, according to some embodiments, installing system 100 becomes very easy and uses off-the-shelf components.

[0094] Figure 1CFIG. 0 shows a schematic diagram of a sensing system 102 according to an embodiment, the sensing system being configured to and capable of capturing an image of a scene (e.g., a vehicle cabin) including one or more objects (e.g., driver 111 and / or passenger 112), and analyzing the captured image to estimate the quality of one or more objects. In some cases, the sensing system 102 may be Figure 1A and Figure 1B system 100. According to an embodiment, the system 102 includes: an imaging device 120 configured to and capable of capturing sensing data of one or more objects (e.g., objects 111, 112 in scene 105); and a control unit 150 configured to analyze the captured sensing data to determine the quality of one or more objects.

[0095] Optionally, the imaging device 120 and the control unit 150 are integrated together in a single device. In some cases, the imaging device 120 and the control unit 150 are separately integrated in different devices.

[0096] According to one embodiment, the imaging device 120 may be a ToF (time-of-flight) imaging device, which includes one or more ToF sensors (e.g., continuous wave modulation (CWM) sensors or other types of ToF sensors) for obtaining 3D data of a scene and one or more sensors for obtaining 2D of the scene.

[0097] According to one embodiment, the imaging device 120 may be a stereoscopic imaging device, which includes one or more stereoscopic imagers for obtaining 3D data of a scene and one or more imagers for obtaining 2D of the scene.

[0098] According to one embodiment, the imaging device 120 may be a structured light imaging device, which includes one or more imagers for obtaining 3D data of a scene and one or more imagers for obtaining 2D of the scene, as described below herein in Figure 1D FIG.

[0099] Specifically, in an embodiment, the imaging device 120 includes an illumination module 130 configured to illuminate a scene 105 and an imaging module 123 configured to capture 2D and / or 3D images of the scene. In some cases, the imaging module 123 includes one or more imagers, such as different types of cameras or video cameras, such as cameras 126, 122. For example, camera 126 can capture 3D images or 3D video images of the scene (e.g., for measuring the depth of the scene and the depth of objects in the scene), while camera 122 can capture 2D images of the scene (e.g., raw visual images). For example, camera 126 can be a stereo camera having two or more lenses, e.g., each lens having a separate image sensor, and camera 122 can be a 2D camera. Alternatively or in combination, camera 126 can be a 3D camera adapted to capture the reflection of the diffused light elements of a structured light pattern reflected from an object present in the scene. In some cases, the imaging module 123 can include a single camera configured to capture 2D and 3D images of the scene.

[0100] The illumination module 130 is configured to illuminate the scene 105 using one or more illumination sources such as illumination sources 132, 134. In some embodiments, the illumination module 130 is configured to illuminate the scene using wide beam light (e.g., high intensity floodlight) to achieve good visibility of the scene (e.g., the interior of a vehicle), and accordingly for capturing standard images of the scene. In some embodiments, the illumination module is configured to alternately illuminate the scene using structured light and unstructured light (e.g., floodlight), and accordingly capture 2D and 3D images of the scene. For example, the imaging module 123 can capture one or more 2D images in floodlight and continuously capture 3D images in structured light to obtain alternative depth frames and video frames of the interior of the vehicle. For example, illumination source 132 can be a wide beam illumination source, and illumination source 134 can be a structured light source. In some cases, the 2D and 3D images are captured by a single imager. In some cases, the 2D and 3D images are captured by multiple synchronous images. It should be understood that embodiments of the present invention can use any other type of illumination source and imager to obtain a visual map (e.g., 2D image) and a depth map (e.g., 3D image) of the interior of the vehicle.

[0101] In some embodiments, the 2D and 3D images are properly aligned (e.g., synchronized) with each other such that each point (e.g., pixel) in one image can be found in the other image respectively. This can occur automatically depending on the way the structure is built, or an additional alignment step between the two different modes is required.

[0102] According to one embodiment, the structured light pattern may be composed of a plurality of diffused light elements, such as points, lines, shapes, and / or combinations thereof. According to some embodiments, one or more light sources (such as light source 134) may be lasers or the like, which are configured to emit coherent or incoherent light such that the structured light pattern is a coherent or incoherent structured light pattern.

[0103] According to some embodiments, the illumination module 130 is configured to illuminate a selected portion of the scene.

[0104] In an embodiment, the light source 134 may include one or more optical elements for generating a pattern such as, for example, a spot pattern that uniformly covers the field of view. This can be achieved by using one or more beam splitters, which include optical elements such as diffractive optical elements (DOEs), split mirrors, one or more diffusers, or any type of beam splitter configured to divide a single laser spot into multiple spots. Other patterns such as points, lines, shapes, and / or combinations thereof may be projected onto the scene. In some cases, the illumination unit does not include a DOE.

[0105] According to some embodiments, the imager 126 may be a CMOS or CCD sensor. For example, the sensor may include a two-dimensional array of photosensitive or light-responsive elements, such as a two-dimensional array of photodiodes or a two-dimensional array of charge-coupled devices (CCDs), where each pixel of the imager 126 measures the time it takes for light to travel (to the object and back to the focal plane array) from the illumination module 130.

[0106] In some cases, the imaging module 123 may also include one or more optical bandpass filters, for example, for allowing only light having the same wavelength as the illumination unit to pass through.

[0107] The imaging device 120 may optionally include a buffer that is communicatively coupled to the imager 126 to receive image data measured, captured, or otherwise sensed or acquired by the imager 126. The buffer may temporarily store the image data until the image data is processed.

[0108] According to an embodiment, the imaging device 120 is configured to estimate sensed data, including, for example, a visual image of the scene (e.g., a 2D image) and depth parameters, such as the distance of a detected object to the imaging device. For example, the measured sensed data is analyzed by one or more processors (such as processor 152) to extract 3D data (e.g., a depth map) including the distance of the detected object to the imaging device based on the obtained 3D data and extract the pose / orientation of the detected object from the visual image, and combine the two types of data to determine the quality of the object in the scene 105, as will be described in further detail herein.

[0109] The control board 150 may include one or more processors 152, a memory 154, and a communication circuit 156. The components of the control board 150 may be configured to transmit, store, and / or analyze the captured sensed data. Specifically, one or more processors are configured to analyze the captured sensed data to extract visual data and depth data.

[0110] Figure 1D A schematic diagram of a sensing system 103 according to an embodiment is shown. The sensing system is configured to and capable of capturing a structured light image of a vehicle cabin that includes one or more objects (e.g., a driver 111 and / or a passenger 112) and analyzing the captured image to estimate the quality of one or more objects. In some cases, the sensing system 103 may be Figure 1A and Figure 1B a system 100 of. According to an embodiment, the system 103 includes: a structured light imaging device 124 configured to and capable of capturing sensed data of one or more objects (e.g., objects 111, 112 in a scene 105); and a control unit 150 configured to analyze the captured sensed data to determine the quality of one or more objects.

[0111] Optionally, the imaging device 124 and the control unit 150 are integrated together in a single device. In some cases, the imaging device 120 and the control unit 150 are separately integrated in different devices.

[0112] In an embodiment, the structured light imaging device 124 includes a structured light illumination module 133 and an imaging sensor 125 (e.g., a camera, an infrared camera, etc.) for capturing an image of a scene. The structured light illumination module is configured to project a structured light pattern (e.g., a modulated structured light) onto the scene 105 in one or more spectra, for example. The imaging sensor 125 is adapted to capture the reflected diffused light elements of the structured light pattern reflected from the objects present in the scene. Thus, the imaging sensor 125 may be adapted to operate in the spectrum applied by the illumination module 133 to capture the reflected structured light pattern.

[0113] According to an embodiment, the imaging sensor 125 may include an imager 127 that includes one or more lenses for focusing the reflected light and image from the scene onto the imager 127.

[0114] According to an embodiment, the imaging sensor 125 can capture a visual image of a scene (e.g., a 2D image) and an image including a reflected light pattern, and the image including the reflected light pattern can be processed by one or more processors to extract a 3D image for further measuring the depth of the scene and the depth of an object in the scene by the change encountered when the quantitatively emitted light signal bounces back from one or more objects in the scene, and identifying the object and / or the distance of the scene from the imaging device using the characteristics of the reflected light pattern in at least one pixel of the sensor.

[0115] In an embodiment, the depth data and the visual data (e.g., 2D image) derived from the analysis of the image captured by the imaging sensor 125 are time-synchronized. In other words, since the quality classification is derived from the analysis of a common image captured by the same imaging sensor (of the imaging system), they can also be inherently time (temporally) synchronized, thus further simplifying the correlation of the derived data with the objects in the scene.

[0116] The illumination module 133 is configured to project a structured light pattern on the scene 105, for example, in one or more spectra such as near-infrared light emitted by the illumination source 135. The structured light pattern can be composed of a plurality of diffused light elements. According to some embodiments, the illumination module 133 may include one or more light sources, such as a single coherent or incoherent light source 135, such as a laser, etc., which is configured to emit coherent light such that the structured light pattern is a coherent structured light pattern.

[0117] According to some embodiments, the illumination module 133 is configured to illuminate a selected portion of the scene.

[0118] In an embodiment, the illumination module 133 may include one or more optical elements for generating a pattern such as a spot pattern that uniformly covers the field of view. This can be achieved by using one or more beam splitters, which include optical elements such as diffractive optical elements (DOEs), split mirrors, one or more diffusers, or any type of beam splitter configured to divide a single laser spot into multiple spots. Other patterns such as points, lines, shapes, and / or combinations thereof can be projected on the scene. In some cases, the illumination unit does not include a DOE.

[0119] Specifically, the illumination source 135 can be controlled to produce or emit light in a plurality of spatial or two-dimensional patterns, e.g., modulated light. The illumination can take the form of any of a variety of wavelengths or wavelength ranges of electromagnetic energy. For example, the illumination can comprise electromagnetic energy at wavelengths in the optical range or portion of the electromagnetic spectrum, said wavelengths including those in the human visible range or portion (e.g., about 390 nm - 750 nm) and / or the near infrared (NIR) (e.g., about 750 nm - 1400 nm) or infrared (e.g., about 750 nm - 1 mm) portion and / or the near ultraviolet (NUV) (e.g., about 400 nm - 300 nm) or ultraviolet (e.g., about 400 nm - 122 nm) portion of the electromagnetic spectrum. The specific wavelengths are exemplary and are not intended to be limiting. Other wavelengths of electromagnetic energy can be employed. In some cases, the illumination source 135 wavelength can be any in the range of 830 nm or 840 nm or 850 nm or 940 nm.

[0120] According to some embodiments, the imager 127 can be a CMOS or CCD sensor. For example, the sensor can include a two-dimensional array of photosensitive or light-responsive elements, e.g., a two-dimensional array of photodiodes or a two-dimensional array of charge-coupled devices (COD), where each pixel of the imager 127 measures the time it takes for light to travel (to the object and back to the focal plane array) from the illumination source 135.

[0121] In some cases, the imaging sensor 125 can also include one or more optical bandpass filters, e.g., for allowing only light having the same wavelength as the illumination module 133 to pass through.

[0122] The imaging device 124 can optionally include a buffer communicatively coupled to the imager 127 to receive image data measured, captured, or otherwise sensed or acquired by the imager 127. The buffer can temporarily store the image data until the image data is processed.

[0123] According to an embodiment, the imaging device 124 is configured to estimate sensed data, including e.g., a visual image of a scene and depth parameters, such as the distance of a detected object to the imaging device. For example, the measured sensed data is analyzed by one or more processors (e.g., processor 152) to extract 3D data (e.g., a depth map) including the distance of the detected object to the imaging device from the pattern image and the pose / orientation of the detected object from the visual image, and to combine the two types of data to determine the quality of the object in the scene 105, as will be described in further detail herein.

[0124] The control board 150 may include one or more processors 152, a memory 154, and a communication circuit 156. The components of the control board 150 may be configured to transmit, store, and / or analyze the captured sensing data. Specifically, one or more processors (e.g., processor 152) are configured to analyze the captured sensing data to extract visual data and depth data.

[0125] Optionally, the imaging device 124 and the control unit 150 are integrated together in a single device or system (e.g., system 100). In some cases, the imaging device 124 and the control unit 150 are separately integrated in different devices.

[0126] Figure 2A is a block diagram of a processor 152 operating in one or more of the systems 100, 101, and 102 shown in Figure 1A , 1B , 1C, and 1D according to an embodiment. In the Figure 2A illustrated example, the processor 152 includes a capture module 212; a depth map module 214; an estimation module, such as a pose estimation module 216; an integration module 218; a feature extraction module 220; a filter module 222; a quality prediction module 224; a 3D image data memory 232; a 2D image data memory 234; a depth map representation data memory 236; an annotation data memory 238; a skeleton model data memory 240; and a measurement data memory 242. In alternative embodiments not shown, the processor 152 may include additional and / or different and / or fewer modules or data memories. Similarly, the functions performed by the various entities of the processor 152 may be different in different embodiments.

[0127] In some aspects, the modules may be implemented in software (e.g., subroutines and code). In some aspects, some or all of the modules may be implemented in hardware (e.g., an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), a programmable logic device (PLD), a controller, a state machine, gated logic, discrete hardware components, or any other suitable device) and / or a combination of both. Additional features and functions of these modules in accordance with various aspects of the subject technology are further described herein.

[0128] Optionally, the modules may be integrated into one or more cloud-based servers.

[0129] The capture module 212 obtains an image of a scene (e.g., a passenger compartment inside a vehicle) of one or more objects (e.g., one or more passengers or a driver in a vehicle) in the scene. In one embodiment, the processor 152 instructs one or more sensors (e.g., Figure 1C the imaging device 120 shown in Figure 1DThe imaging device 124) shown captures an image of a scene to further extract 2D and 3D data (e.g., images) of an object in the scene or the scene itself. As an example, the capture module 212 can obtain, for example, a sequence of visual images (e.g., raw 2D images) and a sequence of images including depth data (e.g., 3D images) of one or more objects in the scene synchronously and / or sequentially. According to an embodiment, the 2D image is analyzed to determine the pose and orientation of the object, while the 3D image is analyzed to generate a depth map representation of the object, as will be explained in detail below.

[0130] In one embodiment, as described above, the capture module 212 obtains a 3D image of an object illuminated by an illuminator and / or an image obtained by a stereo sensor and / or a ToF sensor, where the illuminator projects structured light with a specific illumination pattern onto the object. The captured image of the object provides useful information for the future generation of a depth map. For example, the captured image of an object illuminated with structured light includes specific pattern features corresponding to the illumination pattern projected onto the object. The pattern features can be stripes, lines, dots, or other geometric shapes, and include uniform or non-uniform characteristics such as shape, size, and intensity. In some cases, the image is captured by other sensors such as a stereo sensor or a ToF sensor, and the depth data is presented in a different manner. Figure 3B An exemplary captured image 310 illuminated with a specific structured light (e.g., dots) is described. The captured image 310 includes two occupants 315 (e.g., the driver) and 325 (a child passenger) sitting in the front seat of a vehicle. In some cases, the captured image 310 and associated image data (e.g., the intensity, depth, and gradient of each pixel) are stored in the 3D image data memory 232, and the captured visual image (e.g., 2D image) and associated image data are stored in the 2D image data memory 234, as described more fully below.

[0131] The depth map module 214 retrieves the captured 3D image of the illuminated object from the image 3D data memory 242 and generates a depth map representation of the object based on the captured image (e.g., the pattern image) of the illuminated object. As described above, the depth map representation of an object refers to an image containing information about different parts of the surface of the object and / or the distance between the scene and a specified viewing point. The specified viewing point can be the location of the sensor that captured the image of the object. In an embodiment, the depth map representation is stored at the depth map representation data memory 236, as described more fully below. Exemplary depth map representations are further described below with reference to Figure 3A Further describe exemplary depth map representations.

[0132] In one embodiment, the depth map module 214 identifies and analyzes pattern features for deriving depth information of a captured image. Based on the identified and analyzed pattern features associated with an object, the depth map module 214 generates a depth map representation of the object. An example of depth information can be the geometric deformation of an object due to depth differences of each pixel on the object in the captured image. The "depth" of a pixel on an object refers to the distance between the pixel on the actual object and a specified viewing point (e.g., the position of the sensor).

[0133] In some embodiments, the depth map module 214 generates a depth map representation of an object in a captured image based on triangulation between a light pattern and an image sensor, and the depth of the object illuminated by the light pattern can be extracted. The detected pattern refers to the pattern projected onto the object and rendered in the captured image, and the reference pattern refers to the original illumination pattern provided by the illuminator. For structured light with an illumination pattern projected onto an object, the pattern detected in the captured image of the object is a distorted version of the original illumination pattern of the structured light. The distorted version of the original pattern includes shifts and other distortions due to the depth of the object. By comparing the detected pattern with the original illumination pattern or comparing portions of the detected pattern with corresponding portions of the original illumination pattern, the depth map module 214 identifies the offsets or distortions and generates a depth map representation of the object.

[0134] Figure 3A An example of a captured image 335 including reflected light pattern spots according to an embodiment is shown. For illustration, each spot of the reflected light pattern spots is colored using a grayscale color, where each color represents the distance of the spot from a reference point (e.g., the camera). For example, the scale 282 includes a grayscale color for a distance of approximately 40 cm from the camera, and the color continuously changes to a black scale within a distance of approximately 140 cm from the camera, and so on, with the color scale changing according to the distance. Thus, analyzing Figure 3B the multi-pattern spots 281 in the captured image 285 shown to obtain a depth representation image 287 of the scene as shown in Figure 3A For example, the cluster of reflected points on the driver's leg (presented by the ellipse 345) is typically approximately 20 - 50 cm from the camera, while the driver's center of mass (presented by the ellipse 355) is farther from the camera (approximately 50 - 80 cm). According to an embodiment, the depth map module 214 receives and analyzes an image of a vehicle cabin (e.g., image 335) to extract a depth map representation based on the distance between the detected pattern in the captured image and the sensor, the depth map representation including depth values for each reflected pattern position (e.g., pixel) in the captured image.

[0135] The pose estimation module 216 retrieves from the 2D image data memory 234 the original image (2D image) of the illuminated occupants (usually one or more persons) of the captured vehicle, and analyzes the original image to identify one or more persons in the image and further estimate their poses. In some cases, the identification includes generating a graphical representation, such as a skeleton of points superimposed on each identified person in the captured image. In some embodiments, the image including the superimposed skeleton is stored at the annotated data memory 238.

[0136] In one embodiment, the pose estimator module 216 uses a DNN (deep neural network) to identify one or more persons in each retrieved image, and superimposes (e.g., marks) a plurality of annotations, such as selected key point positions, at the identified objects. If the object is an identified person (e.g., a passenger or a driver), the key points represent body landmarks (e.g., joint body points) detected at the body image of the captured person. According to an embodiment, the detected key points can be graphically represented as the key points or the framework of the skeleton of the identified person. According to an embodiment, each key point of the skeleton includes coordinates (x, y) at the human body image. In some cases, the skeleton is formed by linking every two key points by connecting lines between two lines as shown in Figure 4A and Figure 4B .

[0137] The integration module 218 obtains the formed skeleton (e.g., 2D skeleton) and the depth map representation of each object, and combines (e.g., blends) them to obtain a skeleton model, e.g., a 3D skeleton including 3D data of each object. In an embodiment, the integration process includes computationally combining the formed skeleton (2D skeleton) and the depth map representation to obtain a skeleton model that includes data of each key point in the skeleton model in the (x, y, z) coordinate system. In an embodiment, the skeleton model includes depth data related to each joint key point at the formed skeleton model, e.g., the position (x, y) of each point of the person in the scene and the distance (z) of that point from the corresponding image sensor in the (x, y, z) coordinate system. In other words, each key point of the formed skeleton has coordinates in the 2D image. Since the captured 2D and 3D images are co-registered with each other, according to an embodiment, the 3D values of the same coordinates in the 3D map can be obtained. Thus, the Z value (e.g., distance) of some key points or each key point is obtained. An example of the combination process is shown in Figure 6 .

[0138] In some cases, the skeleton model data is stored at the skeleton model data memory 240.

[0139] In an embodiment, the feature extraction module 220 is configured to and capable of analyzing the skeleton model data and extracting one or more data measurements of each relevant identified person at the scene. Generally, the extracted measurements include data related to the imaged person, as well as output export values (e.g., features) and non-redundant information about the person that are intended to provide information. Specifically, the extracted features of an imaged occupant (e.g., a person) in a vehicle may include the measured lengths of the occupant's body parts, such as the occupant; torso; length of the shoulders; hip width and pelvic position, etc.

[0140] Generally, estimating the mass of a seated person, even by the human eye, is much more difficult than estimating the mass of a standing person because the main body parts of the person (e.g., legs, knees, or hands) are hidden and / or not fully presented. Specifically, estimating the mass of an object of interest such as a person based on body part measurements (e.g., skeleton measurements) in a vehicle is thus challenging because the person's skeleton is visible in a highly non-standard pose (e.g., sitting or "bent knee" pose). According to an embodiment, it is necessary to identify these non-standard poses (e.g., "bent knee" pose) and avoid using them in the mass estimation process in order to produce an accurate mass estimate. According to an embodiment, the filter module 222 is configured to address this problem by obtaining an image including the skeleton model data of the object from the skeleton model data memory 232 and filtering out one or more of the obtained images based on a predefined filtering criterion (e.g., a predefined filtering criterion). The remaining valid images including the skeleton model data (e.g., valid skeleton model data) can be saved at the skeleton model data memory 240 for further determination of the mass of the object.

[0141] In some cases, the predefined filtering criterion includes specific selection rules that define the valid postures, poses, or orientations of the object and further discard "abnormal" poses. According to an embodiment, an "abnormal" pose can be defined as the body pose (e.g., marked by the skeleton model) or body part of the object that does not reflect or present all or almost all of the main parts of the object. Non-limiting examples of the filtering criterion include: the defined spatial relationship between the skeleton features of the identified object and / or the identified abnormal poses; short imaged body part lengths, object image positions located away from high-density regions.

[0142] According to an embodiment, the defined spatial relationships between the skeletal features of the identified object include, for example, predefined relationships between object parts. Specifically, in the case where the object is an occupant sitting in a vehicle, the criteria include the defined spatial relationships between the occupant's body parts, such as the relationship between the occupant's shoulders and torso or hands, the relationship between the torso and knees in a sitting posture, etc. In some cases, (e.g., in a sitting posture) the spatial relationships between the measured skeletal occupant body organs (e.g., knees; shoulders; hands) are measured and compared with predefined body proportion parameters. For example, as Figure 5B shown in, the spatial relationship between the shoulders and torso of the driver 522 does not match the predefined proportion parameter, so the image 501 will be discarded. In some cases, data related to the occupant's body proportions and characteristics is stored at the annotation data memory 238 and retrieved by the filtering module 222 to discard or confirm the captured image and / or skeletal model of the occupant accordingly.

[0143] According to an embodiment, the high-density filtering criteria include generating a high-density model (e.g., a high-dimensional space vector such as an eight-dimensional vector) based on the measured parameters of one or more vehicle parameters. The high-density model may include, for each key point (body joint) of the identified person, an allowable region where the key point may be located in the captured image. If the high-density model identifies that the key point is outside this region, this image is discarded. The allowable region for each joint is provided by analyzing images with a good "standard" sitting posture.

[0144] In some cases, the generated high-density parameters are stored at the sensing system (e.g., system 100, 102, or 103) (e.g., at the processor 152 or storage device 154, or at a remote processor or database such as a cloud data memory). Then, each generated skeleton is placed in and / or compared with the generated high-dimensional space to determine the position of the skeleton in space relative to the high-density model. Thus, images including skeletons that are not within a predetermined distance from the high-density region in this space are discarded. For example, skeletons generated that are located far from the density center will be filtered out.

[0145] According to an embodiment, the quality prediction module 224 obtains a valid image of an object from the skeleton model data storage 240 and analyzes the valid image to determine the quality of the object. In some embodiments, the analysis includes inserting features of the extracted valid skeleton data into a regression module, such as a pre-trained regression module configured to and capable of estimating quality. In some cases, the pre-trained regression module may use "decision trees" trained according to, for example, the XGBoost method, where each decision tree represents the measured quality of the object in each captured image based on the measured features of the object. For example, each formed tree may include data on the features of the captured occupant (e.g., the length of the occupant's shoulders; the length of the torso; the length of the knees measured based on the valid image). According to an embodiment, the pre-trained regression module is used to optimize the quality estimation process of the occupant to provide the most accurate quality prediction (e.g., estimated prediction) for each captured object (e.g., a person). It should be understood that according to an embodiment, other types of pre-trained methods may be used.

[0146] In some embodiments, as described more fully below, the measured quality of each object is stored at the quality measurement data storage 242.

[0147] The 3D image data storage 232 of the processor 152 stores 3D images of a specific object (e.g., a person) or a scene (e.g., a vehicle cabin) captured and the image data related to the captured images. In an embodiment, the captured 3D image stored in the 3D image data storage 232 may be an image including specific pattern features corresponding to the illumination pattern projected onto the object. For example, the image may include one or more reflected light spots as shown in Figure 3A . In other embodiments, the 3D image may be an image obtained from a stereo camera or a ToF sensor or any known 3D capture device or method.

[0148] The depth map data storage 234 of the processor 152 stores the depth map representation of the object generated by the depth map module 214 and related data. For example, the depth map data storage 234 stores the original depth map representation and related depth data as well as the enhanced depth map representation and related depth data. As described above, the original depth map representation refers to the depth map representation derived from the originally captured image.

[0149] The depth map representation data storage 236 of the processor 152 stores the depth map representation of the object generated by the depth map module 214 and related data. For example, the depth map data storage 234 stores the original depth map representation and related depth data as well as the enhanced depth map representation and related depth data. Specifically, in some cases, the related data may include the image representation of the light pattern of the image according to the measured distance of each image pixel from the image sensor.

[0150] The labeled data memory 238 of the processor 152 stores the skeletal representation of the object and related data generated by the pose estimation module 216. For example, the labeled data memory 234 stores the original 2D image of each object and the related superimposed skeleton. According to one embodiment, the labeled data memory 234 may also store the related data of each pixel or key point at the skeleton, such as one or more confidence levels. The confidence level may be defined as the intensity level of the key point heat map, recognized by, for example, the pose estimation module 216. For example, the pose estimation module 216 may include or use a DNN to provide a "probability heat map" for some or each key point at the captured image. In an embodiment, the "probability heat map" of each key point may be stored, for example, at the labeled data memory 238. For each skeleton point, the DNN (e.g., the pose estimation module 216) represents the confidence, relevance, and accuracy of the position of the key point of the generated skeleton by adjusting the intensity of the maximum point in the probability map. For example, as Figure 4A shown, for each key point at the skeleton 411 (e.g., key points 442, 443, 444, 452, 453, 462, 463, 464, 465, 472, 473, and 474), a probability heat map related to the confidence rating of the DNN is generated. The probability heat map may also be used, for example, together with the density standard score to determine the confidence level of the skeleton key point and accordingly approve or discard the captured image.

[0151] In some cases, the original image is divided according to the recognized objects in the image. For example, an image of a captured vehicle cabin is divided into one or more images according to each vehicle seat (e.g., the front seat or the rear seat).

[0152] Figure 4A and Figure 4B respectively show examples of images 410, 420 of the captured vehicle interior passenger compartments 412, 422, which include two imaged occupants, namely the driver 404 and the passenger 406 in the image 410 and the driver 416 and the passenger 418 in the image 422. In an embodiment, the images 410, 420 are captured by an imaging device mounted on the front section of the vehicle (e.g., on or near the front center mirror). In an embodiment, the pose estimator module 242 retrieves the original visual images of the captured illuminated objects (e.g., occupants 404, 406, 416, and 418) from the 2D image data memory 244 and generates skeletons on the people in the images (e.g., skeletons 411 and 421 on the captured images 410, 420). For example, as Figure 4AAs shown, the object on the left side of the captured image 410 is recognized as a person (e.g., passenger 406) sitting on the front passenger seat of the vehicle, where the skeleton 411 is superimposed on the center body of the passenger. In an embodiment, the pose estimation module 216 is configured to identify and locate the main parts / joints (e.g., shoulders, ankles, knees, wrists, etc.) of the passenger's body by detecting landmarks (e.g., key points) on the identified object and linking the identified landmarks with connecting lines. For example, as Figure 4A shown, the skeleton generation process performed by the pose estimation module 246 includes: identifying key points 442, 443 linked by line 444 for estimating the passenger's shoulders; identifying key points 452, 453 for estimating the passenger's torso; identifying key points 462, 463, 464, 465 for estimating the passenger's knees; identifying key points 472, 473, 474 for estimating the passenger's right hand, and identifying key points 482, 483 for estimating the passenger's left hand. According to an embodiment, the key points are obtained by a DNN trained to identify specific key points.

[0153] In some cases, for each identified key point, a "probability map" is applied to obtain a confidence level that defines the accuracy of the identified key point.

[0154] In some cases, when the vehicle includes multiple seats (e.g., rear seats, front seats, driver seats, infant seats, etc.), and the captured image includes multiple occupants sitting on different seats, the module can, for example, separately identify each seat and the occupants sitting on the identified seats, in order to generate skeletons for each object (e.g., passengers and / or drivers) accordingly. For example, as Figure 4B shown, the driver seat and the driver can be identified, and thus, a second skeleton can be superimposed on the identified driver.

[0155] According to an embodiment, once a skeleton representation is generated for one or more objects (e.g., for each object) by one or more processors (e.g., processor 152), one or more skeleton characteristics of the object are analyzed to estimate the quality of the object. For example, as Figure 4C-4G shown, an image 480 of the rear seat of the vehicle interior compartment including two occupants 482, 484 is captured. For example, the image 480 can be one frame of multiple captured frames of the vehicle compartment. According to an embodiment, the skeleton 486 superimposed on the occupant 484 is analyzed to obtain the body (e.g., skeleton) characteristics of the occupant, such as the length of the shoulders ( Figure 4C ), hips ( Figure 4D ), torso ( Figure 4E ), legs ( Figure 4F ) of the occupant 484, and the center of mass ( Figure 4G)。In some cases, the mass of the occupant is estimated based on these five measured skeleton parts. In other embodiments, different and / or additional body organs of the occupant or elements in the occupant's surroundings may be measured.

[0156] Figure 4H-4K Shows the estimated mass of one or more occupants in a vehicle according to an embodiment as a function of the respective measured body characteristic features of the occupant, such as the shoulder ( Figure 4H ), torso ( Figure 4I ), hip ( Figure 4J ), and leg ( Figure 4K ). Figure 4H-4K The vertical lines at each graph of Figure 4I represent one or more bodies moving around the identified object. The mass estimate of each body part (e.g.,

[0157] the torso length estimate of Figure 4H-4K ) includes noise and inter-image variations of the same object (e.g., the torso of a person). Advantageously, combining different measurements of different body parts of the same object results in an accurate mass estimate.

[0158] In one embodiment, the pose estimation module 216 processes each image using one or more filters obtained, for example, from the filter data memory to check and generate a confidence level. The confidence level is based on the reliability and / or accuracy of the formed skeleton and is particularly used to verify the reliability and accuracy of each identified key point. In some cases, the confidence level may be determined based on the confidence level rating measured by the pose estimation module 216 (e.g., DNN) and the density standard score measured using the pose density model.

[0159] According to an embodiment, the pose density model obtains the skeleton of each image from the pose estimation module 216 and places the skeleton configuration of each object in a high-dimensional space to discard any configuration within a predetermined distance from the high-density region in this space. In some cases, this distance is determined by the Euclidean distance between the 8-vector of the key points of the current frame and the average point calculated from the complete training data. In one embodiment, the confidence probability is configured based on the local density of the skeleton in the skeleton space. In some cases, temporal smoothing is performed on the obtained estimates to reduce noise and fluctuations.

[0160] Figure 2BFIG. 250 is a flow chart according to one embodiment, which shows steps of capturing one or more images 252 of one or more objects (e.g., objects 254, 255 in scene 256) and estimating the quality of the objects. In some cases, the scene may be the interior of a vehicle with one or more seats, and the objects are one or more passengers and / or a driver sitting on these seats. As Figure 2B shown, an imaging device 262 including one or more illuminators 265 provides structured light with a specific illumination pattern (e.g., a spot pattern) to the objects 254, 255 in the scene 256, and a sensor 266 captures one or more images of the objects 254, 255 in the scene 256. In other embodiments, the device 262 may be or may include a stereo imager or a ToF imager. For example, the imaging device may be a ToF imaging device, and the illuminator includes an illumination source configured to project light onto the scene, and the sensor is a ToF sensor configured to capture a plurality of images including reflections of the modulated structured light pattern from one or more objects in the scene. In some cases, the imaging device 262 may be a stereo imaging device including a stereo imager known in the art.

[0161] In various embodiments, the projected light pattern may be, for example, a spot pattern that uniformly covers the scene or a selective part of the scene. When the light is projected into the scene, the spots from the light pattern fall on one or more objects of interest. In some cases, the light is projected by the illuminator 265 using a diffractive optical element (DOE) to split a single laser spot into multiple spots, as Figure 1B described. Other patterns such as points, lines, shapes, and / or combinations thereof may be projected onto the scene. In some cases, the illumination unit does not include a DOE.

[0162] In some cases, each reflected light pattern (e.g., a spot) is covered by one or more of the ToF sensor pixels 266. For example, each spot may be covered by a 5×5 pixel window.

[0163] In one embodiment, the processor 152 may instruct the illuminator 265 to illuminate the objects 254, 265 with a specific modulated structured light. One or more reflected pattern images 260 and clear images 270 (e.g., visual images excluding the reflected light pattern) are provided to the processor 152 to generate a depth map representation 264 and a skeletal model representation 266 of the objects 254, 255. To generate the skeletal model representation 266 of each of the objects 254, 255, the processor 152 first identifies the pose and / or orientation (272) of the captured objects in the original image (270) by associating each point in the scene space 256 with a specific part of the object. For example, in the case where the objects 254, 255 are two vehicle passengers (e.g., persons), each point or selected points in the passenger image are linked to a specific body organ, such as the legs, torso, etc. Then, the processor 152 filters the identified points (274) by checking the reliability of each identified object based on, for example, a measured confidence level (as described above) and applying the confidence level of each identified point in the space. Thereafter, in some cases, the processor 152 splits the captured image into one or more images (276) according to the identified pose and / or orientation and / or confidence level of the identified objects to generate a skeletal representation (278) of each identified object. In some cases, the position and orientation of the objects can be detected and measured by applying the OpenPose algorithm and / or other DNN algorithms (e.g., DensePose configured to extract body poses) to the images.

[0164] According to an embodiment, to generate the depth map representation 264 of the objects 254, 255, the processor 152 analyzes the rendered reflected pattern features and / or ToF data and / or stereo data in the captured image 260 to obtain the depth of each reflected pattern from a reference point, e.g., distance. In some cases, the pattern is a spot-shaped pattern, and the generated depth map representation 264 includes a grid of points superimposed on the captured image 252, where each point indicates the depth of the surface of the image, as Figure 3A shown. The processor 152 then integrates (e.g., combines) the depth map representation (264) with the skeletal annotation representation (278) to obtain the skeletal model representation (266) of each object. According to an embodiment, the skeletal model representation (266) of each object is then analyzed by the processor 152 to extract object features (268), e.g., the length or width of a body part in the case where the identified object is a person. Non-limiting examples of the extracted features may include the length of the body parts of the object, such as the shoulder, torso, knee length.

[0165] In some embodiments, the processor 152 filters the skeleton model rendered images according to predefined filtering criteria to obtain one or more valid skeleton model renders (269). In some cases, the filtering criteria are based on the measured confidence ratings of each identified point and one or more selection rules as described herein with respect to Figure 2A and the like. In some cases, a rank is assigned to each analysis frame, reflecting the accuracy and reliability of the shape and position of the identified object.

[0166] According to an embodiment, based on the extracted features, the processor 152 determines the quality of each object (280). For example, the extracted features of each captured image are inserted into a quality model (e.g., a pre-trained regression quality model), which receives the extracted object features of each acquired image over time (t) to determine the quality (280) or quality classification (282) of each object in the scene. In an embodiment, the quality model takes into account previous quality predictions, such as obtained from previous image processing steps, to select the most accurate quality prediction result. In some embodiments, the quality model also takes into account the measured rank of each skeleton model, and optionally also takes into account the provided confidence level, to obtain the most accurate quality prediction result.

[0167] In some embodiments, a temporal filter is activated to stabilize and remove outliers, such that a single prediction is provided at each timestamp. For example, temporal filtering can include removing invalid images and determining quality predictions based on previous valid frames. If the desired output is a continuous quality value (e.g., which can include any numerical value, e.g., 5, 97.3, 42.1, 60, etc.), then it is the output of the temporal filter.

[0168] Alternatively or in combination, the quality classification (282) of each identified object (e.g., objects 254, 255) in the scene can be determined according to a number of predefined quality categories such as child, adolescent, adult. For example, a vehicle passenger weighing 60 kg and / or between 50 - 65 kg will be classified as "small adult" or "adolescent", while a child weighing 25 kg or within the range of 25 kg will be classified as "child".

[0169] Figure 5A and 5B illustrates examples of captured images 500, 501 of the interior passenger compartment of a vehicle filtered out according to a predefined filtering criterion according to an embodiment. As Figure 5A and Figure 5BAs shown, each of the acquired images 500, 501 includes a graphical skeleton representation formed by a number of lines superimposed on the main body part of the occupant. According to an embodiment, each of these skeletons is analyzed to distinguish images including "abnormal" or "invalid" postures, and selected images (e.g., valid images including valid postures) are retained, and the selected images will be used for further accurate occupant mass measurement for calcification.

[0170] Figure 5A The captured image 500 of the passenger 512 sitting on the passenger front seat of the vehicle is shown, and the captured image will be filtered out based on predefined filtering criteria, for example, due to the measured short torso. Specifically, as shown in the image 500, the passenger 512 is tilted forward relative to the imaging sensor, and thus, the torso length 516 of the passenger of the measured skeleton 518 (e.g., the length between the neck and the pelvis measured between the skeleton points 511 and 513) is shorter than the shoulder width. Since it is difficult and inaccurate to estimate the mass based on such positions shown in the image 500 (due to the measured short torso defined in the predefined filtering criteria), this image will be filtered out from the captured images of the vehicle cabin.

[0171] Figure 5B The captured image 501 of the driver 522 sitting on the driver's seat of the vehicle and the shaped skeleton 528 superimposed on the driver's upper body are shown. According to an embodiment, the image 501 will be filtered out from the list of captured images of the vehicle cabin because the captured image of the driver's body emphasized by the skeleton 528 is located away from the high-density area. Specifically, the body image of the driver is located in a low-density area in the "skeleton configuration space" (e.g., the skeleton fits at a location far from the density center), which means that the body of the identified person is tilted from the "standard" sitting posture, and thus the joints are further away than the allowed positions.

[0172] Figure 6 is a flowchart 600 according to an embodiment, which shows that by combining Figure 3B the depth map representation 335 of the occupant shown in Figure 4A and 4B the 2D skeleton representation image 422 of the occupant shown in Figure 3B the skeleton model (3D skeleton model 650) of the occupant shown in Figure 6As shown, the captured image 325 renders pattern features on the captured occupant (e.g., a person). The depth map representation 335 of the occupant is derived from the captured image 325 of the occupant, where the depth map representation 335 of the occupant provides depth information of the occupant, and the 2D skeleton representation 422 provides information about the pose, orientation, and dimensions of the occupant. The skeleton model 650 is created by combining the depth map representation 335 of the occupant and the 2D skeleton representation 422 of the object. According to an embodiment, the skeleton model 650 is created by applying depth values (e.g., calculated from the nearest depth points around the point) to each skeleton key point. Alternatively or in combination, the average depth in the skeleton region can be provided as a single constant number. This number can be used as a physical "scale" for each provided skeleton, as further explained with respect to Figure 8 Further explained.

[0173] Figure 7A is a schematic high-level flowchart of a method 700 for measuring the mass of one or more occupants in a vehicle according to an embodiment. For example, the method may include determining, in real time, the mass of one or more occupants sitting on a vehicle seat according to one or more mass classification categories, and correspondingly outputting one or more signals to activate and / or provide information associated with the activation of one or more of the vehicle units or applications. Some stages of the method 700 may be performed at least in part by at least one computer processor, such as by the processor 152 and / or the vehicle computing unit. A corresponding computer program product may be provided, which includes a computer-readable storage medium having a computer-readable program included therein and configured to perform the relevant stages of the method 700. In other embodiments, the method includes steps different from or additional to those described in connection with FIG. 7. Additionally, in various embodiments, the steps of the method may be performed in an order different from the order described in connection with Figure 7A described. In some embodiments, some steps of the method are optional, such as the filtering process.

[0174] At step 710, according to an embodiment, a plurality of images including one or more visual images are obtained, and the visual images are, for example, a 2D image sequence and a 3D image sequence of a vehicle cabin. The obtained 2D and 3D image sequences include images of one or more occupants (e.g., a driver and / or a passenger sitting in the rear and / or back seats of the vehicle). According to some embodiments, the 3D image is an image including a reflected light pattern and / or ToF data and / or any stereoscopic data, and the 2D image is a clear original visual image that does not include additional data such as a reflected light pattern. In some embodiments, an image sensor located in the vehicle cabin, such as at the vehicle front section shown in Figure 1A captures a plurality of images (e.g., 2D and 3D images) synchronously and / or sequentially. In some cases, the images are obtained and processed in real time.

[0175] At step 720, one or more pose detection algorithms are applied to the obtained 2D image sequence to detect the pose and orientation of an occupant in a vehicle cabin. Specifically, the pose detection algorithm is configured to identify and / or measure features such as position; orientation; body organs; the length and width of the occupant. For example, the position and orientation of an object can be detected and measured by applying the OpenPose algorithm to the image and / or dense pose. Specifically, according to an embodiment, a neural network such as a DNN (Deep Neural Network) is applied over time (t) to each obtained 2D image to generate (e.g., superimpose) a skeleton layer on each identified occupant. The skeleton layer can include a plurality of key point positions describing the joints of the occupant. In other words, the key points represent body landmarks (e.g., joint body points) detected at the captured body image forming the skeleton representation as shown in Figure 4A and 4B According to an embodiment, each key point of the skeleton representation includes the identified coordinates (x, y) at the occupant body image for extracting features of the identified occupant.

[0176] In some embodiments, the pose estimation method can further be used to identify the occupant and / or the seat of the occupant in each obtained 2D image.

[0177] In some embodiments, the pose estimation method is configured to extract one or more features of the occupant and / or the environment around the occupant, such as the body parts of the occupant and the seat position of the occupant.

[0178] In some embodiments, the identified occupants are separated from each other to obtain a separate image of each identified occupant. In some embodiments, each separate image includes the identified occupant and optionally the environment around the occupant, such as the seat of the occupant.

[0179] In some embodiments, each obtained 2D image in the 2D image sequence is divided based on the number of identified occupants in the image, thereby generating a separate skeleton for each identified occupant.

[0180] In some embodiments, a confidence level is assigned to each estimated key point in space (e.g., the vehicle cabin).

[0181] At step 730, according to an embodiment, the 3D image sequence is analyzed to generate a depth map representation of the occupant. For example, a 3D image of an object captured with structured light illumination includes specific pattern features corresponding to the illumination pattern projected onto the object. The pattern features can be stripes, lines, dots, or other geometric shapes and include uniform or non-uniform characteristics such as shape, size, and intensity. Figure 3A An exemplary captured image illuminated with a specific structured light (e.g., dots) is described in

[0182] At step 740, according to an embodiment, the 3D graphical representation and the skeleton annotation layer of each occupant of each image are combined to obtain, for example, a skeleton model (3D skeleton model) of each occupant. Generally, the generated skeleton model is used to identify the orientation / pose / distance of the occupant in the image obtained from the imaging device. Specifically, the skeleton model includes data such as the 3D key point (x, y, z) representation of the occupant relative to the X-Y-Z coordinate system, where the (x, y) points represent the positions of the body joint surfaces of the occupant in the obtained image, and (z) represents the distance of the relevant (x, y) key point surface from the image sensor.

[0183] For example, Figure 7B According to an embodiment, an image 780 including a combined 3D layer and a skeleton layer of a vehicle interior passenger compartment 782 is shown. Image 780 shows a passenger 785 sitting in a vehicle seat and a plurality of reflected light patterns (e.g., dots) 788 for estimating the distance (e.g., depth) of each relevant body part from the image sensor. The image also includes a skeleton 790 formed by connecting a number of selected key points at the passenger's body by connecting lines.

[0184] It should be emphasized that although Figure 7A steps 730 and 740 include obtaining a reflected light pattern image to obtain depth data for each image, the present invention may include obtaining 3D images and / or extracting depth data by any type of 3D system, device, and method (e.g., stereo cameras and / or ToF sensors known in the art).

[0185] At step 750, the skeleton model is analyzed to extract one or more features of the occupant. In an embodiment, the extracted features may include data such as the measured pose and / or orientation of each occupant in the vehicle. In some embodiments, the features may also include the length of one or more body parts of the occupant (e.g., the main body parts of the occupant, such as shoulders, hips, torso, legs, body, etc.). Advantageously, the generated skeleton model provides the "true length" (e.g., or actual length) of each body part, rather than the "projected length" that can be obtained when only a 2D image of a person is obtained. The analysis based on 3D data improves the accuracy of quality estimation because the "projected length" is very limited in providing quality estimation (e.g., angle-sensitive, etc.). For example, as Figure 7B shown, the analysis includes an obtained image 780 of the reflected light pattern including points and a skeleton superimposed on the image of a passenger sitting in a front or rear vehicle seat to estimate the lengths of human body parts such as shoulders, hips, torso, legs, body, etc.

[0186] At step 760, according to an embodiment, one or more skeletal models of each occupant are analyzed to filter out (e.g., remove or discard) one or more skeletal models based on, for example, predefined filtering criteria, and an effective skeletal model of the occupant (e.g., suitable for mass estimation) is obtained. The predefined filtering criteria include selection rules that define the required postures and orientations for estimating the mass of the occupant. For example, the predefined filtering criteria include selection rules that define "abnormal" or "invalid" postures or orientations of the occupant. An "abnormal" posture or orientation can be defined as a posture or orientation of the occupant in which, due to, for example, an improper sitting posture of the occupant or due to the angle of the occupant relative to the imaging of the image sensor, a complete or nearly complete skeletal representation is not presented or imaged. In some cases, an improper posture can be related to a posture in which the occupant is not sitting upright, such as in a bent position. Thus, the analysis of these "abnormal" skeletal representations is used to discard postures defined as "abnormal" (e.g., inaccurate or incorrectly measured), and thus these skeletons are ignored in the mass estimation process. Non-limiting examples of the filtering criteria include defined spatial relationships between the skeletal features of the identified object and / or the identified abnormal postures. Non-limiting examples of the discarded postures are shown in Figure 5A and 5B as shown.

[0187] In some cases, a pose density model method can be used to filter the analyzed images. According to an embodiment, the pose density model method includes placing each object skeleton configuration in a high-dimensional space and discarding any configuration within a predetermined distance from the high-density regions in this space.

[0188] At step 770, according to an embodiment, the effective skeletal model of the occupant is analyzed to estimate the mass of the occupant. In some embodiments, the analysis process includes inserting the features of the extracted effective skeletal model into a measurement model (e.g., a pre-trained regression model), which is configured to estimate the mass of the occupant at time (t) based on the current and previous (t-i) mass measurement values. In some cases, the estimation model is a machine learning estimation model, which is configured to determine the mass and / or mass classification of the occupant. In some cases, the measurement model is configured to provide a continuous value of the predicted mass, or to perform a coarser estimation and classify the occupant according to mass categories (e.g., child, short adult, normal person, tall adult).

[0189] Alternatively or in combination, the effective skeletal model of the occupant is processed to classify each occupant according to a predetermined mass classification. For example, a passenger weighing approximately 60 kg, e.g., in the range of 50 - 65 kg, will be classified into the "short adult" subcategory, while a child weighing approximately 25 kg, e.g., in the range of 10 - 30 kg, will be classified into the "child" subcategory.

[0190] Figure 7CIt is a schematic flowchart of a method 705 for estimating the mass of one or more occupants in a vehicle according to an embodiment. The method 705 presents all the steps of the aforementioned method 700, but also includes, at step 781, classifying the identified occupants according to one or more measured mass subcategories (e.g., child, small adult, normal person, large adult).

[0191] At step 782, an output, such as an output signal, is generated based on the measured and determined mass or mass classification of each identified occupant. For example, an output signal including an estimated mass and / or mass classification can be transmitted to an airbag control unit (ACU) to determine whether the airbag should be inhibited or deployed, and if deployed, at various output levels.

[0192] According to other embodiments, the output including the mass estimate can control the vehicle's HVAC (heating, ventilation, and air conditioning) system; and / or optimize the vehicle's electronic stability control (ESC) according to the measured mass of each of the vehicle occupants.

[0193] Figure 8 It is a schematic flowchart of a method 800 for measuring the mass of one or more occupants in a vehicle according to another embodiment. For example, the method can include determining, in real time, the mass of one or more occupants sitting on vehicle seats according to one or more mass classification categories, and correspondingly outputting one or more signals to activate one or more vehicle units. Some stages of the method 800 can be performed at least in part by at least one computer processor, such as by the processor 152 and / or the vehicle computing unit. A corresponding computer program product can be provided, which includes a computer-readable storage medium having a computer-readable program included therein and configured to perform the relevant stages of the method 800. In other embodiments, the method includes steps different from or additional to those described in Figure 8 description. Additionally, in various embodiments, the steps of the method can be performed in an order different from the order described in Figure 8 description.

[0194] At step 810, according to an embodiment, a plurality of images including one or more visual images are obtained, where the visual images are, for example, a 2D image sequence and a 3D image sequence of a vehicle cabin. The obtained 2D and 3D image sequences include images of one or more occupants (e.g., a driver and / or a passenger sitting in the rear of the vehicle and / or in the back seat). According to an embodiment, the 3D image can be any type of stereoscopic image, such as an image captured by a stereoscopic camera. Alternatively or in combination, the 3D image can be captured by a ToF image sensor. Alternatively or in combination, the 3D image may include a reflected light pattern. The 2D image can be, for example, a clear visual image that does not include a reflected light pattern. In some embodiments, an image sensor located in the vehicle cabin, such as at the vehicle front section shown in Figure 1A captures a plurality of images (e.g., 2D and 3D images) synchronously and / or sequentially.

[0195] According to an embodiment, the 3D image may include a depth map representation of an occupant. For example, a 3D image of an object captured with structured light illumination may include specific pattern features corresponding to the illumination pattern projected onto the object. The pattern features can be stripes, lines, dots, or other geometric shapes, and include uniform or non-uniform characteristics, such as shape, size, and intensity. Figure 3A Exemplary captured images illuminated with a specific structured light (e.g., dots) are described in

[0196] In some cases, the images are obtained and processed in real time. In some cases, the 2D image and the 3D image can be captured by a single image sensor. In some cases, the 2D image and the 3D image can be captured by different image sensors.

[0197] At step 820, one or more detection algorithms are applied to the obtained 2D image sequence, such as a pose detection and / or a posture detection algorithm to detect the pose and orientation of the occupants in the vehicle cabin. Specifically, the pose detection algorithm is configured to generate a skeleton representation (e.g., a 2D skeleton representation) or a 2D skeleton model of each occupant to identify and / or measure features such as position; orientation; body parts; the length and width of the occupant. For example, the position and orientation of an object can be detected and measured by applying the OpenPose algorithm to the image. Specifically, according to an embodiment, a neural network such as a DNN (Deep Neural Network) is applied to each obtained 2D image over time (t) to generate (e.g., superimpose) a skeleton layer (e.g., a 2D skeleton representation) on each identified occupant. The skeleton layer may include the positions of a plurality of key points that describe the joints of the occupant. In other words, the key points represent in forming as Figure 4A and 4BBody landmarks (e.g., joint body points) detected at the captured body image represented by the skeleton shown. According to an embodiment, each key point of the skeleton representation includes the identified coordinates (x, y) at the occupant body image for extracting features of the identified occupant.

[0198] In some embodiments, the pose estimation method can further be used to identify the occupant and / or the seat of the occupant in each acquired 2D image.

[0199] In some embodiments, the pose estimation method is configured to extract one or more features of the occupant and / or the environment around the occupant, such as the body parts of the occupant and the seat position of the occupant.

[0200] In some embodiments, the identified occupants are separated from each other to obtain separate images of each identified occupant. In some embodiments, each separate image includes the identified occupant and optionally the environment around the occupant, such as the seat of the occupant.

[0201] In some embodiments, each acquired 2D image in the 2D image sequence is divided based on the number of identified occupants in the image, thereby generating a separate skeleton for each identified occupant.

[0202] In some embodiments, a confidence level is assigned to each estimated key point in the space (e.g., vehicle cabin).

[0203] At step 830, according to an embodiment, the 3D image (e.g., depth map) is analyzed to extract one or more distance or depth values regarding the scene or an object in the scene (e.g., occupant) or the distance of the seat of the vehicle from a reference point such as an image sensor. These depth values need to be extracted because objects located far apart from each other in the captured 2D image are wrongly seen as having the same size. Thus, to measure the actual size of the occupants in the vehicle, one or more extracted depth values can be used as a reference scale, such as a scale factor or a normalization factor, to adjust the absolute values of the skeleton model. In some cases, one or more distance values can be extracted by, for example, measuring the average depth value of the features of the occupant (e.g., skeleton values, such as hip, width, shoulder, torso, and / or other body organs) in pixels. In some cases, a single scale factor is extracted. In some cases, scale factors are extracted for each occupant and / or for each acquired image.

[0204] At step 840, the 2D skeleton model is analyzed to extract one or more features of the occupant. In an embodiment, the extracted features can include data such as the measured pose and / or orientation of each occupant in the vehicle. In some embodiments, the features can also include the lengths of one or more body organs of the occupant (e.g., the main body parts of the occupant, such as shoulders, hips, torso, legs, body, etc.).

[0205] At step 850, according to an embodiment, one or more 2D skeleton models of each occupant are analyzed to filter out (e.g., remove or discard) one or more 2D skeleton models based on, for example, one or more extracted features and predefined filtering criteria, thereby obtaining a valid 2D skeleton model of the occupant (e.g., suitable for weight estimation). The predefined filtering criteria include selection rules that define the required postures and orientations for estimating the mass of the occupant. For example, the predefined filtering criteria include selection rules that define "abnormal" postures or orientations of the occupant. An "abnormal" posture or orientation can be defined as a posture or orientation of the occupant in which, due to, for example, an improper sitting posture of the occupant or the angle of the occupant relative to the imaging of the image sensor, a complete or nearly complete skeleton representation is not presented or imaged. In some cases, an improper posture can be related to a posture in which the occupant is not sitting upright, such as in a bent position. Thus, the analysis of these "abnormal" skeleton model representations is used to discard postures defined as "abnormal" (e.g., inaccurate or mismeasured), and thus these skeletons are deleted. Non-limiting examples of filtering criteria include defined spatial relationships between the skeleton features of the identified object and / or the identified abnormal postures. Non-limiting examples of the discarded postures are shown in Figure 5A and 5B are shown.

[0206] In some cases, a pose density model method can be used to filter the analyzed images. According to an embodiment, the pose density model method includes placing each object skeleton configuration in a high-dimensional space and discarding any configuration within a predetermined distance from the high-density regions in this space.

[0207] At step 860, the measured scale factor for each occupant or each image, for example, is applied to the valid 2D skeleton model of the associated occupant accordingly to obtain a scaled 2D skeleton model of the occupant (e.g., a properly scaled 2D skeleton model of the occupant). The scaled 2D skeleton model of the occupant includes information about the distance of the skeleton model from the viewpoint (e.g., the image sensor).

[0208] At step 870, according to an embodiment, the scaled skeleton model of the occupant is analyzed to estimate the mass of the occupant. In some embodiments, the analysis process includes inserting the features of the extracted scaled 2D skeleton model into a measurement model, such as a pre-trained regression model configured to estimate the mass of the occupant. In some cases, the measurement model is a machine learning estimation model configured to determine the mass and / or mass classification of the occupant. In some cases, the measurement model is configured to provide a continuous value of the predicted mass, or to perform a coarser estimation and classify the occupant according to mass categories (e.g., child, short adult, normal person, tall adult).

[0209] Alternatively or in combination, an effective skeletal model of the occupant is processed to classify each occupant according to a predetermined mass classification. For example, a passenger weighing approximately 60 kg, such as in the range of 50 - 65 kg, will be classified into the "small adult" subcategory, while a child weighing approximately 25 kg, such as in the range of 10 - 30 kg, will be classified into the "child" subcategory.

[0210] Figure 9A Figure 901 shows the variation of the mass prediction results (Y-axis) of one or more occupants in a vehicle cabin with the actual measured mass (X-axis) of these occupants, based on the analysis of captured images over time and filtering out invalid images of the occupants in the vehicle, according to an embodiment. Specifically, each captured image is analyzed, and a mass prediction is generated for the identified valid images, while discarding the invalid images. In an embodiment, each point in graph 901 represents a frame captured and analyzed according to the embodiment. As can be clearly shown from the figure, the predicted mass of the occupant is within the range of the actual measured mass of the occupant. For example, based on this method and system, the predicted mass of an occupant with a corresponding predicted mass of 100 Kg is in the range of 80 - 120 kg (and the average value is approximately 100 kg).

[0211] Figure 9B Figure 902 shows the presentation of the mass prediction percentage for mass classification according to an embodiment. For example, the mass prediction for the mass classification of 0 - 35 kg is 100%, 25 - 70 is 95.9%, and 60+ is 94%.

[0212] Figure 9C Another example of Figure 903 shows the variation of the mass prediction results (Y-axis) of one or more occupants sitting in a vehicle cabin with the actual measured mass (X-axis) of these occupants, based on the analysis of captured images of the occupants in the vehicle, such as images 910, 920, and 930, according to an embodiment. As Figure 9C shown, in some cases, some invalid images, such as image 910, are not filtered out, and thus affect the accuracy of the mass prediction. Generally, invalid images, such as 2D image 910, will be automatically filtered out, for example, in real time, to analyze only the images including the standard positions of the occupants (e.g., valid images), and thus obtain an accurate mass prediction.

[0213] In some cases, the identification of non-standard positions of the occupant (e.g., the position shown in image 910) can be used to activate or deactivate one or more vehicle units, such as airbags. For example, the identification of the occupant bending or moving their head away from the road based on the pose estimation model as described herein can be reported to the vehicle's computer and / or processor, and accordingly, the vehicle airbag or hazard warning device can be activated.

[0214] It should be understood that embodiments of the present invention may include mass estimation and / or mass determination of an occupant in a vehicle. For example, the system and method can provide a quick and accurate estimation of the occupant.

[0215] In additional embodiments, the processing unit may be a digital processing device that includes one or more hardware central processing units (CPUs) that perform the functions of the device. In other additional embodiments, the digital processing device further includes an operating system configured to execute executable instructions. In some embodiments, the digital processing device is optionally connected to a computer network. In additional embodiments, the digital processing device is optionally connected to the Internet such that the digital processing device can access the World Wide Web. In yet additional embodiments, the digital processing device is optionally connected to a cloud computing infrastructure. In other embodiments, the digital processing device is optionally connected to an intranet. In other embodiments, the digital processing device is optionally connected to a data storage device.

[0216] According to the description herein, by way of non-limiting example, suitable digital processing devices include server computers, desktop computers, laptop computers, notebook computers, subnotebook computers, netbook computers, netbook tablet computers, set-top box computers, handheld computers, Internet appliances, mobile smart phones, tablet computers, personal digital assistants, video game consoles, and vehicles. Those skilled in the art will recognize that many smart phones are suitable for use in the systems described herein. Those skilled in the art will also recognize that select televisions with optional computer connectivity are suitable for use in the systems described herein. Suitable tablet computers include tablet computers with booklet, tablet, and convertible configurations known to those skilled in the art.

[0217] In some embodiments, the digital processing device includes an operating system configured to execute executable instructions. For example, an operating system is software that includes programs and data, manages the hardware of the device, and provides services for the execution of application programs. Those skilled in the art will recognize that, by way of non-limiting example, suitable server operating systems include FreeBSD, OpenBSD, Linux, Mac OS X Windows and Those skilled in the art will recognize that, by way of non-limiting example, suitable personal computer operating systems include Mac and, for example, UNIX-like operating systems such as. In some embodiments, the operating system is provided by cloud computing. Those skilled in the art will also recognize that, by way of non-limiting example, suitable mobile smart phone operating systems include OS, Research In BlackBerry Windows OS, Windows OS, and

[0218] In some embodiments, the device includes a storage device and / or a memory device. The storage and / or memory device is one or more physical devices for temporarily or permanently storing data or programs. In some embodiments, the device is a volatile memory and requires power to maintain the stored information. In some embodiments, the device is a non-volatile memory and retains the stored information when the digital processing device is not powered. In additional embodiments, the non-volatile memory includes flash memory. In some embodiments, the non-volatile memory includes dynamic random access memory (DRAM). In some embodiments, the non-volatile memory includes ferroelectric random access memory (FRAM). In some embodiments, the non-volatile memory includes phase change random access memory (PRAM). In other embodiments, the device is a storage device, including, by way of non-limiting example, CD-ROM, DVD, flash memory device, disk drive, tape drive, optical disk drive, and cloud computing-based storage devices. In additional embodiments, the storage device and / or memory device is a combination of devices such as those disclosed herein.

[0219] In some embodiments, the digital processing device includes a display for sending visual information to the user. In some embodiments, the display is a cathode ray tube (CRT). In some embodiments, the display is a liquid crystal display (LCD). In additional embodiments, the display is a thin film transistor liquid crystal display (TFT-LCD). In some embodiments, the display is an organic light emitting diode (OLED) display. In various additional embodiments, the OLED display is a passive matrix OLED (PMOLED) or an active matrix OLED (AMOLED) display. In some embodiments, the display is a plasma display. In other embodiments, the display is a video projector. In yet additional embodiments, the display is a combination of devices such as those disclosed herein.

[0220] In some embodiments, a digital processing device includes an input device that receives information from a user. In some embodiments, the input device is a keyboard. In some embodiments, the input device is a pointing device, and as non-limiting examples, the pointing device includes a mouse, a trackball, a touchpad, a joystick, a game controller, or a stylus. In some embodiments, the input device is a touchscreen or a multi-touch screen. In other embodiments, the input device is a microphone for capturing voice or other sound input. In other embodiments, the input device is a camera that captures motion or visual input. In yet other embodiments, the input device is a combination of devices such as those disclosed herein.

[0221] In some embodiments, the systems disclosed herein include one or more non-transitory computer-readable storage media encoded with a program, the program including instructions executable by an operating system of an optionally networked digital processing device. In additional embodiments, the computer-readable storage media are tangible components of the digital processing device. In yet additional embodiments, the computer-readable storage media may optionally be removable from the digital processing device.

[0222] In some embodiments, as non-limiting examples, the computer-readable storage media include CD-ROMs, DVDs, flash memory devices, solid state memories, disk drives, tape drives, optical disc drives, cloud computing systems and services, and the like. In some cases, the program and instructions are permanently, substantially permanently, semi-permanently, or non-transitorily encoded on the media. In some embodiments, the systems disclosed herein include at least one computer program or its use. A computer program includes a sequence of instructions that are executable in a CPU of a digital processing device and are written to perform a particular task. The computer-readable instructions may be implemented as program modules, such as functions, objects, application programming interfaces (APIs), data structures, and the like, that perform particular tasks or implement particular abstract data types. Based on the disclosure provided herein, those skilled in the art will recognize that computer programs may be written in various versions of various languages.

[0223] The functionality of the computer-readable instructions can be combined or distributed in various environments as needed. In some embodiments, a computer program includes a sequence of instructions. In some embodiments, a computer program includes multiple sequences of instructions. In some embodiments, a computer program is provided from one location. In other embodiments, a computer program is provided from multiple locations. In various embodiments, a computer program includes one or more software modules. In various embodiments, a computer program includes, in whole or in part, one or more web applications, one or more mobile applications, one or more stand-alone applications, one or more web browser plugins, extensions, add-ons, or plug-ins, or combinations thereof. In some embodiments, a computer program includes a mobile application provided to a mobile digital processing device. In some embodiments, the mobile application is provided to the mobile digital processing device at the time of manufacture. In other embodiments, the mobile application is provided to the mobile digital processing device via the computer network described herein.

[0224] In view of the disclosure provided herein, mobile applications are generated using hardware, languages, and development environments known in the art by techniques known to those skilled in the art. Those skilled in the art will recognize that mobile applications are written in a variety of languages. As non-limiting examples, suitable programming languages include C, C++, C#, Objective-C, Java TM , Javascript, Pascal, Object Pascal, Python TM , Ruby, VB.NET, WML, and XHTML / HTML with or without CSS, or combinations thereof.

[0225] Suitable mobile application development environments can be obtained from several sources. As non-limiting examples, commercially available development environments include AirplaySDK, alcheMo, Celsius, Bedrock, FlashLite,., NET Compact Framework, Rhomobile, and WorkLight mobile platforms. As non-limiting examples, other development environments are available for free, including but not limited to Lazarus, MobiFlex, MoSyn, and Phonegap. In addition, as non-limiting examples, mobile device manufacturers distribute software development kits that include, but are not limited to, iPhone and iPad (iOS) SDK, Android TM SDK, SDK, BREW SDK, OS SDK, Symbian SDK, webOS SDK, and Mobile SDK。

[0226] Those skilled in the art will recognize that, as non-limiting examples, several commercial forums can be used to distribute mobile applications that include app stores, Android TM Market, App World, app stores for handheld devices, app catalogs for webOS, for mobile markets for featured services for devices, applications, and DSi Shop.

[0227] In some embodiments, the systems disclosed herein include software, server, and / or database modules, or their uses. Given the disclosure provided herein, software modules are created by techniques known to those skilled in the art, using machines, software, and languages known to those skilled in the art. The software modules disclosed herein are implemented in a variety of ways. In various embodiments, software modules include files, code segments, programming objects, programming constructs, or combinations thereof. In additional various embodiments, software modules include multiple files, multiple code segments, multiple programming objects, multiple programming constructs, or combinations thereof. In various embodiments, as non-limiting examples, one or more software modules include web applications, mobile applications, and stand-alone applications. In some embodiments, software modules are in one computer program or application. In other embodiments, software modules are in more than one computer program or application. In some embodiments, software modules are hosted on one machine. In other embodiments, software modules are hosted on more than one machine. In additional embodiments, software modules are hosted on a cloud computing platform. In some embodiments, software modules are hosted on one or more machines in one location. In other embodiments, software modules are hosted on one or more machines in more than one location.

[0228] In some embodiments, the systems disclosed herein include one or more databases or the uses of such databases. Given the disclosure provided herein, those skilled in the art will recognize that many databases are suitable for storing and retrieving information as described herein. In various embodiments, as non-limiting examples, suitable databases include relational databases, non-relational databases, object-oriented databases, object databases, entity-relationship model databases, associative databases, and XML databases. In some embodiments, databases are Internet-based. In additional embodiments, databases are network-based. In yet additional embodiments, databases are cloud computing-based. In other embodiments, databases are based on one or more local computer storage devices.

[0229] In the above description, an embodiment is an example or implementation of the present invention. Various occurrences of "an embodiment", "embodiments", or "some embodiments" do not necessarily refer to the same embodiment.

[0230] Although various features of the present invention may be described in the context of a single embodiment, the features may also be provided individually or in any suitable combination. Conversely, although the present invention may be described herein in the context of separate embodiments for clarity, the present invention may also be implemented in a single embodiment.

[0231] References in the specification to "some embodiments", "embodiments", "an embodiment", or "other embodiments" mean that the particular features, structures, or characteristics described in connection with the embodiments are included in at least some embodiments of the present invention, but not necessarily in all embodiments.

[0232] It should be understood that the language and terminology used herein should not be construed as limiting and are for descriptive purposes only.

[0233] The principles and uses of the teachings of the present invention can be better understood with reference to the accompanying description, drawings, and examples.

[0234] It should be understood that the details set forth herein are not to be construed as limitations on the application of the present invention.

[0235] Furthermore, it should be understood that the present invention can be implemented or practiced in various ways and can be implemented in embodiments other than those outlined in the above description.

[0236] It should be understood that the terms "comprising", "including", "consisting", and their grammatical variants do not exclude the addition of one or more components, features, steps, or integers or groups thereof, and these terms should be construed as designating components, features, steps, or integers.

[0237] If the specification or claims refer to "additional" elements, the presence of more than one additional element is not excluded.

[0238] It should be understood that where the claims or specification refer to "a" or "an" element, such reference should not be construed as meaning only one of such element. It should be understood that where the specification states that a component, feature, structure, or characteristic "may", "might", or "could" be included, the inclusion of a particular component, feature, structure, or characteristic is not required. Where applicable, although state diagrams, flowcharts, or both may be used to describe embodiments, the present invention is not limited to those illustrations or corresponding descriptions. For example, a process need not move through each illustrated box or state, or move in exactly the same order as illustrated and described. The methods of the present invention can be implemented by performing or completing selected steps or tasks manually, automatically, or a combination thereof.

[0239] The descriptions, examples, methods, and materials presented in the claims and the specification should not be construed as limiting, but rather as illustrative only. Unless otherwise defined, the meanings of technical and scientific terms used herein are those commonly understood by one of ordinary skill in the art to which this invention pertains. The present invention may be practiced or tested using methods and materials equivalent or similar to those described herein.

[0240] Although the present invention has been described with reference to a limited number of embodiments, these should not be construed as limiting the scope of the invention, but rather as exemplifications of some preferred embodiments. Other possible variations, modifications, and applications are also within the scope of the present invention. Accordingly, the scope of the present invention should not be limited by what has been described so far, but rather by the appended claims and their legal equivalents.

[0241] All publications, patents, and patent applications mentioned in this specification are hereby incorporated by reference in their entirety into this specification, and similarly, each individual publication, patent, or patent application is specifically and individually indicated to be incorporated by reference herein. Additionally, the citation or identification of any reference document in this application should not be construed as an admission that such reference document is available as prior art to the present invention. With respect to the use of section headings, the section headings should not be construed as necessarily limiting.

Claims

1. A method for estimating the mass of one or more occupants in a vehicle cabin, the method comprising: Providing a processor configured to: Obtain a plurality of images of the one or more occupants, wherein the plurality of images includes a 2D (two-dimensional) image sequence and a 3D (three-dimensional) image sequence of the vehicle cabin captured by an image sensor; Applying a pose detection algorithm to each of the obtained 2D image sequences to obtain one or more skeletal representations of the one or more occupants; Combining one or more 3D images of the 3D image sequence with the one or more skeletal representations of the one or more occupants to obtain at least one skeletal model for each of the one or more occupants, wherein the skeletal model includes information about the distances of one or more key points of the skeletal model from a viewpoint; Analyzing the one or more skeletal models to extract one or more features of each of the one or more occupants; Processing the one or more extracted features of the skeletal model to estimate the mass of each of the one or more occupants; Wherein the processor is configured to and capable of filtering out one or more skeletal models based on predefined filtering criteria to obtain valid skeletal models, and wherein the predefined filtering criteria includes specific selection rules defining valid poses or orientations of the one or more occupants, wherein the predefined filtering criteria includes defined spatial relationships between skeletal features, and wherein the defined spatial relationships between skeletal features includes defined spatial relationships between occupant body parts.

2. The method according to claim 1, wherein the predefined filtering criteria is based on the measured confidence levels of one or more key points in the 2D skeletal representation.

3. The method according to claim 2, wherein the confidence level is based on the measured probability heatmap of the one or more key points.

4. The method according to claim 1, wherein the predefined filtering criteria is based on a high-density model.

5. The method according to claim 1, wherein the processor is configured to generate one or more output signals, the one or more output signals including the estimated mass of each of the one or more occupants.

6. The method according to claim 5, wherein the output signal is associated with the operation of one or more of the units of the vehicle.

7. The method according to claim 6, wherein the units of the vehicle are selected from: Airbag; Electronic Stability Control (ESC) unit; Seatbelt.

8. The method according to claim 1, wherein the 2D image sequence is a visual image of the cabin.

9. The method according to claim 1, wherein the 3D image sequence is one or more of the following: Reflected light pattern image; Stereo image.

10. The method according to claim 1, wherein the image sensor is selected from: Time-of-Flight (ToF) camera; Stereo camera.

11. The method according to claim 1, wherein the pose detection algorithm is configured to identify the pose or orientation of the one or more occupants in the obtained 2D images.

12. The method according to claim 1, wherein the pose detection algorithm is configured to: Identify a plurality of key points of the one or more occupant body parts in at least one 2D image of the 2D image sequence; Link pairs of the detected plurality of key points to generate a skeletal representation of the occupant in the 2D image.

13. The method according to claim 12, wherein the key points are joints of the occupant's body.

14. The method according to claim 1, wherein the pose detection algorithm is the OpenPose algorithm.

15. The method according to claim 1, wherein the one or more extracted features are of the occupant: One or more of shoulder length; torso length; knee length; pelvic position; hip width.

16. A method for estimating the mass of one or more occupants in a vehicle cabin, the method comprising: Providing a processor configured to: Obtain a plurality of images of the one or more occupants, wherein the plurality of images includes a 2D (two-dimensional) image sequence and a 3D (three-dimensional) image sequence of the vehicle cabin captured by an image sensor; Apply a pose detection algorithm to each of the obtained 2D image sequences to obtain one or more skeletal representations of the one or more occupants; Analyze one or more 3D images of the 3D image sequence to extract one or more depth values of the one or more occupants; Apply the extracted depth values to the skeletal representation accordingly to obtain a scaled skeletal representation of the one or more occupants, wherein the scaled skeletal model includes information about the distance of the skeletal model from the viewpoint; Analyze the scaled skeletal representation to extract one or more features of each of the one or more occupants; Process the one or more extracted features to estimate the mass or body mass classification of each of the one or more occupants; Wherein the processor is configured to filter out one or more skeletal representations based on predefined filtering criteria to obtain a valid skeletal representation, and wherein the predefined filtering criteria includes specific selection rules defining a valid pose or orientation of the one or more occupants, wherein the predefined filtering criteria includes defined spatial relationships between skeletal features, and wherein the defined spatial relationships between skeletal features include defined spatial relationships between occupant body parts.

17. The method according to claim 16, wherein the predefined filtering criteria is based on a measured confidence level of one or more key points in the 2D skeletal representation.

18. The method according to claim 17, wherein the confidence level is based on a measured probability heat map of the one or more key points.

19. The method according to claim 16, wherein the predefined filtering criteria is based on a high-density model.

20. The method according to claim 16, wherein the processor is configured to generate one or more output signals, the one or more output signals including the estimated mass or body mass classification of each of the one or more occupants.

21. The method according to claim 20, wherein the output signal corresponds to the operation of one or more of the units of the vehicle.

22. A system for estimating the mass of one or more occupants in a vehicle cabin, the system comprising: a sensing device, the sensing device comprising: a lighting module, the lighting module comprising one or more lighting sources configured to illuminate the vehicle cabin; at least one imaging sensor configured to capture a sequence of 2D (two-dimensional) images and a sequence of 3D (three-dimensional) images of the vehicle cabin; and at least one processor configured to: apply a pose detection algorithm to each of the acquired 2D image sequences to obtain one or more skeletal representations of the one or more occupants; combine one or more 3D images of the 3D image sequence with the one or more skeletal representations of the one or more occupants to obtain at least one skeletal model for each of the one or more occupants, wherein the skeletal model includes information about the distances of one or more key points in the skeletal model from the viewpoint; analyze the one or more skeletal models to extract one or more features of each of the one or more occupants; and process the one or more extracted features of the skeletal model to estimate the mass of each of the one or more occupants; wherein the processor is configured to filter out one or more skeletal models based on predefined filtering criteria to obtain valid skeletal models, and wherein the predefined filtering criteria include specific selection rules defining valid poses or orientations of the one or more occupants, wherein the predefined filtering criteria include defined spatial relationships between skeletal features, and wherein the defined spatial relationships between skeletal features include defined spatial relationships between occupant body parts.

23. The system according to claim 22, wherein the predefined filtering criteria are based on the measured confidence levels of one or more key points in the 2D skeletal representation.

24. The system according to claim 23, wherein the confidence level is based on the measured probability heat map of the one or more key points.

25. The system according to claim 22, wherein the predefined filtering criteria are based on a high-density model.

26. The system according to claim 22, wherein the sensing device is selected from: a ToF sensing device; a stereoscopic sensing device.

27. The system according to claim 22, wherein the sensing device is a structured light pattern sensing device, and the at least one lighting source is configured to project modulated light onto the vehicle cabin in a predefined structured light pattern.

28. The system according to claim 27, wherein the predefined structured light pattern is composed of a plurality of diffused light elements.

29. The system according to claim 28, wherein the shape of the light element is one or more of the following: a point; a line; a stripe; or a combination thereof.

30. The system according to claim 22, wherein the processor is configured to generate one or more output signals, the one or more output signals including the estimated mass or body mass classification of each of the one or more occupants.

31. The system according to claim 30, wherein the output signal corresponds to the operation of one or more of the units of the vehicle.

32. The system according to claim 31, wherein the units of the vehicle are selected from: Airbag; Electronic Stability Control (ESC) unit; Seat belt.

33. A non-transitory computer-readable storage medium storing computer program instructions that, when executed by a computer processor, cause the processor to perform the following steps: Obtain a sequence of 2D (two-dimensional) images and a sequence of 3D (three-dimensional) images of one or more occupants, wherein the 3D images have a plurality of pattern features according to an illumination pattern; Apply a pose detection algorithm to each of the obtained 2D image sequences to obtain one or more skeletal representations of the one or more occupants; Combine one or more 3D images of the 3D image sequence with the one or more skeletal representations of the one or more occupants to obtain at least one skeletal model for each of the one or more occupants, wherein the skeletal model includes information about the distances of one or more key points in the skeletal model from the viewpoint; Analyze the one or more skeletal models to extract one or more features of each of the one or more occupants; Process the one or more extracted features of the skeletal model to estimate the mass or body mass classification of each of the one or more occupants; Wherein the processor is configured to filter out one or more skeletal models based on predefined filtering criteria to obtain valid skeletal models, and wherein the predefined filtering criteria include specific selection rules defining valid poses or orientations of the one or more occupants, wherein the predefined filtering criteria include defined spatial relationships between skeletal features, and wherein the defined spatial relationships between skeletal features include defined spatial relationships between occupant body parts.

Citation Information

Patent Citations

  • Optical weight sensor for vehicular safety restraint systems

    US5988676A

  • Video surveillance systems, devices and methods with improved 3D human pose and shape modeling

    US20130250050A1