Intelligent cabin monitoring system based on light time-of-flight camera
The intelligent cockpit monitoring system based on optical time-of-flight cameras solves the problems of light sensitivity, low accuracy, and high privacy risks in existing cockpit monitoring solutions. It achieves high-precision, non-contact, multi-dimensional monitoring and functional linkage, reducing the R&D and mass production costs for automakers.
Patent Information
- Application Number
- CN202511713041.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-20
- Publication Date
- 2026-02-17
AI Technical Summary
Existing cockpit monitoring solutions are inadequate in terms of light sensitivity, low accuracy, high privacy risks, and fragmented functions, making it difficult to meet the high-level requirements of smart cockpits.
The intelligent cockpit monitoring system, based on a time-of-flight camera, includes a ToF camera module, a main control processing unit, a cockpit control unit, and a local encrypted storage unit. It achieves 3D point cloud acquisition, preprocessing, analysis, and control command transmission through MIPI CSI-2 and SPI interfaces. It integrates identity recognition, gesture interaction, health monitoring, and safety early warning functions, and uses Gaussian filtering, RANSAC algorithm, and 3D face recognition model for accurate monitoring.
It achieves high-precision multi-dimensional monitoring in complex environments, reduces the risk of privacy leaks, lowers the R&D and mass production costs for automakers, and forms an intelligent closed loop of perception-decision-execution.
Smart Images

Figure CN121545138A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of intelligent cockpit, and particularly to an intelligent cockpit monitoring system based on a light time-of-flight camera. BACKGROUND
[0002] With the continuous improvement of the intelligent level of automobiles, intelligent cockpits have gradually moved away from the traditional function superposition mode and evolved towards active perception and adaptive interaction. In this development process, accurate and real-time perception of the state of the occupants in the cockpit is the core foundation for realizing the high-level functions of intelligent cockpits and directly affects the interactive experience and driving safety. At present, the mainstream cockpit monitoring solutions rely on 2D RGB cameras or near-infrared cameras to build, but such solutions are limited by the technical principles and have many inherent defects, making it difficult to meet the advanced needs of intelligent cockpits.
[0003] In terms of environmental adaptability, a 2D vision system passively receives ambient light for imaging. In strong light scenes such as midday sunlight, the light intensity can reach 100,000 lux, and image overexposure and backlight shadows are easy to occur, resulting in blurred facial features and gesture outlines. In weak light scenes such as night without light sources, the light intensity is only 0.1 lux, and the image will produce a large number of noise points, and in severe cases, the target cannot be recognized. Although some solutions attempt to optimize through light supplementation, the optimization effect is limited when the light supplementation intensity is insufficient, and too high light supplementation intensity will cause visual interference to the occupants, and it is always impossible to achieve truly all-weather stable monitoring. In terms of monitoring accuracy, a 2D image can only reflect the plane coordinates and cannot distinguish the positional relationship of objects in three-dimensional space, which is easy to cause misjudgment. For example, it will misjudge the action of the hand close to the central control screen as a touch operation, and misjudge the posture of the head of the rear passenger as not wearing a seat belt. At the same time, the 2D vision system cannot accurately capture the millimeter-level micro-motions of the chest due to breathing and heartbeat, resulting in a life sign monitoring error rate of more than 5 times / minute, which is much higher than the standard of an error of less than or equal to 2 times / minute required for practicality. In terms of privacy protection, a 2D RGB image contains sensitive information such as the facial details of the occupants, the wearing features, and the in-vehicle environment. Even if a local storage method is used, these data still face the risk of being stolen by hackers or uploaded by the vehicle manufacturer in violation of regulations. Some solutions protect privacy through blurring, but this method will further reduce the recognition accuracy, forming a contradiction between privacy protection and monitoring accuracy.
[0004] In addition, the existing system also has the problem of functional fragmentation. The system does not design a standardized processing flow for three-dimensional point clouds, and often splits identity recognition, gesture control and health monitoring into independent modules. Different sensors are used in each module, which leads to the problem that data cannot be interchanged between modules. At the same time, the system can only achieve a single response of a single function, for example, only adjusting the seat after recognizing the driver, and cannot link to adjust the air conditioner and the sound; only issuing a sound alarm when detecting fatigue driving, without forming a complete closed loop of perception-decision-execution; and the hardware of the scheme has poor compatibility with the vehicle system, and the data interface does not comply with the vehicle protocol, which causes the vehicle manufacturer to face the problems of high production cost and long cycle in the integration process.
[0005] In summary, the existing cockpit monitoring scheme cannot balance environmental adaptability, monitoring accuracy, privacy protection and functional linkage, and therefore a new monitoring scheme is urgently needed to break through the above bottlenecks. SUMMARY
[0006] The purpose of the present application is to provide an intelligent cockpit monitoring system based on a light time-of-flight camera, which solves the problems of light sensitivity, low accuracy, high privacy risk and functional fragmentation of the existing 2D vision system, and can realize high-precision, non-contact cockpit multi-dimensional monitoring.
[0007] To achieve the above purpose, the present application provides an intelligent cockpit monitoring system based on a light time-of-flight camera, which comprises a ToF camera module, a main control processing unit, a cockpit control unit and a local encrypted storage unit. The ToF camera module is electrically connected to the main control processing unit through a MIPI CSI-2 interface and is used to collect three-dimensional point clouds of the entire cockpit. The main control processing unit is used to receive three-dimensional point clouds, and to issue control instructions and warning signals after preprocessing and analyzing the three-dimensional point clouds. The cockpit control unit is electrically connected to the main control processing unit through a vehicle Ethernet, and is used to receive control instructions and warning signals and drive the actuators in the cockpit to act. The local encrypted storage unit is used to store the feature templates of users, and is connected to the main control processing unit through a SPI interface to store the point clouds preprocessed by the main control processing unit.
[0008] Preferably, the ToF camera module comprises an infrared light source, an optical assembly, a ToF sensor and a module controller. The infrared light source is used to emit modulated infrared light. The optical assembly comprises an aspherical condenser lens and a narrow-band optical filter. The aspherical condenser lens is used to focus the infrared light rays. The bandwidth of the narrow-band optical filter is for filtering out the interference of ambient light; The ToF sensor is configured to receive the infrared light signal reflected by the cabin, convert the light signal into an electrical signal, and generate a three-dimensional point cloud containing distance information based on a phase difference method.
[0009] Preferably, the module controller further transmits the standardized data to the host processing unit through a MIPI CSI-2 interface, and receives a parameter configuration instruction sent by the host processing unit through an I2C interface.
[0010] Preferably, the host processing unit comprises a data preprocessing module, a core analysis engine, and a decision output module. The data processing module is configured to preprocess the collected three-dimensional point cloud. The core analysis engine is configured to receive the preprocessed three-dimensional point cloud and perform multi-dimensional monitoring analysis on the derived point cloud. The decision output module is configured to generate a control instruction and a warning signal based on the analysis result of the core analysis engine.
[0011] Preferably, the data processing module pre-processes the collected three-dimensional point cloud in the following manner: Gaussian filtering is performed on the three-dimensional point cloud to remove salt and pepper noise in the three-dimensional point cloud by calculating the weighted average value of the neighborhood pixels; A static plane in the cabin is fitted based on the RANSAC algorithm, the static plane includes a seat backrest and an instrument panel, a plane fitting error is set, and the point cloud with a depth value exceeding the threshold range is determined as a dynamic occupant point cloud; The occupant point cloud is converted from a camera coordinate system to a cabin coordinate system, and the deviation caused by the installation angle of the ToF camera module is eliminated through point cloud registration; The derived point cloud is filtered according to the type of the monitoring target, and the derived point cloud includes a head point cloud, a chest point cloud, and a hand point cloud.
[0012] Preferably, the core analysis engine performs multi-dimensional monitoring analysis on the derived point cloud through an identity recognition module, a posture and behavior analysis module, a vital sign monitoring module, and an attention state analysis module. The identity recognition sub-module is configured to identify the occupant, and the specific identification process is as follows: First, the voxel grid downsampling algorithm is used to reduce the resolution of the head point cloud, and then the downsampled head point cloud is input into the MobileFaceNet-3D model to extract the feature vector of the head point cloud through NPU inference. Then, the feature vector is matched with the user feature template in the local encrypted storage unit in the Euclidean distance: If the matching value of the user feature template is less than or equal to If the matching value of the user feature template is less than or equal to Then, it is determined that the user identity recognition is successful, and the user ID is recorded. If the matching value of the user feature template is less than or equal to The frame's matching value is greater than If so, it is determined that the user's identity has not been identified.
[0013] Preferably, the posture behavior analysis module is used to analyze occupant posture behavior, and the specific analysis process is as follows: First, the 3D space is divided into several voxel meshes to form voxel feature maps; then, the voxel feature maps are input into the HRNet-3D model and output. The three-dimensional coordinates of each joint point cover the head, neck, shoulder, elbow, wrist, fingertips, torso, hip, knee, and ankle. Next, gesture recognition is performed by calculating the displacement trajectory and angle changes of the wrist and fingertip joints. Then, by calculating the distance between key points of the torso and the seat back plane, it is determined whether the occupant is leaning excessively forward. Key torso points include the base of the neck, sternum, and lower back. When the distance between the two is less than... And the duration exceeds If so, it is judged that the body is leaning forward excessively; By calculating the distance between the shoulder acromion and the seatbelt buckle point, it can be determined whether the occupant is not wearing a seatbelt. The shoulder acromion is the key point on the shoulder. When the distance between the two is greater than [missing information], it indicates that the occupant is not wearing a seatbelt. At that time, it was determined that the occupant was not wearing a seat belt.
[0014] Preferably, the vital signs monitoring module is used for monitoring the vital signs of occupants, and the specific monitoring process is as follows: First, the point cloud of the thoracic cavity region is located. Then, the average depth value of the thoracic cavity region is collected to generate a data structure containing... The time series signal of each data point was used to extract respiratory and heartbeat signals through bandpass filtering. Then, the peak detection algorithm is used to calculate the respiratory rate and heart rate. The user defines the abnormal threshold for vital signs. When the respiratory rate and heart rate exceed the threshold range, it is judged as abnormal and an early warning is issued.
[0015] Preferably, the attention state analysis module is used for driver attention assessment, and the specific process is as follows: First, calculate the head orientation vector based on the coordinates of key head points, including the top of the head and the chin. Then, calculate the angle between the head orientation vector and the vehicle's direction of travel. Simultaneously, the PERCLOS value was calculated by analyzing the density changes of the point cloud in the eye region. The PERCLOS value represents the percentage of eyelid closure time per minute. When the included angle And the duration exceeds Or when the PERCLOS value is greater than At that time, it was determined to be a case of inattention; When the PERCLOS value is greater than At that time, it was determined to be fatigued driving, and a graded warning was issued.
[0016] Preferably, the decision output module generates control commands and early warning signals based on the analysis results obtained from the core analysis engine. The specific process is as follows: When identity recognition is successful, the decision output module sends a parameter configuration command to the cockpit control unit. The parameter configuration command includes the three-dimensional coordinates of the seat, the horizontal and vertical angles of the rearview mirror, and the air conditioning temperature. When a valid gesture is recognized, the decision output module generates the actuator control command corresponding to the configuration command; when a dangerous behavior is detected, it sends commands to vibrate the seat and flash the red icon on the dashboard. When respiratory rate and heart rate are detected to be abnormal vital signs that exceed the threshold range, the decision output module displays warning text through the HUD and plays voice prompts through the speaker. When abnormal attention is detected, a tiered warning system is activated: Level 1 warning is seat vibration on one side, Level 2 warning is seat vibration plus voice prompt, and Level 3 warning is seat vibration, voice prompt, red flashing on the instrument panel and warning icon displayed on the HUD. When the head is facing the angle between the vector and the vehicle's direction of travel And the duration is equal to Or when the PERCLOS value is less than or equal to The warning was lifted at that time.
[0017] Therefore, the present invention employs the above-mentioned intelligent cockpit monitoring system based on an optical time-of-flight camera, and the beneficial effects are as follows: (1) The three-dimensional point cloud based on ToF technology can actively resist ambient light interference. Under automotive-grade temperature range and wide illumination conditions, the recognition accuracy of each function decreases less. It can solve the problem of strong light overexposure and weak light failure of two-dimensional system. It can also use the reflection intensity to help distinguish objects in the cabin, further improving the recognition stability in complex environments.
[0018] (2) The head, chest cavity and hand derived point clouds obtained by the present invention through preprocessing can be processed as needed, which can balance recognition accuracy and computing power consumption to meet the real-time requirements of the vehicle system. At the same time, the original point cloud will be quickly destroyed, only the encrypted feature vector is retained and the details cannot be restored. It also supports users to actively clear the template, and eliminates privacy leakage throughout the entire chain.
[0019] (3) The present invention integrates functions such as identity recognition, gesture interaction, health monitoring, and security warning in a single system without redundant sensors. It forms an intelligent closed loop of perception, decision-making and execution through a standard bus. The algorithm is lightweight and adapted to the vehicle's computing power, which can reduce the cost and cycle of R&D and mass production for car companies.
[0020] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Attached Figure Description
[0021] Figure 1 This is an overall system block diagram of an embodiment of an intelligent cockpit monitoring system based on an optical time-of-flight camera according to the present invention; Figure 2 This is a block diagram of the main control processing unit structure of an embodiment of an intelligent cockpit monitoring system based on an optical time-of-flight camera according to the present invention; Figure 3 This is a schematic diagram of temperature adaptability testing of an embodiment of an intelligent cockpit monitoring system based on an optical time-of-flight camera according to the present invention. Figure 4 This is a schematic diagram illustrating the time consumption percentage of each stage in the entire process of an embodiment of an intelligent cockpit monitoring system based on an optical time-of-flight camera according to the present invention. Detailed Implementation
[0022] The technical solution of the present invention will be further described below with reference to the accompanying drawings and embodiments.
[0023] Unless otherwise defined, the technical or scientific terms used in this invention shall have the ordinary meaning as understood by one of ordinary skill in the art to which this invention pertains.
[0024] This invention employs a four-layer architecture consisting of a perception layer, a processing layer, an execution layer, and a security layer. The interaction of point cloud data and instructions is achieved between these layers through standardized interfaces and protocols. For example... Figure 1 As shown, an intelligent cockpit monitoring system based on a Time-of-Flight (ToF) camera includes a ToF camera module, a main control processing unit, a cockpit control unit, and a local encrypted storage unit. The ToF camera module is electrically connected to the main control processing unit via a MIPI CSI-2 interface and can be installed above the rearview mirror or in the center of the roof. Its field of view covers the driver, front passenger, and the left or right rear seats, enabling real-time acquisition of 3D point clouds across the entire cockpit. The output resolution is [resolution missing]. Depth and intensity images with adjustable frame rates from 30 to 60 fps.
[0025] The main control processing unit receives 3D point clouds, preprocesses and analyzes them, and then issues control commands and warning signals. The cockpit control unit is electrically connected to the main control processing unit via an in-vehicle Ethernet network. It receives control commands and warning signals and drives actuators within the cockpit, such as seat motors, air conditioning controllers, HUDs, or seat vibration modules. A local encrypted storage unit stores the user's feature templates and connects to the main control processing unit via an SPI interface to store the preprocessed point cloud. This invention's local encrypted storage unit uses an AES-256 encryption chip, storing only anonymized 3D facial feature templates and not retaining the original depth image. It supports user-triggered data clearing commands.
[0026] The ToF camera module includes an infrared light source, optical components, a ToF sensor, and a module controller. The infrared light source employs an 8-channel VCSEL array laser to emit near-infrared light with a modulation wavelength of 940nm and a modulation frequency of 20MHz. Its output power ranges from 5 to 15mW and can dynamically adapt to the reflectivity of the cabin's interior materials to avoid overexposure under direct strong light and prevent insufficient signal in low-light environments.
[0027] The optical components include a focal length of 3.5mm and a field of view of... Aspherical condenser lens and bandwidth of A narrowband filter. The aspherical focusing lens is used to focus infrared light, while the narrowband filter only allows light with a wavelength of 940nm to pass through, thus filtering out interference from ambient light such as sunlight and interior lights.
[0028] The ToF sensor receives infrared light signals reflected from the cockpit and converts them into electrical signals. Simultaneously, it calculates the round-trip time difference of the light signal using a phase difference method to generate a 3D point cloud containing distance information. This invention uses a CMOS image sensor based on the iToF principle, achieving high single-pixel depth measurement accuracy. With a thickness of 3mm and a dynamic measurement range between 0.3m and 5m, it can meet the monitoring needs of different seat distances in the cockpit.
[0029] The module controller integrates an ARM Cortex-M4 core, primarily responsible for sending emission control commands to the infrared light source, receiving electrical signals output from the ToF sensor, and performing noise reduction and depth format conversion on the electrical signals. The module controller also transmits standardized data to the main control processing unit via the MIPI CSI-2 interface and receives parameter configuration commands from the main control processing unit via the I2C interface.
[0030] like Figure 2 As shown, the main control processing unit uses an automotive-grade SoC that supports ASIL-B functional safety. It integrates three key modules: a data preprocessing module, a core analysis engine, and a decision output module. The data processing module is used to preprocess the acquired 3D point cloud data. The specific process is as follows: Step 1: Filter the 3D point cloud with a kernel of 3. The Gaussian filtering process in step 3 removes salt-and-pepper noise from the 3D point cloud by calculating the weighted average of neighboring pixels, while preserving the edge details of the object.
[0031] Step 2: Fit the static plane within the cockpit using the RANSAC algorithm. The static plane includes the seat back and instrument panel. During the fitting process, set the plane fitting error to 5mm and extend the depth value beyond the plane. The point cloud within the range is determined as a dynamic occupant point cloud, achieving separation between dynamic targets and static background.
[0032] Step 3: Perform point cloud registration and cockpit coordinate system normalization to transform the occupant point cloud from the camera coordinate system to the cockpit coordinate system. The cockpit coordinate system has the midpoint of the driver's seat rail as its origin, with the X-axis along the vehicle's forward direction and the Z-axis perpendicular to the ground. The camera coordinate system has the optical center of the ToF camera module as its origin. During the transformation, point cloud registration is used to eliminate the deviation caused by the ToF camera module's installation angle, ensuring the coordinate consistency of multiple frames of point cloud data in the same coordinate system.
[0033] Step 4: Select derived point clouds based on the type of monitoring target. The derived point clouds include head point clouds, chest point clouds, and hand point clouds. The selected head point clouds are between 0.2 and 0.4 m deep and located above the torso. The chest point clouds are located in the upper part of the torso and are between 0.3 and 0.5 m deep. The hand point clouds are located within 5 cm around the wrist and finger joints.
[0034] The core analysis engine includes an identity recognition module, a posture and behavior analysis module, a vital signs monitoring module, and an attention state analysis module. It receives preprocessed 3D point clouds and performs multi-dimensional monitoring and analysis on the derived point clouds. The specific process is as follows: The identity recognition submodule first uses a voxel grid downsampling algorithm to reduce the resolution of the head point cloud from 1000 to 10000. The downsampled head point cloud is then input into the lightweight 3D face recognition model MobileFaceNet-3D. The NPU inference extracts a 128-dimensional feature vector from the head point cloud (depth range 0.2-0.4m, located above the torso). The feature vector is then matched with the user feature template in the local encrypted storage unit using Euclidean distance: if the user feature template is continuous... The matching value of the frame is less than or equal to If the user identification is successful, the user ID is recorded; if it is continuous... The frame's matching value is greater than If the user identity is not identified, this submodule supports storing 5 user feature templates.
[0035] The posture and behavior analysis submodule first performs voxelization on the occupant point cloud. It divides the 3D space into several voxel grids, records whether each voxel grid contains point clouds, and forms a voxel feature map. Then, the voxel feature map is input into the 3D skeleton keypoint detection model HRNet-3D, which outputs the data in real time. The system uses the three-dimensional coordinates of several joints, covering the head, neck, shoulders, elbows, wrists, fingertips, torso, hips, knees, and ankles. Gesture recognition is then performed by calculating the displacement trajectories and angular changes of the wrist and fingertip joints. Furthermore, the system determines whether the occupant is leaning excessively forward by calculating the distance between key torso points and the seat back plane. Key torso points include the root of the neck, sternum, and lower back. When the distance between these points is less than [a certain value], [the system determines whether the occupant is leaning forward excessively]. And the duration exceeds When assessing excessive forward leaning, the system calculates the distance between the shoulder acromion and the seatbelt buckle point to determine if the occupant is not wearing a seatbelt. At that time, it was determined that the occupant was not wearing a seat belt.
[0036] The vital signs monitoring module first locates the point cloud of the chest cavity region (depth range 0.3-0.5m, located in the upper half of the torso), then acquires the average depth value of the chest cavity region at a sampling frequency of 5Hz, generating a data set containing... The time series signal of each data point is processed by using a bandpass filter of 0.1-0.5Hz to extract respiratory signals and a bandpass filter of 0.8-3.0Hz to extract heartbeat signals. Then, a peak detection algorithm is used to calculate respiratory rate and heart rate. Users can define abnormal thresholds for vital signs, such as a heart rate greater than 100 beats / minute or less than 60 beats / minute. When the respiratory rate and heart rate exceed the threshold range, it is judged as abnormal and an alert is issued.
[0037] The attention state analysis module calculates the head orientation vector based on the coordinates of key head points, including the top of the head and the chin, and then calculates the angle between the head orientation vector and the vehicle's direction of travel. Simultaneously, the PERCLOS value was calculated by analyzing the density changes of the point cloud in the eye region. The PERCLOS value represents the percentage of eyelid closure time per minute. When the included angle And the duration exceeds Or when the PERCLOS value is greater than When the PERCLOS value is greater than a certain value, it is considered inattentive; when the PERCLOS value is greater than a certain value, it is considered inattentive. At that time, it was determined to be fatigued driving, and a graded warning was issued.
[0038] The decision output module is used to generate control commands and early warning signals based on the analysis results of the core analysis engine. The specific process is as follows: When identity recognition is successful, the decision output module sends a parameter configuration command to the cockpit control unit. The parameter configuration command includes the three-dimensional coordinates of the seat. , Coordinates at scope, Coordinates at scope, Coordinates at Range. Horizontal angle of the rearview mirror. exist to Range and vertical angle exist to Range, air conditioning temperature exist to The range also includes the audio volume in the range of 0 to 30 dB, with the actuator completing the adjustment within 3 seconds.
[0039] When a valid gesture is recognized, the decision output module generates actuator control commands corresponding to the configuration instructions, such as a right hand sliding up or down corresponding to the air conditioner fan speed. The system will send a command to vibrate the seat at a frequency of 50 Hz for 0.5 seconds and flash a red icon on the dashboard when dangerous behavior is detected.
[0040] When respiratory rate and heart rate are detected to be outside the threshold range and the vital signs are abnormal, the decision output module displays a warning message on the HUD, such as suggesting that the car stop and rest if the heart rate is too high. At the same time, a voice prompt is played through the speaker at a volume of 20dB for 2 seconds.
[0041] When abnormal attention is detected, a tiered warning system is activated: Level 1 warning involves unilateral seat vibration for 0.3 seconds; Level 2 warning involves seat vibration combined with a voice prompt to remind the driver to concentrate; Level 3 warning involves seat vibration, a voice prompt, a red flashing indicator on the instrument panel, and a warning icon displayed on the HUD. Each warning level is spaced 3 seconds apart. The angle between the head orientation vector and the vehicle's direction of travel is also considered. And the duration reached Or when the PERCLOS value is less than or equal to The warning was lifted at that time.
[0042] Example 1 To verify whether the system's performance in actual vehicles (gasoline and electric vehicles) meets the design requirements, multi-scenario real-vehicle tests were conducted, covering four core dimensions: environmental adaptability, real-time performance, reliability, and privacy protection. The specific test results are as follows.
[0043] I. Conduct environmental adaptability testing to verify the system's stable operation under complex environments such as extreme temperatures, strong light, and weak light, ensuring that it meets the all-weather working requirements of vehicle-mounted scenarios.
[0044] ① Temperature adaptability test: First, place the test vehicles in different locations. Low-temperature constant temperature environment chamber and High-temperature constant-temperature environment chamber, continuous operation for 4 hours, monitoring system accuracy and hardware status: like Figure 3As shown, in In the low-temperature constant-temperature environment chamber, the system operated without any crashes or functional interruptions, with only a slight decrease in the accuracy of core functions. Specifically, the accuracy of identity recognition improved from that at room temperature. Down to Gesture recognition accuracy at room temperature Down to The decrease in accuracy for both items was less than This meets the cabin monitoring needs in low-temperature environments.
[0045] exist In the high-temperature constant temperature environment chamber, the temperature of the main control SoC is stably maintained within the high-temperature constant temperature environment chamber. Far lower than the high-temperature constant temperature environment chamber The safety threshold is met, with no risk of overheating; the depth measurement error of the ToF camera module is reduced from that at room temperature. Increase to The accuracy rate is still within an acceptable range; the overall system function triggering accuracy rate reaches [percentage missing]. No false triggering or missed triggering occurred due to high temperature.
[0046] ② Light adaptability test: Under strong midday light (100,000 lux) and no external light source at night (0.1 lux), the system's adaptability to different light conditions was tested. When the midday sun shines directly into the cockpit, the system is less affected by ambient light, and the gesture recognition error rate is lower than under normal temperature and low light conditions. Increase to The head angle calculation error in attention monitoring has increased from that under normal temperature and low light conditions. Increase to The core functionality remains stable and reliable.
[0047] When there is no external light source at night, the ToF camera module uses active light to supplement illumination, and the accuracy of vital sign monitoring is not affected, with the respiratory rate measurement error remaining within a certain range. beats / minute, with heart rate measurement error maintained at [value missing]. The frequency is [number] times per minute, which is completely consistent with the monitoring accuracy under normal temperature conditions.
[0048] II. Conduct real-time testing, using an accuracy of [insert accuracy here]. A high-precision timer records the entire process time from ToF data acquisition to actuator response, ensuring that the design requirements for millisecond-level response of the vehicle system are met.
[0049] ① Overall process time statistics: The entire system process is divided into five stages: ToF camera module data acquisition, point cloud preprocessing, core analysis, command generation and transmission, and actuator response. The time consumption of each stage and the total delay are shown in Table 1 and 2. Figure 4 As shown: Table 1. Total Process Time Statistics
[0050] The test results show that the total latency of the entire system is only 88ms, which is less than the 100ms threshold for in-vehicle real-time performance. This enables a rapid closed loop of perception, decision-making, and execution, avoiding interaction stuttering or security risks caused by latency.
[0051] Third, reliability testing: the system's stability under real-world driving scenarios, such as long-term operation, high-speed driving, and bumpy roads, is verified through actual vehicle road driving.
[0052] ① Continuous operation test: The test vehicle was run continuously for 72 hours in urban road and highway scenarios to monitor system stability: The system did not crash or lose data throughout the entire process, and processed more than 770,000 frames of raw 3D point cloud data collected by the ToF camera module; the overall function trigger accuracy reached [percentage missing]. The false trigger rate is only Furthermore, the false triggers are all concentrated in gesture recognition, such as misjudging natural hand movements as volume adjustment. The false judgment rate can be further reduced by optimizing the model threshold.
[0053] ② Anti-interference test: Under the scenario of high-speed vehicle travel and bumpy road surface (simulating a third-class highway, amplitude 5cm), the system's vibration resistance and anti-interference capability were tested: During high-speed travel and bumpy road conditions, the ToF camera module achieved vibration reduction through the metal bracket, and the point cloud noise rate was only [missing information]. less than The design threshold does not affect the accuracy of subsequent point cloud segmentation and feature extraction; after the cockpit control unit receives the command, the actuators (seat motor, vibration module, etc.) move without delay, and there is no loss of command or execution deviation due to bumps.
[0054] IV. Privacy Protection Test: Using professional tools, we attempted to read and recover system data to verify the effectiveness of the end-to-end data security mechanism.
[0055] ① Data extraction test: Using professional tools such as memory readers and storage chip analyzers, we attempted to read data from the main control memory and the local encrypted storage unit. Only the encrypted 128-dimensional facial feature vector could be read from the encrypted storage unit. It was in garbled form and could not be restored to head point cloud, facial details or original depth image. No original point cloud or derived point cloud data was detected in the main control memory (because the original point cloud was forcibly destroyed within 100ms after preprocessing), to avoid the risk of sensitive data leakage.
[0056] ② Data destruction test: After triggering the system's original data deletion and feature template clearing commands, data recovery software such as DiskGenius is used to scan the main control memory and storage unit: the original point cloud data has been completely destroyed by overwriting 0 values, with no residue; after the feature template in the encrypted storage unit is cleared, the data cannot be recovered, ensuring that users can independently control their private data.
[0057] V. Test Conclusions: Based on the comprehensive results of multi-scenario real-vehicle tests, the intelligent cockpit monitoring system based on the ToF camera module meets the following design requirements: In terms of environmental adaptability, to Under temperature range and illumination range of 0.1 lux to 100,000 lux, the decrease in core function accuracy is less than It can operate stably around the clock. In terms of real-time performance, the total latency of the entire process is less than 100ms (88ms), meeting the real-time response requirements of the vehicle system.
[0058] In terms of reliability, it can run continuously for 72 hours without failure, with point cloud noise rate of less than 3% in high-speed and bumpy scenarios, and function trigger accuracy of 97%. In terms of privacy protection, data extraction only yields encrypted feature vectors, and the original data is completely destroyed, with no risk of leakage.
[0059] Therefore, the intelligent cockpit monitoring system based on an optical time-of-flight camera adopted in this invention has achieved the above-mentioned performance, meets the intelligent cockpit monitoring requirements of both fuel vehicles and electric vehicles, and has value for mass production applications in real vehicles.
[0060] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can still be made to the technical solutions of the present invention, and these modifications or equivalent substitutions cannot cause the modified technical solutions to deviate from the spirit and scope of the technical solutions of the present invention.
Claims
1. An intelligent cockpit monitoring system based on an optical time-of-flight camera, characterized in that: It includes a ToF camera module, a main control processing unit, a cockpit control unit, and a local encrypted storage unit. The ToF camera module is electrically connected to the main control processing unit through the MIPI CSI-2 interface and is used to collect three-dimensional point clouds of the entire cockpit area. The main control processing unit is used to receive 3D point clouds, preprocess and analyze them, and then issue control commands and warning signals. The cockpit control unit is electrically connected to the main control processing unit via vehicle Ethernet to receive control commands and warning signals and drive the actuators in the cockpit. The local encrypted storage unit is used to store the user's feature templates and is connected to the main control processing unit via the SPI interface to store the preprocessed point clouds.
2. The intelligent cockpit monitoring system based on an optical time-of-flight camera according to claim 1, characterized in that: The ToF camera module includes an infrared light source, optical components, a ToF sensor, and a module controller. The infrared light source emits modulated infrared light. The optical components include an aspherical condenser lens and a narrowband filter. The aspherical condenser lens focuses the infrared light, and the narrowband filter has a bandwidth of [missing information]. Used to filter out interference from ambient light; The ToF sensor is used to receive infrared light signals reflected by the cockpit and convert the light signals into electrical signals. At the same time, it calculates the round-trip time difference of the light signals based on the phase difference method to generate a three-dimensional point cloud containing distance information. The module controller is used to send light emission control commands to the infrared light source, receive the electrical signals output by the ToF sensor, and perform noise reduction and format conversion processing on the electrical signals.
3. The intelligent cockpit monitoring system based on an optical time-of-flight camera according to claim 2, characterized in that: The module controller also transmits standardized data to the main control processing unit via the MIPI CSI-2 interface and receives parameter configuration commands sent by the main control processing unit via the I2C interface.
4. The intelligent cockpit monitoring system based on an optical time-of-flight camera according to claim 1, characterized in that: The main control processing unit includes a data preprocessing module, a core analysis engine, and a decision output module. The data processing module is used to preprocess the acquired 3D point cloud, and the core analysis engine is used to receive the preprocessed 3D point cloud and perform multi-dimensional monitoring and analysis on the derived point cloud. The decision output module is used to generate control commands and early warning signals based on the analysis results of the core analysis engine.
5. The intelligent cockpit monitoring system based on an optical time-of-flight camera according to claim 4, characterized in that, The data processing module performs the following preprocessing steps on the acquired 3D point cloud: Gaussian filtering is applied to the 3D point cloud to remove salt-and-pepper noise by calculating the weighted average of neighboring pixels. Based on the RANSAC algorithm, the static planes in the cockpit are fitted. The static planes include the seat back and the instrument panel. The plane fitting error is set, and the point clouds with depth values exceeding the threshold range are judged as dynamic occupant point clouds. Transform the occupant point cloud from the camera coordinate system to the cockpit coordinate system, and eliminate the deviation caused by the installation angle of the ToF camera module through point cloud registration; Derivative point clouds are selected based on the type of monitoring target. These derived point clouds include head point clouds, chest point clouds, and hand point clouds.
6. The intelligent cockpit monitoring system based on an optical time-of-flight camera according to claim 5, characterized in that: The core analysis engine performs multi-dimensional monitoring and analysis of the derived point cloud through an identity recognition module, a posture and behavior analysis module, a vital signs monitoring module, and an attention state analysis module. The identity recognition submodule is used for occupant identification, and the specific identification process is as follows: First, a voxel lattice downsampling algorithm is used to reduce the resolution of the head point cloud. Then, the downsampled head point cloud is input into the MobileFaceNet-3D model, and the feature vector of the head point cloud is extracted through NPU inference. Finally, the feature vector is matched with the user feature template in the local encrypted storage unit using Euclidean distance. If user feature templates are continuous The matching value of the frame is less than or equal to If the user identity is successfully identified, the user ID is recorded. If continuous The frame's matching value is greater than If so, it is determined that the user's identity has not been identified.
7. The intelligent cockpit monitoring system based on an optical time-of-flight camera according to claim 6, characterized in that: The posture behavior analysis module is used to analyze occupant posture behavior. The specific analysis process is as follows: First, the 3D space is divided into several voxel meshes to form voxel feature maps; then, the voxel feature maps are input into the HRNet-3D model and output. The three-dimensional coordinates of each joint point cover the head, neck, shoulder, elbow, wrist, fingertips, torso, hip, knee, and ankle. Next, gesture recognition is performed by calculating the displacement trajectory and angle changes of the wrist and fingertip joints. Then, by calculating the distance between key points of the torso and the seat back plane, it is determined whether the occupant is leaning excessively forward. Key torso points include the base of the neck, sternum, and lower back. When the distance between the two is less than... And the duration exceeds If so, it is judged that the body is leaning forward excessively; By calculating the distance between the shoulder acromion and the seatbelt buckle point, it can be determined whether the occupant is not wearing a seatbelt. The shoulder acromion is the key point on the shoulder. When the distance between the two is greater than [missing information], it indicates that the occupant is not wearing a seatbelt. At that time, it was determined that the occupant was not wearing a seat belt.
8. The intelligent cockpit monitoring system based on an optical time-of-flight camera according to claim 7, characterized in that: The vital signs monitoring module is used for monitoring the vital signs of occupants. The specific monitoring process is as follows: First, the point cloud of the thoracic cavity region is located. Then, the average depth value of the thoracic cavity region is collected to generate a data structure containing... The time series signal of each data point was used to extract respiratory and heartbeat signals through bandpass filtering. Then, the peak detection algorithm is used to calculate the respiratory rate and heart rate. The user defines the abnormal threshold for vital signs. When the respiratory rate and heart rate exceed the threshold range, it is judged as abnormal and an early warning is issued.
9. The intelligent cockpit monitoring system based on an optical time-of-flight camera according to claim 8, characterized in that: The attention state analysis module is used for driver attention assessment. The specific process is as follows: First, calculate the head orientation vector based on the coordinates of key head points, including the top of the head and the chin. Then, calculate the angle between the head orientation vector and the vehicle's direction of travel. Simultaneously, the PERCLOS value was calculated by analyzing the density changes of the point cloud in the eye region. The PERCLOS value represents the percentage of eyelid closure time per minute. When the included angle And the duration exceeds Or when the PERCLOS value is greater than At that time, it was determined to be a case of inattention; When the PERCLOS value is greater than At that time, it was determined to be fatigued driving, and a graded warning was issued.
10. The intelligent cockpit monitoring system based on an optical time-of-flight camera according to claim 9, characterized in that: The decision output module generates control commands and early warning signals based on the analysis results from the core analysis engine. The specific process is as follows: When identity recognition is successful, the decision output module sends a parameter configuration command to the cockpit control unit. The parameter configuration command includes the three-dimensional coordinates of the seat, the horizontal and vertical angles of the rearview mirror, and the air conditioning temperature. When a valid gesture is recognized, the decision output module generates the actuator control command corresponding to the configuration command; When dangerous behavior is detected, commands are sent to vibrate the seat and flash a red icon on the dashboard; When respiratory rate and heart rate are detected to be abnormal vital signs that exceed the threshold range, the decision output module displays warning text through the HUD and plays voice prompts through the speaker. When abnormal attention is detected, a tiered warning system is activated: Level 1 warning is seat vibration on one side, Level 2 warning is seat vibration plus voice prompt, and Level 3 warning is seat vibration, voice prompt, red flashing on the instrument panel and warning icon displayed on the HUD. When the head is facing the angle between the vector and the vehicle's direction of travel And the duration is equal to Or when the PERCLOS value is less than or equal to The warning was lifted at that time.