Intelligent vehicle face unlocking method and system based on multi-modal fusion
Patent Information
- Application Number
- CN202610698190.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-20
- Publication Date
- 2026-09-29
AI Technical Summary
[0003]为解决现有技术中智能车辆人脸解锁方案存在的静态特征对比易被攻击、活体检测机制简单导致误解锁风险较高的问题,本发明提供了一种基于多模态融合的智能车辆人脸解锁方法及系统,通过融合时序RGB图像、深度图像、近红外图像和热成像等多模态数据,结合多模态一致性检验,设计高鲁棒性的活体检测机制,显著提升人脸解锁的安全性、准确性和用户体验
本发明提供了一种基于多模态融合的智能车辆人脸解锁方法及系统,通过先进行静态人脸特征匹配再进行活体检测的两阶段验证策略,有效降低了计算资源消耗,匹配失败时直接拒绝解锁无需启动多模态传感器,提升了系统效率。其中,通过融合时序RGB图像、深度图像、近红外图像和热成像等多模态数据,并结合多模态一致性检验,构建高鲁棒性的活体检测机制,能够有效防御照片攻击、视频攻击、面具攻击、深度伪造攻击等多种欺骗手段,显著提升人脸解锁的安全性;通过引入环境自适应阈值调整策略,根据环境光照度动态调整活体检测阈值和人脸匹配阈值,使系统能够适应暗光、强光等复杂环境条件,保证了在不同场景下的识别准确率和用户体验;此外,本发明还通过引入四级权限体系,针对不同用户身份提供差异化的车辆控制权限,既保证了主用户和家庭成员的便捷使用,又对临时用户和代驾模式进行了安全限制,提升了系统的灵活性和安全性。
Smart Images

Figure CN122830607A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent vehicle unlocking technology, and in particular to an intelligent vehicle face unlocking method and system based on multimodal fusion. Background Technology
[0002] With the rapid development of intelligent vehicle technology, vehicle unlocking methods are gradually evolving from traditional mechanical keys and remote keys to biometric technologies. Facial recognition technology is increasingly widely used due to its contactless and convenient nature. Current methods primarily involve collecting user facial images and extracting static features for comparison and matching to achieve identity verification and vehicle unlocking. However, static feature comparisons are susceptible to deception techniques such as photo, video, and mask attacks, leading to a high risk of false unlocks. Although liveness detection mechanisms have been introduced to assist facial recognition, existing mechanisms are relatively simple, relying solely on single-modal physiological signal detection. For example, they may rely solely on visible light cameras to detect blinking, mouth opening, head turning, etc., or solely on depth cameras to detect facial depth information. These mechanisms lack dynamic processing and cross-validation of multimodal information, making them vulnerable to attacks such as deepfakes, resulting in a risk of false unlocks. Summary of the Invention
[0003] To address the problems of static feature comparison being vulnerable to attack and the simple liveness detection mechanism leading to a high risk of false unlocking in existing intelligent vehicle face unlocking schemes, this invention provides an intelligent vehicle face unlocking method and system based on multimodal fusion. By fusing multimodal data such as temporal RGB images, depth images, near-infrared images, and thermal imaging, and combining multimodal consistency checks, a highly robust liveness detection mechanism is designed, significantly improving the security, accuracy, and user experience of face unlocking.
[0004] In a first aspect, the present invention provides a method for unlocking intelligent vehicles using facial recognition based on multimodal fusion.
[0005] A method for unlocking intelligent vehicles using facial recognition based on multimodal fusion, comprising: When the vehicle is in the unlocking state, it detects human signals around the vehicle in real time. If a human body is detected approaching, it collects a static frontal face image of the approaching human body and performs feature matching with a preset authorized face image. If the match fails, the unlocking is refused; otherwise, a liveness detection is performed. This involves simultaneously acquiring temporal RGB images, depth images, near-infrared images, and thermal images of the human face. Multimodal dynamic features of the face are extracted through a temporal convolutional network. A liveness confidence score is then output after a multimodal consistency test. Combined with a liveness detection threshold that is adaptively adjusted based on the environment, the face unlock verification is determined to be successful. If successful, the corresponding vehicle control command will be generated based on the verified user's identity and permission level to drive the vehicle to perform the corresponding unlocking operation and function pre-start operation.
[0006] A further technical solution involves the following static facial feature matching process: A static frontal image of a human face close to the body is acquired and preprocessed; the preprocessing includes global and local contrast normalization. A facial feature extraction model is used to extract key feature points from the preprocessed facial image. These points are then compared with the corresponding key feature points of standard facial images pre-stored in the authorized facial feature database. The matching result is determined based on the comparison, including: Using the extracted key feature points of the face image as the query vector, the key feature vectors of all registered users are read from the authorized face feature library stored in the vehicle. The cosine similarity between the query vector and each feature vector in the library is calculated. If the maximum similarity is greater than or equal to the face matching threshold, the match is considered successful. The maximum similarity and the corresponding user ID are taken. Otherwise, the match is considered unsuccessful.
[0007] A further technical solution, wherein the multimodal consistency check and temporal causality verification include: The study investigates the spatial alignment error between the texture and depth images of RGB face images, the correlation between thermal imaging temperature gradient and skin color distribution in RGB face images, and the frequency domain consistency between near-infrared reflectance and texture in RGB face images. Based on the probabilities obtained from the tests, the liveness confidence score is calculated through weighted fusion.
[0008] Further technical solutions for spatial alignment error detection between RGB face image texture and depth image include: A face detection algorithm is used to detect face bounding boxes and multiple key feature points in RGB face images and depth images, respectively; The difference in pixel coordinates of the corresponding key points in the RGB image plane and the depth image plane is calculated to obtain the spatial alignment error vector; Calculate the L2 norm of the spatial alignment error vector. If the L2 norm is less than a set pixel threshold, it is determined to be a natural registration error of a real face; otherwise, it is determined to be a lack of depth information caused by a planar attack. Output the alignment consistency probability. .
[0009] Further technical solutions include the detection of the correlation between thermal imaging temperature gradient and skin color distribution in RGB face images, including: In RGB images, facial skin regions are extracted using a skin color model in the YCrCb color space, and the spatial distribution density map of the skin regions is calculated. Calculate the facial temperature gradient field in thermal imaging images and extract high gradient regions where the temperature gradient amplitude is greater than a set temperature gradient threshold. Calculate the spatial correlation coefficient between the skin density map and the temperature gradient field, and combine it with the correlation coefficient of a real face to output the temperature correlation probability. .
[0010] Further technical solutions, including frequency domain consistency detection of near-infrared reflectance and RGB face image texture, include: Two-dimensional discrete Fourier transforms were performed on the near-infrared image and the RGB grayscale image respectively to obtain the frequency domain amplitude spectrum; Extract the mid-frequency component of the 0.1-0.5 period / pixel frequency band in the frequency domain amplitude spectrum, which corresponds to the facial skin texture features; Calculate the normalized cross-correlation coefficients of the near-infrared and visible light frequency domain amplitude spectra, and combine them with the cross-correlation coefficients of real faces to output the frequency domain consistency probability. .
[0011] Further technical solutions, including adaptively adjusting the liveness detection threshold according to the environment, include: Using an ambient light sensor, the ambient light level of the current vehicle's surroundings is monitored in real time; Based on the comparison between the ambient light intensity and the preset threshold values for each light intensity level, the system determines whether the current environment is a dark environment, a bright environment, or a normal lighting environment, and then determines the corresponding liveness detection threshold based on the different types of environments.
[0012] A further technical solution divides user identity and permission levels into four levels, including: The primary user has the privileges to unlock all vehicle doors, start the vehicle, adjust all personalized settings, and modify vehicle system settings. These privileges are permanently valid. Family member permissions allow unlocking all vehicle doors, starting the vehicle, and adjusting personalized settings, but not modifying system settings; these permissions are permanent. Temporary user privileges allow unlocking the driver's side door, but the vehicle cannot be started. These privileges are time-limited and expire after a default time. The designated driver mode grants permissions to unlock the driver's side door and start the vehicle, but limits the maximum speed, prohibits access to personal data, and automatically records driving routes. These permissions are time-limited and expire after a default time setting.
[0013] Secondly, the present invention provides an intelligent vehicle face unlocking system based on multimodal fusion.
[0014] A smart vehicle face unlocking system based on multimodal fusion includes: The human detection module includes a passive infrared human sensor, which is used to detect human signals around the vehicle in real time when the vehicle is in a state of waiting to be unlocked. The image acquisition and preprocessing module is used to acquire and preprocess a static frontal face image of a person approaching the body if a person is detected to be approaching; and to simultaneously acquire temporal RGB images, depth images, near-infrared images and thermal images of the human face. The face feature matching module is used to match the preprocessed face image with the preset authorized face image. If the match fails, the unlocking will be refused; otherwise, a liveness detection will be performed. The liveness detection module is used to extract multimodal dynamic features of the face through a temporal convolutional network, and then output a liveness confidence score through a multimodal consistency test. Combined with the liveness detection threshold that is adaptively adjusted based on the environment, it determines whether the face unlock verification is successful. The access control module is used to generate corresponding vehicle control commands based on the verified user's identity and access level, so as to drive the vehicle to perform corresponding unlocking operations and function pre-start operations.
[0015] Thirdly, the present invention also provides a vehicle for executing vehicle control commands corresponding to the above-mentioned intelligent vehicle face unlocking method based on multimodal fusion, or including the above-mentioned intelligent vehicle face unlocking system based on multimodal fusion.
[0016] The above one or more technical solutions have the following beneficial effects: This invention provides a method and system for intelligent vehicle face unlocking based on multimodal fusion. By employing a two-stage verification strategy—first performing static facial feature matching and then liveness detection—it effectively reduces computational resource consumption. When matching fails, unlocking is directly refused without activating the multimodal sensors, improving system efficiency. Specifically, by fusing multimodal data such as temporal RGB images, depth images, near-infrared images, and thermal imaging, and combining this with multimodal consistency checks, a highly robust liveness detection mechanism is constructed. This mechanism effectively defends against various deception techniques such as photo attacks, video attacks, mask attacks, and deepfake attacks, significantly improving the security of face unlocking. By introducing an environment-adaptive threshold adjustment strategy, the liveness detection threshold and face matching threshold are dynamically adjusted according to ambient light levels, enabling the system to adapt to complex environmental conditions such as low light and strong light, ensuring recognition accuracy and user experience in different scenarios. Furthermore, this invention introduces a four-level permission system, providing differentiated vehicle control permissions for different user identities. This ensures convenient use for the primary user and family members while imposing security restrictions on temporary users and chauffeur services, enhancing the system's flexibility and security.
[0017] Advantages of additional aspects of the invention will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of the invention. Attached Figure Description
[0018] The accompanying drawings, which form part of this invention, are used to provide a further understanding of the invention. The illustrative embodiments of the invention and their descriptions are used to explain the invention and do not constitute an improper limitation of the invention.
[0019] Figure 1 This is an overall flowchart of the intelligent vehicle face unlocking method based on multimodal fusion in Embodiment 1 of the present invention. Detailed Implementation
[0020] It should be noted that the following detailed descriptions are exemplary and are intended only to describe specific embodiments and to provide further explanation of the invention, and are not intended to limit the scope of exemplary embodiments of the invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. Furthermore, it should be understood that when the terms "comprising" and / or "including" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.
[0021] Example 1 As the background section points out, existing smart vehicle face unlocking solutions based on static feature comparison are vulnerable to various attack methods, and their liveness detection mechanisms are too simplistic and lack robustness: static face feature matching relies solely on a single facial image for identity recognition, allowing attackers to forge user facial images using high-definition photos, video playback, 3D masks, etc., bypassing the identity verification system and leading to a high risk of false unlocks; existing liveness detection mechanisms rely solely on single-modality physiological signal detection, such as detecting blinking only through a visible light camera, which attackers can use to deceive the system by playing videos containing blinking, or detecting facial depth information only through a depth camera, which attackers can use to forge depth information using 3D-printed masks, lacking cross-validation of multimodal information and making them susceptible to new attack methods; existing solutions typically use a single fixed threshold for identity verification and liveness detection, failing to dynamically adjust based on user history and environmental conditions, making it difficult to strike a balance between security and convenience, being too strict for high-trust users and affecting user experience, while being too lenient for low-trust users poses security risks. In addition, existing solutions lack the ability to adapt to different environmental conditions. In low light, image quality deteriorates, leading to inaccurate feature extraction. In strong light, facial overexposure causes loss of feature information. In vehicle vibration, image blurring increases the matching failure rate. These environmental factors can significantly reduce recognition accuracy and affect user experience.
[0022] Based on this, this embodiment proposes a multimodal fusion-based intelligent vehicle face unlocking method. It employs a two-stage verification strategy, first performing static facial feature matching followed by multimodal liveness detection. This is combined with strategies such as global and local contrast normalization, multimodal consistency verification, environmentally adaptive threshold adjustment, and a four-level permission system. This addresses the technical shortcomings of existing technologies, such as susceptibility to attacks on static feature comparison, simplistic liveness detection mechanisms, and weak environmental adaptability, significantly improving the security, accuracy, and user experience of intelligent vehicle face unlocking. Figure 1 As shown, the method proposed in this embodiment specifically includes the following steps: Step S1: When the vehicle is in the unlocking state, human signals around the vehicle are detected in real time. If a human body is detected approaching, a static frontal face image of the approaching human body is collected and matched with the preset authorized face image for feature matching.
[0023] Step S101: When the vehicle is in the unlocking state, the passive infrared human body sensor detects human body signals around the vehicle in real time. If a human body is detected approaching, a static frontal face image of the approaching human body is captured by a depth camera.
[0024] In this embodiment, the vehicle being in an unlocking state means that the vehicle has been turned off and locked, and is in standby mode, waiting for an authorized user to approach and perform the unlocking operation. In this state, the vehicle domain controller is in a low-power standby mode, keeping only the passive infrared human body sensor active, while other high-power modules such as depth cameras, near-infrared cameras, and thermal imaging cameras are turned off to reduce the vehicle's standby power consumption.
[0025] Passive infrared human body sensors are deployed on the exterior of the vehicle, specifically near the driver's side door handle, below the door mirror, or on the B-pillar. These sensors have a detection range of 2 meters and a detection angle of 110°, covering the main approach area on the driver's side. The passive infrared human body sensor determines the presence of a human body by detecting changes in infrared radiation in the environment. The human body temperature is approximately 36 to 37°C, emitting infrared radiation with a wavelength of approximately 10 μs. When a person enters the sensor's detection range, the pyroelectric element inside the sensor detects the change in infrared radiation and outputs a detection signal.
[0026] The passive infrared human body sensor communicates with the vehicle domain controller via the vehicle's CAN bus. When the sensor detects a human body approaching, it sends a human body detection signal to the vehicle domain controller via the CAN bus. This signal contains information such as the detection timestamp and detection confidence level. After receiving the human body detection signal, the vehicle domain controller first determines whether the detection confidence level exceeds a preset detection threshold (e.g., 0.8) to filter out false detections caused by environmental interference such as wind blowing leaves or small animals passing by. If the detection confidence level exceeds the threshold, it is determined that a human body has indeed approached the vehicle. The vehicle domain controller immediately switches from low-power standby mode to normal operating mode and sends a wake-up command and an image acquisition command to the depth camera.
[0027] After receiving the wake-up command, the depth camera first performs hardware initialization, including powering on the image sensor, autofocusing the lens, and preheating the infrared projector. After initialization, the depth camera begins to acquire static frontal facial images close to the human body. The acquisition process includes two parts: RGB image acquisition and depth image acquisition. RGB image acquisition is completed by a visible light image sensor, using automatic exposure and automatic white balance algorithms to ensure clear color images under different lighting conditions. Depth image acquisition can be completed using structured light or time-of-flight technology. A structured light pattern is projected onto the face by an infrared projector, and the reflected structured light pattern is acquired by an infrared image sensor. The depth value of each pixel is calculated based on the degree of distortion of the pattern.
[0028] During the above process, the RGB image and depth image captured by the depth camera are time-stamped and aligned to ensure that the two images are captured at the same time with a timestamp error of no more than 10ms. The synchronized image data is then transmitted to the vehicle domain controller via the vehicle Ethernet. After receiving the image data, the vehicle domain controller stores it in a memory buffer for subsequent image preprocessing and facial feature extraction.
[0029] It's important to note that when capturing static frontal face images, the depth camera uses a face detection algorithm to determine in real-time whether the captured image contains a face and whether the face is facing forward. The face detection algorithm employs the MTCNN multi-task cascaded convolutional neural network based on deep learning. This algorithm can quickly locate the face region in the image and output the face bounding box and the coordinates of five key points, including the center of the left eye, the center of the right eye, the tip of the nose, the left corner of the mouth, and the right corner of the mouth. By calculating the angle between the line connecting the centers of the left and right eyes and the horizontal line, the angle of face rotation can be determined. Only when the rotation angle is less than 15 degrees is it considered a frontal face; otherwise, the user is prompted to adjust their facial orientation. By calculating the relative positions of the five key points and the face bounding box, it can be determined whether the face is complete and occluded. Only when all five key points are within the bounding box and are not occluded is the image considered a valid frontal face image.
[0030] Furthermore, in practical applications, to enhance user experience, the depth camera continuously captures multiple frames of images, selecting the highest quality frame as the static frontal face image. Image quality evaluation metrics include sharpness, brightness, contrast, and face integrity. Sharpness is assessed by calculating the Laplacian variance of the image; a higher value indicates a sharper image. Brightness is assessed by calculating the average grayscale value of the image. Contrast is assessed by calculating the standard deviation of the image; a higher value indicates higher contrast. Face integrity is assessed by the confidence score of the face detection algorithm; a higher confidence score indicates a more complete face. The system selects the frame with the highest overall quality score from the 10 continuously captured images as the static frontal face image for subsequent feature matching.
[0031] Through the above steps, human proximity detection and static frontal face image acquisition in the vehicle's unlocking state are realized, providing high-quality input data for subsequent face feature matching. Among them, the use of a passive infrared human proximity sensor for human proximity detection can significantly reduce the system's standby power consumption and extend the vehicle's battery life compared to continuously running a camera for video monitoring. The use of a depth camera to simultaneously acquire RGB images and depth images provides a data foundation for subsequent multimodal liveness detection.
[0032] Step S102: Preprocess the acquired static frontal face image. The preprocessing includes global contrast normalization and local contrast normalization to obtain a face image with enhanced contrast.
[0033] In this embodiment, image preprocessing is a crucial step in the face recognition process, directly affecting the accuracy of subsequent face feature extraction. Due to the complex and varied environmental conditions in vehicle face unlock scenarios, including low light, strong light, backlight, and sidelight, the acquired face images may suffer from uneven brightness, low contrast, and blurred details. To eliminate the impact of lighting conditions on image quality and improve image contrast and detail, this embodiment employs a preprocessing strategy combining global contrast normalization and local contrast normalization.
[0034] (1) The specific process of global contrast normalization is as follows: First, calculate the global average intensity of all pixels in the face image. Specifically, for the acquired RGB color image... ,in Indicates the first i Line 1 j The intensity of red, Indicates the first i Line 1 j The intensity of green, Indicates the firsti Line 1 j The intensity of blue, and the contrast of the entire image, can be expressed as: ; in, It is the global average intensity of all pixels in the entire image, satisfying: ; At this point, global contrast normalization can be performed: by subtracting the average value (i.e., global average intensity) from each pixel in the image and then rescaling it so that the standard deviation of the pixels is equal to a certain constant s, the image can be prevented from having varying contrast.
[0035] The global contrast normalization process described above can effectively solve the problem of an image being too dark or too bright overall, improving the overall contrast. However, it cannot solve the problem of uneven contrast in local areas of the image. For example, under side lighting, one side of a face may be brighter than the other, and global normalization cannot improve the contrast on both sides simultaneously. Therefore, the following local contrast normalization process is further performed to ensure that the contrast is normalized in each small window, rather than being normalized as a whole across the image.
[0036] As one implementation, considering that images with very low but non-zero contrast typically contain almost no information, dividing by the true standard deviation in this case usually only amplifies sensor noise or compresses artifacts. Therefore, this embodiment also introduces a regularization parameter λ to balance the standard deviation of the evaluation. This introduced parameter can constrain the denominator to be greater than or equal to ε. Specifically, for a given input image X, the output image generated by global contrast normalization... ,available: ; In the above formula, s represents the target global contrast standard deviation.
[0037] Considering that a dataset composed of cropped target objects from large images is unlikely to contain images with constant intensity, it is safe to ignore the small denominator problem by setting the adjustment parameter λ=0 in this case, and in rare cases, ε can be set to a very small value to avoid division by zero. Based on this, randomly cropped small images have more nearly constant intensity, making more extreme regularization more effective.
[0038] (2) The specific process of local contrast normalization is as follows: First, the image is traversed row by row and column by column according to pixel coordinates to determine the position of the current pixel to be processed.
[0039] Next, taking a fixed-size local window centered on the current pixel, calculate the local mean and local standard deviation of the pixels within the window. The local mean is the sum of the values of all pixels within the window, divided by the total number of pixels in the window. The local standard deviation is the sum of the squares of all pixels within the window minus the local mean, divided by the total number of pixels in the window, and finally the square root is taken. The local mean reflects the average brightness of the area surrounding the current pixel, and the local standard deviation reflects the contrast of the area surrounding the current pixel.
[0040] Finally, the center pixels are normalized based on the local mean and standard deviation. This local normalization enhances local detail information, making facial texture features such as wrinkles, blemishes, and pores more clearly visible, thus improving the accuracy of facial feature extraction.
[0041] After traversal, a face image with enhanced details is obtained. The resulting image undergoes global contrast normalization and local contrast normalization, exhibiting a uniform overall brightness distribution and rich local detail information, providing high-quality input data for subsequent face feature extraction.
[0042] It should be noted that the algorithm for global and local contrast normalization is based on the Difference of Gaussian (DOG) filter. The DOG filter extracts edge and texture information from an image by calculating the difference between two Gaussian filtering results at different scales, while suppressing low-frequency illumination variations and high-frequency noise. In this embodiment, global normalization is equivalent to large-scale Gaussian filtering, used to eliminate overall illumination variations, while local normalization is equivalent to small-scale Gaussian filtering, used to enhance local texture details. The combination of the two achieves multi-scale contrast enhancement.
[0043] The above steps enable preprocessing of the acquired static frontal face images. By performing global and local contrast normalization, the influence of lighting conditions on image quality is eliminated, enhancing image contrast and detail information. This provides high-quality input data for subsequent face feature extraction, improving the accuracy and robustness of face recognition.
[0044] Step S103: Use a face feature extraction model to extract key feature points of the contrast-enhanced face image to obtain a query feature vector. Read the feature vectors of all registered users from the authorized face feature library stored in the vehicle. Calculate the cosine similarity between the query feature vector and each feature vector in the library. Take the maximum similarity and the corresponding user identity identifier. If the maximum similarity is greater than or equal to the face matching threshold, the match is considered successful; otherwise, the match is considered unsuccessful and unlocking is refused.
[0045] In this embodiment, facial feature extraction and matching are the core steps for identity recognition. The facial feature extraction model adopts a deep learning-based convolutional neural network architecture, specifically the MobileNet V3 Large backbone network. MobileNet V3 is a lightweight convolutional neural network that, through techniques such as depthwise separable convolution, inverse residual structure, and squeezed excitation module, significantly reduces the number of model parameters and computational cost while maintaining high recognition accuracy, making it suitable for deployment on edge computing platforms such as vehicle domain controllers.
[0046] Specifically, the preprocessed face image is input into the face feature extraction model. After face detection, alignment, and cropping operations, the final face feature vector is extracted. This feature vector is then subjected to L2 normalization so that the vector's magnitude is 1. The normalized feature vector is the query feature vector. This vector contains facial identity information and can be used for subsequent feature matching.
[0047] After obtaining the query feature vector Next, the feature vectors of all registered users are read from the authorized facial feature database stored in the vehicle, and feature matching is performed. The authorized facial feature database is stored in the secure storage area of the vehicle domain controller and encrypted using the AES-256 encryption algorithm to prevent data leakage. The feature database stores 512-dimensional feature vectors of all registered users. Each user corresponds to a unique user identifier, abbreviated as User ID, as well as the user's permission level information. The data structure of the feature database is in key-value pair format, where the key is the User ID and the value is a combination of the feature vector and the permission level.
[0048] Specifically, the vehicle domain controller reads the feature vectors of all registered users from the feature library. Assuming there are N registered users, it calculates the query feature vectors. With each feature vector in the library The cosine similarity is calculated by taking the vectors as a dot product, since the feature vectors have already undergone L2 normalization and have a magnitude of 1. The value of cosine similarity ranges from -1 to +1; the closer the value is to 1, the more similar the two vectors are, i.e., the more similar the two faces are. After calculating all N cosine similarities, the maximum similarity is taken. and corresponding user identification This maximum similarity reflects the degree of similarity between the queried face and the most similar face in the feature database. If the match is greater than or equal to the face matching threshold, the match is considered successful, and the queried face is considered to belong to user ID [user ID]. For authorized users, record the user ID for subsequent permission queries; if If the face is less than the face matching threshold, the matching is deemed to have failed. The system assumes that the queried face does not belong to any registered authorized user and may be a stranger or attacker. The system directly refuses to unlock the device and does not initiate the subsequent multimodal liveness detection process, thereby effectively reducing the consumption of computing resources and the system response time.
[0049] As an implementation method, the face matching threshold is a key parameter that directly affects the system's security and convenience. Setting the threshold too high will increase the false rejection rate, as authorized users may have altered features due to changes in facial expressions, makeup, or wearing glasses, resulting in a similarity score below the threshold and being denied unlocking, thus impacting user experience. Conversely, setting the threshold too low will increase the false recognition rate, as strangers or attackers may be mistakenly identified as authorized users due to facial features similar to an authorized user, leading to false unlocking and posing a security risk. In this embodiment, the face matching threshold is set to 0.65, providing a good user experience while ensuring security. It should be noted that the face matching threshold is not fixed but adaptively adjusted based on environmental conditions and user history.
[0050] Through the above steps, feature extraction and matching of preprocessed face images are achieved. A 512-dimensional face feature vector is extracted using a deep learning model. Cosine similarity calculation is used to achieve rapid matching with the authorized face feature library. By setting a reasonable face matching threshold, the system's security and convenience are balanced. A two-stage verification strategy reduces computational resource consumption while ensuring security.
[0051] Furthermore, to enhance system security, this embodiment introduces a liveness detection mechanism. Even after successful facial feature matching, multimodal liveness detection is performed to verify that the queried face is a real, live person and not a photo, video, mask, or other attack sample. Only if the liveness detection also passes is the final verification successful, and the unlocking operation is executed. This two-stage verification strategy—static feature matching followed by liveness detection—ensures system security while effectively reducing computational resource consumption and system response time by skipping the liveness detection process when matching fails.
[0052] Step S2: If the matching fails, the unlocking is refused; otherwise, a liveness detection is performed. This involves: simultaneously acquiring temporal RGB images, depth images, near-infrared images, and thermal images of the human face; extracting multimodal dynamic features of the face through a temporal convolutional network; outputting a liveness confidence score after a multimodal consistency test; and combining this with a liveness detection threshold that is adaptively adjusted based on the environment to determine whether the face unlock verification is successful.
[0053] In step S201, if the matching is successful, the temporal RGB image, depth image, near-infrared image and thermal image of the human face are acquired simultaneously, and the multimodal dynamic features of the face are extracted through a temporal convolutional network.
[0054] In this embodiment, if the static facial feature matching is successful, the multimodal sensors are activated to synchronously acquire temporal data of the human face for subsequent liveness detection. The multimodal sensors include an RGB image sensor of a depth camera, a depth image sensor, a near-infrared camera, and a thermal imaging camera. These sensors can acquire facial information from different physical dimensions, providing rich data support for liveness detection.
[0055] Among them, temporal RGB images refer to a sequence of multiple color images continuously acquired over a period of time, capable of capturing dynamic changes in the face, including micro-expressions, blinking, mouth movements, and head posture changes. In this embodiment, the RGB image sensor continuously acquires images for 3 seconds at a frame rate of 30 frames per second, resulting in a total of 90 RGB images; Temporal depth images refer to a sequence of multiple depth images acquired continuously over a period of time. They can capture changes in the three-dimensional shape of the face, including depth changes in the eyelids during blinking, subtle facial movements during breathing, and changes in depth distribution caused by changes in head posture. In this embodiment, the depth image sensor continuously acquires images at a rate of 30 frames per second for 3 seconds, resulting in a total of 90 depth images. Temporal near-infrared imaging refers to a sequence of multiple near-infrared images acquired continuously over a period of time. Near-infrared wavelength is 850 nanometers, a wavelength that can penetrate the skin's surface, reflecting the skin's internal structure and blood flow. The near-infrared reflectance of a real human face differs significantly from that of a photograph or a face displayed on a screen. In this embodiment, the near-infrared camera continuously acquires images at a frame rate of 30 frames per second for 3 seconds, resulting in a total of 90 near-infrared images. Temporal thermal imaging refers to a sequence of multiple thermal images acquired continuously over a period of time. Thermal imaging cameras can detect the temperature distribution on a face. The temperature distribution of a real human face exhibits a specific pattern: areas such as the forehead, tip of the nose, and cheeks have higher temperatures, while areas such as the eye sockets and nostrils have lower temperatures. Furthermore, the temperature distribution dynamically changes with respiration and blood flow. Attack samples such as photos, videos, and masks cannot simulate this realistic temperature distribution and dynamic changes. In this embodiment, the thermal imaging camera continuously acquires images at a frame rate of 9 frames per second for 3 seconds, resulting in a total of 27 thermal images.
[0056] Preferably, to ensure time synchronization of multimodal data, all sensors will be clock synchronized before the acquisition begins, using hardware trigger signals or network time protocols to ensure that the acquisition time error of each sensor does not exceed 10 milliseconds.
[0057] After acquisition, four temporal image sequences were obtained: 90 frames of RGB images, 90 frames of depth images, 90 frames of near-infrared images, and 27 frames of thermal imaging images. Linear interpolation of the thermal imaging images also yielded 90 frames. These image sequence data were transmitted to the vehicle domain controller via the vehicle Ethernet. The vehicle domain controller first performed preprocessing on the received image sequences, including image registration, size normalization, and data augmentation. After preprocessing, the four temporal image sequences were input into a temporal convolutional network (TCN) to extract multimodal dynamic features of the face. This TCN is a convolutional neural network architecture specifically designed for processing temporal data. Through one-dimensional convolutional operations, it extracts features along the time dimension, effectively capturing long-term dependencies and dynamic change patterns in temporal data.
[0058] In this embodiment, the temporal convolutional network comprises two branches: the first branch processes RGB and near-infrared images, and the second branch processes depth and thermal imaging images. Each branch first extracts spatial features from each frame of the image through the MobileNetV2 backbone network. MobileNetV2 is a lightweight convolutional neural network that significantly reduces computational cost while maintaining high recognition accuracy through inverted residual structures and linear bottleneck layers.
[0059] Then, the feature matrix is input into the temporal convolutional layer. This layer uses a one-dimensional convolutional kernel to perform convolution along the time dimension, with a kernel size of 3 and a stride of 1, ensuring the output sequence has the same length as the input sequence. The temporal convolutional layer can capture feature changes between adjacent frames and extract temporal dynamic patterns. In this embodiment, the temporal convolutional network contains four layers, each with 256, 512, 512, and 256 channels respectively. Through layer-by-layer convolution, the network can capture temporal dependencies from short-term to long-term.
[0060] Finally, through a global average pooling layer and a fully connected layer, the temporal feature matrix is mapped to a fixed-length feature vector, which is the facial multimodal dynamic feature with a dimension of 512. It contains information on dynamic changes in the face, such as micro-expression changes, blinking, head posture changes, and facial movements caused by breathing.
[0061] Through the above steps, synchronous acquisition and dynamic feature extraction of multimodal temporal data can be achieved. By fusing multimodal data such as RGB images, depth images, near-infrared images and thermal imaging, rich dynamic information of the face is captured. Dynamic features containing temporal dependencies are extracted through temporal convolutional networks, providing a data foundation for subsequent multimodal consistency testing and temporal causality verification.
[0062] Step S202: Perform a multimodal consistency test on the facial multimodal dynamic features. Based on the probabilities obtained from the multimodal consistency test, calculate the liveness confidence score through weighted fusion.
[0063] In this embodiment, multimodal consistency testing is the core step in liveness detection. By performing cross-validation and temporal analysis on multimodal data, it determines whether the queried face is a real live person. Specifically, the multimodal consistency test includes detecting the spatial alignment error between the texture and depth image of the RGB face image, the correlation between the thermal imaging temperature gradient and the skin color distribution of the RGB face image, and the frequency domain consistency between the near-infrared reflectance and the texture of the RGB face image.
[0064] (1) Detection of spatial alignment error between RGB face image texture and depth image.
[0065] A real human face's RGB image and depth image should have a high degree of spatial consistency, meaning that facial key points in the RGB image and their corresponding key points in the depth image should be located in the same spatial position. However, planar attack samples such as photo attacks and video attacks lack realistic depth information; their depth images are either completely flat with no depth variation or their depth information does not match the RGB image, resulting in large spatial alignment errors. Therefore, it is necessary to detect whether the facial texture in the RGB image and the facial depth information in the depth image are spatially aligned. The detection process is as follows: First, a face detection algorithm is used to detect face bounding boxes and multiple key feature points in both RGB and depth images. The face detection algorithm employs the aforementioned MTCNN algorithm, running on both RGB and depth images respectively, outputting face bounding boxes and the coordinates of five key points, including the center of the left eye, the center of the right eye, the tip of the nose, the left corner of the mouth, and the right corner of the mouth.
[0066] Secondly, the difference in pixel coordinates of the corresponding key points in the RGB image plane and the depth image plane is calculated to obtain the spatial alignment error vector. That is, the coordinates of five key points in the RGB image and the corresponding key point coordinates in the depth image are represented by a vector. The spatial alignment error vector is represented by the difference between the two vectors, which reflects the positional deviation of the facial key points in the RGB image and the depth image.
[0067] Next, the L2 norm of the spatial alignment error vector is calculated, which reflects the overall magnitude of the spatial alignment error.
[0068] Finally, the L2 norm is used to determine whether it is a real face. If the L2 norm is less than a set pixel threshold (3 pixels in this embodiment), it is determined to be a natural registration error of a real face. This is because, due to factors such as sensor calibration errors and image registration errors, there will be a small range of alignment errors between the RGB image and the depth image of a real face, but this error is usually less than 3 pixels. If the L2 norm is greater than the set pixel threshold (10 pixels in this embodiment), it is determined to be a lack of depth information caused by a planar attack. The depth image and RGB image of a planar attack sample are severely mismatched, and the alignment error is usually greater than 10 pixels. Based on the L2 norm value e, the alignment consistency probability is output. , can be represented as: ; The alignment consistency probability ranges from 0 to 1. The closer the value is to 1, the higher the alignment consistency and the more likely it is to be a real human face.
[0069] (2) Correlation detection between thermal imaging temperature gradient and skin color distribution in RGB face images.
[0070] In a realistic human face, the temperature distribution and skin color distribution should exhibit a high spatial correlation. Facial skin areas should have higher temperatures and more pronounced temperature gradients, while non-skin areas such as hair and background should have lower temperatures and smaller temperature gradients. However, samples from photo attacks, video attacks, and mask attacks lack a realistic temperature distribution, and the temperature gradient in their thermal imaging images is uncorrelated with the skin color distribution in the RGB images. Therefore, it is necessary to detect whether the temperature gradient distribution in thermal imaging images and the skin color distribution in RGB images have a spatial correlation. The detection process is as follows: First, facial skin regions are extracted from the RGB image using a skin color model in the YCrCb color space, where Y represents luminance and Cr and Cb represent chroma. Human skin chroma values have a relatively stable distribution range in the YCrCb space; by setting thresholds for Cr and Cb, skin regions can be effectively extracted. In this embodiment, the extraction conditions for skin regions are: Cr values between 133 and 173, and Cb values between 77 and 127. Pixels meeting these conditions are identified as skin pixels. After extracting the skin regions, a spatial distribution density map of the skin regions is calculated. Each pixel value in the density map represents the probability that the location is skin, with a value ranging from 0 to 1.
[0071] Secondly, the facial temperature gradient field is calculated in the thermal imaging image. The temperature gradient field reflects the rate of temperature change in space. This is achieved by processing the thermal imaging image with the Sobel gradient operator (which includes gradient operators in the horizontal and vertical directions), then calculating the rate of temperature change of each pixel in the horizontal and vertical directions through convolution operations, and finally calculating the gradient magnitude, which reflects the severity of temperature change at that location.
[0072] Next, high gradient regions with temperature gradient amplitudes greater than a set temperature gradient threshold are extracted; in this embodiment, this threshold is set to 0.5 degrees Celsius per centimeter. High gradient regions correspond to areas of the face with drastic temperature changes, such as the tip of the nose, cheeks, and forehead. These areas experience larger temperature gradients due to blood flow and respiratory activity.
[0073] Next, the spatial correlation coefficient between the skin density map and the temperature gradient field is calculated. The spatial correlation coefficient is calculated using the Pearson correlation coefficient, which is obtained by calculating the covariance of the skin density map and the temperature gradient field, dividing it by the standard deviation of the skin density map multiplied by the standard deviation of the temperature gradient field. The correlation coefficient ranges from -1 to +1; the closer the value is to 1, the stronger the spatial correlation between the two distributions.
[0074] Finally, the correlation coefficient is used to determine whether the face is a real person. A real face should have a correlation coefficient greater than 0.7, indicating a high degree of overlap between the skin area and the high temperature gradient area. Samples from photo attacks, video attacks, or mask attacks typically have correlation coefficients less than 0.3, indicating that the skin area and the temperature gradient area are not correlated. Based on the correlation coefficient, the temperature correlation probability is output. The temperature correlation probability ranges from 0 to 1. The closer the value is to 1, the higher the temperature correlation and the more likely it is to be a real human face.
[0075] (3) Frequency domain consistency detection of near-infrared reflectance and RGB face image texture.
[0076] The near-infrared reflectance and visible light texture of a real human face should exhibit high consistency in the frequency domain, as they both reflect the skin's inherent texture structure, such as wrinkles, blemishes, and pores. However, due to the pixelation effect of the display, the near-infrared and RGB images of screen attack samples show significant differences in the frequency domain. The screen-displayed image exhibits obvious pixel grid textures in the high-frequency range, leading to reduced frequency domain consistency. Therefore, it is necessary to detect whether the reflectance characteristics of the near-infrared image and the texture characteristics of the RGB image are consistent in the frequency domain. The detection process is as follows: First, two-dimensional discrete Fourier transforms are performed on the near-infrared image and the RGB grayscale image, respectively. The Fourier transform can convert an image in the spatial domain to the frequency domain, and the frequency domain image reflects the distribution of different frequency components in the image.
[0077] Secondly, the amplitude spectrum of the frequency domain image is calculated. The frequency domain image is in complex form, containing real and imaginary parts. The amplitude spectrum can be calculated by adding the square of the real part to the square of the imaginary part and then taking the square root. This amplitude spectrum reflects the intensity of each frequency component.
[0078] Next, the mid-frequency component of the frequency band from 0.1 to 0.5 periods per pixel in the frequency domain amplitude spectrum is extracted. This frequency band corresponds to the main frequency range of facial skin texture features. Low-frequency components correspond to overall illumination changes and facial contours, high-frequency components correspond to noise and minor pixel changes, and mid-frequency components correspond to skin texture features. The mid-frequency component is extracted using a bandpass filter with a lower limit frequency of 0.1 periods per pixel and an upper limit frequency of 0.5 periods per pixel.
[0079] Then, the normalized cross-correlation coefficient between the near-infrared frequency domain amplitude spectrum and the visible light frequency domain amplitude spectrum is calculated. The normalized cross-correlation coefficient ranges from 0 to 1, and the closer the value is to 1, the more similar the two are in the frequency domain.
[0080] Finally, the normalized cross-correlation coefficient is used to determine whether it is a real face. A real face should have a cross-correlation coefficient greater than 0.75, indicating a high degree of consistency between the near-infrared and visible light images in the frequency domain. Screen attack samples typically have a cross-correlation coefficient less than 0.5 due to display pixelation effects. Based on the cross-correlation coefficient, the frequency domain consistency probability is output. The frequency domain consistency probability ranges from 0 to 1. The closer the value is to 1, the higher the frequency domain consistency and the more likely it is to be a real human face.
[0081] The above three testing strategies yield corresponding probability values, which are the spatial alignment consistency probabilities. Temperature correlation probability Frequency domain consistency probability These values reflect the consistency and causal relationship of multimodal data in different dimensions. In order to obtain a comprehensive liveness detection result, these probability values need to be weighted and fused to calculate the liveness confidence score L.
[0082] As one implementation method, the weighting coefficients for calculating liveness confidence can be determined based on the importance of defense against different types of adversarial attacks. Spatial alignment detection has the highest defense weight against planar attacks such as photo attacks and video attacks, and is set to 0.25. Temperature correlation detection and causal consistency verification are effective against mask attacks and deepfake attacks, and are set to 0.2. Frequency domain consistency detection is effective against screen attacks, and is set to 0.15. Breathing synchronization verification is effective against static attack samples, and is set to 0.2. Through weighted fusion, the liveness confidence score L comprehensively reflects the probability of the queried face being live, with a value ranging from 0 to 1. The closer the value is to 1, the more likely it is to be a real live person, and the closer the value is to 0, the more likely it is to be an attack sample.
[0083] Through the above steps, consistency verification of multimodal dynamic facial features is achieved. Cross-validation and temporal analysis of multimodal data can be performed from different dimensions, effectively improving the robustness and accuracy of liveness detection. It can defend against various attack methods such as photo attacks, video attacks, mask attacks, screen attacks, and deepfake attacks. A comprehensive liveness confidence score is obtained through weighted fusion, providing a reliable basis for subsequent comprehensive decision-making.
[0084] Step S203: Based on the ambient light intensity monitored in real time by the ambient light sensor, the liveness detection threshold and face matching threshold are adaptively adjusted. The liveness confidence score is compared with the adjusted liveness detection threshold to determine whether the face unlock verification is successful.
[0085] In this embodiment, environmental adaptive threshold adjustment is a key technology for improving the system's environmental adaptability. Because the environmental conditions in vehicle face unlock scenarios are complex and varied, potentially facing low light, strong light, and other lighting conditions, different lighting conditions can significantly impact image quality, feature extraction accuracy, and liveness detection performance. To ensure the system maintains high recognition accuracy and user experience under various environmental conditions, this embodiment introduces an environmental adaptive threshold adjustment strategy, dynamically adjusting the liveness detection threshold and face matching threshold based on ambient light levels. The process is as follows: First, the ambient light level of the vehicle's surroundings is monitored in real time using an ambient light sensor. This sensor is deployed externally, typically on the roof or near the door mirrors, and detects the ambient light intensity. The sensor collects ambient light data once per second and transmits the data to the vehicle's domain controller via the CAN bus.
[0086] Secondly, based on a comparison of the ambient light intensity with preset threshold values for each light intensity level, the current environment is determined to be either a low-light environment, a bright-light environment, or a normal-light environment. In this embodiment, ambient light intensity is divided into three levels: low-light environment corresponds to an illuminance of less than 10 lux, bright-light environment corresponds to an illuminance of greater than 1000 lux, and normal-light environment corresponds to an illuminance between 10 and 1000 lux. These three levels are determined based on the human eye's perception of light intensity and the performance characteristics of the image sensor. Low-light environment corresponds to indoor lighting conditions without artificial light or outdoor lighting conditions at night, resulting in poor image quality and high noise. Bright-light environment corresponds to direct sunlight at midday, where images may be overexposed and details may be lost. Normal-light environment corresponds to indoor lighting conditions or lighting conditions on a cloudy day, resulting in good image quality.
[0087] Based on this, corresponding liveness detection thresholds and face matching thresholds are determined according to different types of environments. In low-light environments, due to poor image quality and decreased accuracy of feature extraction, the confidence level of liveness detection may be low, and the similarity of face matching may also be low. If the thresholds used in normal lighting environments are still applied, the false rejection rate will increase, and authorized users may be denied unlocking due to poor lighting conditions, affecting user experience. Therefore, in low-light environments, the thresholds need to be appropriately reduced, such as lowering the liveness detection threshold from 0.85 to 0.80 and the face matching threshold from 0.65 to 0.60, to improve the pass rate and ensure user experience.
[0088] In strong light environments, although the overall image brightness is high, overexposure may occur, resulting in loss of facial details. Furthermore, strong light makes photo or video attacks easier to execute because attackers can use the bright light to mask flaws in the attack sample. Therefore, in strong light environments, the liveness detection threshold needs to be appropriately increased, such as from 0.85 to 0.90, to improve security and prevent attack samples from passing verification. The face matching threshold remains unchanged at 0.65 because strong light has a relatively small impact on face feature matching; its main effect is on liveness detection.
[0089] Under normal lighting conditions, the image quality is good, and the effects of feature extraction and liveness detection are both good. The thresholds remain at their default values, with a liveness detection threshold of 0.85 and a face matching threshold of 0.65, which ensures both security and user experience.
[0090] Finally, the liveness confidence score L is compared with the adjusted liveness detection threshold to determine whether the face unlock verification passes. If the liveness confidence score L is greater than or equal to the adjusted liveness detection threshold, the liveness detection is considered successful, and the queried face is considered a real live person rather than an attack sample, proceeding to the subsequent comprehensive decision-making steps. If the liveness confidence score L is less than the adjusted liveness detection threshold, the liveness detection is considered unsuccessful, and the queried face is considered potentially an attack sample. The system refuses to unlock and records a failure log, which includes information such as the failure time, failure reason, liveness confidence score, and ambient light level, for subsequent security auditing and system optimization.
[0091] As another implementation method, the environmental adaptive threshold adjustment strategy in this embodiment, in addition to considering the light intensity dimension, can also be extended to other environmental dimensions in practical applications, such as the vehicle vibration amplitude dimension and the user's historical success rate dimension, which will not be elaborated here.
[0092] Through a multi-dimensional environment-adaptive threshold adjustment strategy, the system can dynamically adjust the verification threshold based on real-time environmental conditions and user behavior, maximizing user experience while ensuring security, and enabling the system to adapt to various complex and ever-changing application scenarios.
[0093] This step enables adaptive threshold adjustment based on ambient light intensity. By dynamically adjusting the liveness detection threshold and face matching threshold, the system can adapt to different lighting conditions such as low light and strong light. This improves user experience, reduces false rejection rate, and enhances the system's environmental adaptability and robustness while ensuring safety.
[0094] Step S3: If successful, generate the corresponding vehicle control command based on the verified user's identity and permission level to drive the vehicle to perform the corresponding unlocking operation and function pre-start operation.
[0095] Step S301: If the verification is successful, query the corresponding permission level according to the matched user identity identifier. The permission level includes main user permission, family member permission, temporary user permission and chauffeur mode permission. Generate the corresponding vehicle control command according to the permission level to drive the vehicle to perform the corresponding unlocking operation and function pre-start operation.
[0096] In this embodiment, different user identities have different vehicle usage needs and security risks. By introducing a four-level permission system, differentiated vehicle control permissions are provided for different user identities, which not only ensures convenient use for the main user and family members, but also imposes security restrictions on temporary users and chauffeur mode, thereby improving the flexibility and security of the system.
[0097] Specifically, the four-level permission system includes primary user permissions, family member permissions, temporary user permissions, and designated driver mode permissions, among which: The primary user has the highest level of access, corresponding to the vehicle owner or primary user. Primary user permissions include unlocking all vehicle doors, starting the vehicle, adjusting all personalized settings such as seat position, rearview mirror angle, air conditioning temperature, and audio volume, modifying vehicle system settings such as security settings, network settings, and user management. These permissions are permanent and have no expiration date. When the primary user uses the face unlock function for the first time, face registration is required. The system will collect the primary user's facial image, extract facial feature vectors, store them in the authorized facial feature database, and mark them as having primary user permissions. The primary user can also add or delete other users and modify their permission levels through the in-vehicle system interface.
[0098] Family member permissions are the second highest level of access, corresponding to the vehicle owner's family members such as spouse and children. Family member permissions allow actions such as unlocking all vehicle doors, starting the vehicle, and adjusting personalized settings such as seat position, rearview mirror angle, air conditioning temperature, and audio volume. However, they cannot modify vehicle system settings such as security settings, network settings, or user management. These permissions are permanent and have no expiration date. The main difference between family member permissions and primary user permissions is the inability to modify system settings. This is to prevent accidental system configuration errors caused by family members or to prevent family members from deleting the primary user's permissions. Family members also need to register their faces when using the face unlock function for the first time. The primary user adds family members through the in-vehicle system interface. The system collects the family member's facial image, extracts facial feature vectors, stores them in the authorized facial feature database, and marks them as family member permissions.
[0099] Temporary user permissions are restricted permissions, suitable for users who temporarily use the vehicle, such as friends or guests. Temporary user permissions only allow unlocking the driver's side door; other doors cannot be unlocked, the vehicle cannot be started, and personalized settings cannot be adjusted. These permissions are time-limited and expire after 2 hours by default. The purpose of temporary user permissions is to allow temporary users to briefly enter the vehicle when the owner is not present, such as while waiting for the owner to return in a parking lot. They can enter the car to rest, but cannot start the vehicle or perform other operations, preventing vehicle theft or accidental operation. Adding temporary user permissions requires the primary user or a family member to do so through the in-vehicle system interface. The system will collect the temporary user's facial image, extract facial feature vectors, store them in the authorized facial feature database, and mark the user as having temporary user permissions. The system will also record the effective time of the permissions. After 2 hours, the system will automatically delete the temporary user's feature vectors, and the permissions will expire.
[0100] The designated driver mode permission is a special restricted level of permission, corresponding to users in special scenarios such as designated drivers. The operations that can be performed in designated driver mode include unlocking the driver's side door and starting the vehicle, but are subject to several restrictions, including a maximum speed limit of 60 km / h, prohibition of access to personal data in the vehicle system such as contacts, navigation history, and multimedia files, and automatic recording of driving trajectory including GPS location, speed, and time. This permission is time-limited and expires after 2 hours by default. The design purpose of designated driver mode permission is to allow designated drivers to drive the vehicle when the owner is under the influence of alcohol or fatigued, but by limiting the maximum speed, prohibiting access to personal data, and recording driving trajectory, it protects the owner's privacy and vehicle safety, preventing designated drivers from speeding, stealing personal information, or deviating from the planned route. Adding designated driver mode permission requires the main user to do so through the vehicle system interface. The system will collect the designated driver's facial image, extract facial feature vectors, store them in the authorized facial feature database, and mark it as designated driver mode permission, while recording the effective time of the permission. After 2 hours, the system will automatically delete the designated driver's feature vectors, and the permission will expire.
[0101] In step S103, if the facial feature match is successful, the system will record the matched user's identity identifier. This user identifier uniquely corresponds to a specific user in the authorized facial feature database. In this step, the system uses the user identifier... The system queries the corresponding permission level. This query is performed in the authorized facial feature database, which stores each user's feature vector, user identifier, and permission level information. The user identifier allows for quick retrieval of the corresponding permission level. After retrieving the permission level, the system generates the corresponding vehicle control commands based on that level. Vehicle control commands are a set of instructions used to control vehicle actuators, including door lock control commands, vehicle start control commands, personalized setting adjustment commands, and function pre-start commands.
[0102] As another implementation method, during the generation of vehicle control commands, the system also performs comprehensive decision-making, taking into account multiple factors such as face matching similarity, liveness confidence score, and thermal imaging verification results to determine the confidence level of the verification and adopt different control strategies based on the confidence level. This comprehensive decision-making process is as follows: if the face matching similarity is greater than or equal to the face matching threshold, the liveness confidence score L is greater than or equal to the liveness detection threshold, and the thermal imaging verification is passed, then it is determined to be a high-confidence pass. The system directly queries the permission level based on the matched user identity and generates the corresponding vehicle control command, executing the complete unlocking operation and function pre-start operation.
[0103] Through the above steps, differentiated vehicle control based on a four-level permission system can be achieved. By querying the permission level according to the user's identity, corresponding vehicle control commands are generated, which not only ensures convenient use for the main user and family members, but also imposes safety restrictions on temporary users and chauffeur mode. Through comprehensive decision-making logic, security and user experience are balanced, and the system's flexibility, security and intelligence level are improved.
[0104] Example 2 This embodiment provides an intelligent vehicle face unlocking system based on multimodal fusion, specifically including: The human detection module includes a passive infrared human sensor, which is used to detect human signals around the vehicle in real time when the vehicle is in a state of waiting to be unlocked. The image acquisition and preprocessing module is used to acquire and preprocess a static frontal face image of a person approaching the body if a person is detected to be approaching; and to simultaneously acquire temporal RGB images, depth images, near-infrared images and thermal images of the human face. The face feature matching module is used to match the preprocessed face image with the preset authorized face image. If the match fails, the unlocking will be refused; otherwise, a liveness detection will be performed. The liveness detection module is used to extract multimodal dynamic features of the face through a temporal convolutional network, and then output a liveness confidence score through a multimodal consistency test. Combined with the liveness detection threshold that is adaptively adjusted based on the environment, it determines whether the face unlock verification is successful. The access control module is used to generate corresponding vehicle control commands based on the verified user's identity and access level, so as to drive the vehicle to perform corresponding unlocking operations and function pre-start operations.
[0105] Example 3 This embodiment provides a vehicle for executing vehicle control commands corresponding to the above-described intelligent vehicle face unlocking method based on multimodal fusion, or includes the above-described intelligent vehicle face unlocking system based on multimodal fusion.
[0106] The steps and methods involved in Examples 2 and 3 above correspond to those in Example 1. For specific implementation details, please refer to the relevant description section of Example 1.
[0107] The above description is only a preferred embodiment of the present invention. Although the specific implementation of the present invention has been described in conjunction with the accompanying drawings, it is not intended to limit the scope of protection of the present invention. Those skilled in the art should understand that, based on the technical solution of the present invention, various modifications or variations that can be made by those skilled in the art without creative effort are still within the scope of protection of the present invention.
Claims
1. A method for unlocking intelligent vehicles using facial recognition based on multimodal fusion, characterized in that, include: When the vehicle is in the unlocking state, it detects human signals around the vehicle in real time. If a human body is detected approaching, it collects a static frontal face image of the approaching human body and performs feature matching with a preset authorized face image. If the match fails, the unlocking is refused; otherwise, a liveness detection is performed. This involves simultaneously acquiring temporal RGB images, depth images, near-infrared images, and thermal images of the human face. Multimodal dynamic features of the face are extracted through a temporal convolutional network. A liveness confidence score is then output after a multimodal consistency test. Combined with a liveness detection threshold that is adaptively adjusted based on the environment, the face unlock verification is determined to be successful. If successful, the corresponding vehicle control command will be generated based on the verified user's identity and permission level to drive the vehicle to perform the corresponding unlocking operation and function pre-start operation.
2. The intelligent vehicle face unlocking method based on multimodal fusion as described in claim 1, characterized in that, The process of static facial feature matching is as follows: A static frontal image of a human face close to the body is acquired and preprocessed; the preprocessing includes global and local contrast normalization. A facial feature extraction model is used to extract key feature points from the preprocessed facial image. These points are then compared with the corresponding key feature points of standard facial images pre-stored in the authorized facial feature database. The matching result is determined based on the comparison, including: Using the extracted key feature points of the face image as the query vector, the key feature vectors of all registered users are read from the authorized face feature library stored in the vehicle. The cosine similarity between the query vector and each feature vector in the library is calculated. If the maximum similarity is greater than or equal to the face matching threshold, the match is considered successful. The maximum similarity and the corresponding user ID are taken. Otherwise, the match is considered unsuccessful.
3. The intelligent vehicle face unlocking method based on multimodal fusion as described in claim 1, characterized in that, The multimodal consistency test and temporal causality verification include: The study investigates the spatial alignment error between the texture and depth images of RGB face images, the correlation between thermal imaging temperature gradient and skin color distribution in RGB face images, and the frequency domain consistency between near-infrared reflectance and texture in RGB face images. Based on the probabilities obtained from the tests, the liveness confidence score is calculated through weighted fusion.
4. The intelligent vehicle face unlocking method based on multimodal fusion as described in claim 3, characterized in that, Spatial alignment error detection between RGB face image texture and depth image, including: A face detection algorithm is used to detect face bounding boxes and multiple key feature points in RGB face images and depth images, respectively; The difference in pixel coordinates of the corresponding key points in the RGB image plane and the depth image plane is calculated to obtain the spatial alignment error vector; Calculate the L2 norm of the spatial alignment error vector. If the L2 norm is less than a set pixel threshold, it is determined to be a natural registration error of a real face; otherwise, it is determined to be a lack of depth information caused by a planar attack. Output the alignment consistency probability. .
5. The intelligent vehicle face unlocking method based on multimodal fusion as described in claim 3, characterized in that, Correlation detection between thermal imaging temperature gradient and skin color distribution in RGB face images, including: In RGB images, facial skin regions are extracted using a skin color model in the YCrCb color space, and the spatial distribution density map of the skin regions is calculated. Calculate the facial temperature gradient field in thermal imaging images and extract high gradient regions where the temperature gradient amplitude is greater than a set temperature gradient threshold. Calculate the spatial correlation coefficient between the skin density map and the temperature gradient field, and combine it with the correlation coefficient of a real face to output the temperature correlation probability. .
6. The intelligent vehicle face unlocking method based on multimodal fusion as described in claim 3, characterized in that, Frequency domain consistency detection of near-infrared reflectance and RGB face image texture includes: Two-dimensional discrete Fourier transforms were performed on the near-infrared image and the RGB grayscale image respectively to obtain the frequency domain amplitude spectrum; Extract the mid-frequency component of the 0.1-0.5 period / pixel frequency band in the frequency domain amplitude spectrum, which corresponds to the facial skin texture features; Calculate the normalized cross-correlation coefficients of the near-infrared and visible light frequency domain amplitude spectra, and combine them with the cross-correlation coefficients of real faces to output the frequency domain consistency probability. .
7. The intelligent vehicle face unlocking method based on multimodal fusion as described in claim 1, characterized in that, The liveness detection threshold is adaptively adjusted according to the environment, including: Using an ambient light sensor, the ambient light level of the current vehicle's surroundings is monitored in real time; Based on the comparison between the ambient light intensity and the preset threshold values for each light intensity level, the system determines whether the current environment is a dark environment, a bright environment, or a normal lighting environment, and then determines the corresponding liveness detection threshold based on the different types of environments.
8. The intelligent vehicle face unlocking method based on multimodal fusion as described in claim 1, characterized in that, User access levels are divided into four levels, including: The primary user has the privileges to unlock all vehicle doors, start the vehicle, adjust all personalized settings, and modify vehicle system settings. These privileges are permanently valid. Family member permissions allow unlocking all vehicle doors, starting the vehicle, and adjusting personalized settings, but not modifying system settings; these permissions are permanent. Temporary user privileges allow unlocking the driver's side door, but the vehicle cannot be started. These privileges are time-limited and expire after a default time. The designated driver mode grants permissions to unlock the driver's side door and start the vehicle, but limits the maximum speed, prohibits access to personal data, and automatically records driving routes. These permissions are time-limited and expire after a default time setting.
9. A smart vehicle face unlocking system based on multimodal fusion, characterized in that, include: The human detection module includes a passive infrared human sensor, which is used to detect human signals around the vehicle in real time when the vehicle is in a state of waiting to be unlocked. The image acquisition and preprocessing module is used to acquire and preprocess a static frontal face image of a person approaching the human body if a human body is detected approaching. It also simultaneously acquires temporal RGB images, depth images, near-infrared images, and thermal images of the human face; The face feature matching module is used to match the preprocessed face image with the preset authorized face image. If the match fails, the unlocking will be refused; otherwise, a liveness detection will be performed. The liveness detection module is used to extract multimodal dynamic features of the face through a temporal convolutional network, and then output a liveness confidence score through a multimodal consistency test. Combined with the liveness detection threshold that is adaptively adjusted based on the environment, it determines whether the face unlock verification is successful. The access control module is used to generate corresponding vehicle control commands based on the verified user's identity and access level, so as to drive the vehicle to perform corresponding unlocking operations and function pre-start operations.
10. A vehicle, characterized in that, Used to execute vehicle control commands corresponding to the intelligent vehicle face unlocking method based on multimodal fusion as described in any one of claims 1-8, or including the intelligent vehicle face unlocking system based on multimodal fusion as described in claim 9.