Photographing method and system based on gaze depth and scene verification, vehicle and medium
Patent Information
- Application Number
- CN202610541685.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-04-22
- Publication Date
- 2026-08-18
AI Technical Summary
[0003]然而,现有技术方案仍存在显著不足:一方面,传统DMS及车载监控系统的核心目标聚焦于安全预警,其注视检测仅能判断视线方向或是否分神,缺乏对驾驶员注视距离或深度的量化感知能力;另一方面,这些系统未建立与环境场景的语义关联机制,无法区分驾驶员是关注导航标识、风景还是潜在危险源,导致无法理解注视背后的真实意图
首先,在意图识别层面,不仅解析驾驶员的注视方向,更关键地引入基于双眼视差计算的注视深度距离,结合场景语义分割生成的三维空间标签,将二维视线精确映射为三维空间中的具体目标。只有当该目标属于高价值记录类别时,才判定满足第一触发条件,从而有效过滤无意义的随机扫视,确保捕捉的是驾驶员真正关注的有价值内容。其次,在安全决策层面,利用车辆状态数据动态评估驾驶任务的认知负荷与操作负荷。仅当两项负荷指数均低于安全阈值时,才判定满足第二触发条件。这一机制强制系统优先保障驾驶安全,避免在驾驶员精力紧张或操作复杂时进行自动拍摄干扰。最终,只有当有价值的注视目标与安全的驾驶状态同时满足时,系统才生成拍摄指令并控制成像器。本方案既实现了基于真实意图的智能触发,又确保了拍摄行为不会危及行车安全,达成了精准、智能且安全的影像拍摄效果。
Smart Images

Figure CN122601968A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of automotive engineering technology, specifically to a shooting method, system, vehicle, and medium based on depth of gaze and scene verification. Background Technology
[0002] With the rapid development of intelligent vehicle technology, in-vehicle image acquisition systems have evolved from simple dashcams into multifunctional sensing terminals integrating safety protection, interactive entertainment, and scene recording. Currently, in-vehicle camera technology is mainly applied in areas such as Advanced Driver Assistance Systems (ADAS), Driver Monitoring Systems (DMS), and dashcams. Among these, DMS systems use in-vehicle cameras to capture real-time images of the driver's face and utilize image processing algorithms to identify behavioral characteristics such as closed eyes, yawning, and head-down posture, enabling early warning functions for safety risks such as fatigued driving and inattention. Such systems have been widely adopted in industry regulations and have become an important means of improving active safety performance.
[0003] However, existing technical solutions still have significant shortcomings: On the one hand, the core objective of traditional DMS and vehicle monitoring systems focuses on safety warnings, and their gaze detection can only determine the direction of the gaze or whether the driver is distracted, lacking the ability to quantitatively perceive the driver's gaze distance or depth. On the other hand, these systems have not established a semantic association mechanism with the environmental scene, and cannot distinguish whether the driver is paying attention to navigation signs, scenery, or potential hazards, resulting in an inability to understand the true intention behind the gaze. In addition, during vehicle operation, dynamic factors such as road vibration and changes in lighting further degrade image quality, and existing systems fail to comprehensively consider the collaborative judgment of driver intention and scene value while meeting driving safety load constraints. Therefore, without gaze distance perception and scene semantic verification mechanisms, existing solutions cannot accurately identify and trigger valuable shooting behavior driven by the driver's gaze depth intention under dynamic driving safety load constraints. Summary of the Invention
[0004] This invention provides a shooting method, system, vehicle, and medium based on gaze depth and scene verification, which can accurately identify and capture moments when the driver has real shooting value, and realize intelligent and precise shooting of in-vehicle and out-of-vehicle images.
[0005] This invention provides a shooting method based on gaze depth and scene verification, the method comprising: Acquire the driver's facial image inside the target vehicle, the interior and exterior scene images of the target vehicle, and the vehicle status data of the target vehicle; The gaze direction vector and the gaze depth distance calculated based on the driver's binocular disparity are extracted from the facial image; Semantic segmentation and target recognition are performed on the images of the interior and exterior scenes of the vehicle to generate a set of scene semantic tags containing spatial location information; Using the gaze direction vector and the gaze depth distance, the specific semantic target corresponding to the driver's gaze point is located in a three-dimensional spatial reference system constructed using the spatial position information. When the specific semantic target belongs to a preset high-value record category, it is determined that the first triggering condition is met. Based on the vehicle status data, the cognitive load index and operational load index of the current driving task of the target vehicle are determined. When both the cognitive load index and the operational load index are lower than the preset safety threshold, it is determined that the second triggering condition is met. When the first triggering condition and the second triggering condition are met, a shooting instruction for the specific semantic target is generated, and the imager is controlled to shoot according to the position and motion state of the specific semantic target through the shooting instruction.
[0006] Optionally, the step of resolving the gaze direction vector and the gaze depth distance calculated based on the driver's binocular disparity from the facial image includes: Detect the coordinate positions of the centers of the left and right pupils in the facial image, and calculate the pixel position deviation values of the centers of the left and right pupils in the horizontal direction; The gaze depth distance is determined based on the preset actual distance parameters between the driver's binoculars and the equivalent focal length parameters of the camera that acquires the facial image, the inverse relationship between pixel position deviation value and gaze depth distance; In a sequence of facial image frames composed of multiple temporally consecutive facial images, if the spatial position difference between the calculated gaze depth distances in two adjacent facial images is less than a preset proximity distance threshold, then the gaze actions in the two adjacent facial images are classified as the same gaze event, and the duration of the same gaze event is accumulated to determine the gaze duration.
[0007] Optionally, the shooting method based on gaze depth and scene verification further includes: If an explicit trigger signal is received from manual operation, voice command, or remote terminal, the shooting command is generated. If no explicit trigger signal is received and the current mode is detected to be automatic, then when determining the first trigger condition, if it is detected that the driver's gaze duration toward the specific semantic target exceeds a preset first duration threshold, the step of determining whether the second trigger condition is met is executed.
[0008] Optionally, the step of locating the specific semantic target corresponding to the driver's gaze point in a three-dimensional spatial reference frame constructed using the gaze direction vector and the gaze depth distance, and determining that the first triggering condition is met when the specific semantic target belongs to a preset high-value record category, includes: The gaze point position determined based on the gaze direction vector is transformed from the facial image coordinate system to the vehicle space coordinate system with the target vehicle as the reference. The vehicle space partition where the gaze point is located is determined based on the converted gaze point location. The vehicle space partition includes at least the windshield viewing area, the window viewing area, and the rear passenger area. When the gaze point is located in the perspective area of the vehicle window, the gaze behavior is determined to be a distant view gaze, a medium view gaze, or a close view gaze based on the gaze depth distance. Check whether the proportion of pixel area occupied by a specific semantic target belonging to the high-value record category in the vehicle exterior scene image area corresponding to the vehicle space partition where the gaze point is located exceeds a preset proportion threshold. Check whether the gaze type corresponding to the gaze depth distance is consistent with the spatial attributes of the actual semantic labels in the vehicle exterior scene image. Among them, the distant gaze type corresponds to the distant semantic label, the mid-range gaze type corresponds to the mid-range semantic label, and the near gaze type corresponds to the near semantic label. When the pixel area ratio exceeds the percentage threshold, the gaze type is consistent with the spatial attributes of the actual semantic tag, and the confidence level of the scene semantic tag set is higher than the preset confidence level, the first triggering condition is determined to be met.
[0009] Optionally, based on the vehicle status data, the cognitive load index and operational load index of the current driving task of the target vehicle are determined. When both the cognitive load index and the operational load index are lower than a preset safety threshold, the second triggering condition is determined to be met, including: Based on the cognitive load index and operational load index of the current driving task of the target vehicle, the driver's driving load status is determined to be in the low load range, medium load range, or high load range. When the driving load is in the low load range, if the driver's continuous gaze duration on the specific semantic target exceeds a preset second duration threshold, it is determined that the second triggering condition is met. When the driving load is in the medium load range, if the driver's continuous gaze duration on the specific semantic target exceeds a preset third duration threshold, it is determined that the second triggering condition is met, wherein the third duration threshold is greater than the second duration threshold. When the driving load is in the high load range, it is determined that the second triggering condition is not met.
[0010] Optionally, generating a shooting command for the specific semantic target, and controlling the imager to take a picture based on the position and motion state of the specific semantic target using the shooting command, includes: Based on the vehicle space partition where the specific semantic target is located and the distance attribute of the specific semantic target, the target imager is selected from multiple imagers on the target vehicle to perform this shooting task; The coordinate system used to acquire facial images is converted to the device coordinate system where the target imager is located, and the horizontal rotation angle parameter and vertical pitch angle parameter required for the target imager to frame the specific semantic target are calculated in the device coordinate system. The focusing distance parameter required for the target imager to achieve focusing is calculated based on the gaze depth distance; Control commands, including the horizontal rotation angle parameter, the vertical pitch angle parameter, and the focusing distance parameter, are sent to the drive device associated with the target imager to drive the target imager to complete attitude adjustment and focus setting. After the target imager has stabilized, an image acquisition command is sent to the target imager to capture static images or record dynamic videos.
[0011] Optionally, calculating the focusing distance parameter required for the target imager to focus based on the gaze depth distance includes: When the depth of gaze is greater than fifty meters, the optical magnification of the target imager is set to the maximum telephoto magnification. When the viewing depth distance is greater than 10 meters and less than or equal to 50 meters, the optical magnification of the target imager is set to the standard focal length magnification. When the viewing depth distance is less than or equal to ten meters, the optical magnification of the target imager is set to the minimum wide-angle magnification, wherein, The maximum telephoto magnification is greater than the standard focal length magnification, and the standard focal length magnification is greater than the minimum wide-angle magnification.
[0012] The present invention also provides a shooting system based on gaze depth and scene verification, the system comprising: The acquisition module is used to acquire facial images of the driver inside the target vehicle, images of the interior and exterior scenes of the target vehicle, and vehicle status data of the target vehicle. The calculation module is used to parse the gaze direction vector and the gaze depth distance calculated based on the driver's binocular disparity from the facial image; The recognition module is used to perform semantic segmentation and target recognition on the in-vehicle and out-of-vehicle scene images, and generate a scene semantic tag set containing spatial location information; The first judgment module is used to locate the specific semantic target corresponding to the driver's gaze point in a three-dimensional spatial reference system constructed using the gaze direction vector and the gaze depth distance. When the specific semantic target belongs to a preset high-value record category, it is determined that the first triggering condition is met. The second judgment module is used to determine the cognitive load index and operational load index of the current driving task of the target vehicle based on the vehicle status data. When both the cognitive load index and the operational load index are lower than the preset safety threshold, it is determined that the second triggering condition is met. The output module is used to generate a shooting instruction for the specific semantic target when the first triggering condition and the second triggering condition are met, and to control the imager to shoot according to the position and motion state of the specific semantic target through the shooting instruction.
[0013] The present invention also provides a vehicle, the vehicle including a memory and a processor, the memory storing a computer program, the processor executing the computer program to implement the shooting method based on gaze depth and scene verification as described in any of the preceding claims.
[0014] The present invention also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the shooting method based on gaze depth and scene verification as described in any of the preceding claims.
[0015] The present invention has at least the following beneficial effects: First, at the intent recognition level, it not only analyzes the driver's gaze direction but, more importantly, introduces gaze depth distance calculated based on binocular parallax. Combined with 3D spatial labels generated by scene semantic segmentation, it accurately maps 2D gaze to specific targets in 3D space. Only when the target belongs to a high-value record category is the first trigger condition met, effectively filtering meaningless random glances and ensuring that only valuable content truly of interest to the driver is captured. Second, at the safety decision-making level, it dynamically assesses the cognitive and operational load of the driving task using vehicle status data. Only when both load indices are below a safety threshold is the second trigger condition met. This mechanism forces the system to prioritize driving safety, avoiding automatic shooting interference when the driver is under stress or operating complex tasks. Finally, only when both a valuable gaze target and a safe driving state are simultaneously met does the system generate a shooting command and control the imager. This solution achieves intelligent triggering based on genuine intent while ensuring that the shooting behavior does not jeopardize driving safety, resulting in accurate, intelligent, and safe image capture. Attached Figure Description
[0016] The accompanying drawings are provided to further understand the technical solutions of the present invention and constitute a part of the specification. They are used together with the embodiments of the present invention to explain the technical solutions of the present invention, and do not constitute a limitation on the technical solutions of the present invention.
[0017] Figure 1 This is a flowchart of a shooting method based on depth of gaze and scene verification. Figure 2 This is a flowchart of step S102 in a shooting method based on depth of gaze and scene verification; Figure 3 This is a flowchart of step S104 in a shooting method based on depth of gaze and scene verification; Figure 4 This is a flowchart of step S105 in a shooting method based on depth of gaze and scene verification. Figure 5 This is a flowchart of step S106 in a shooting method based on depth of gaze and scene verification; Figure 6 This is another step in a shooting method based on depth of gaze and scene verification; Figure 7 This is a schematic diagram of a shooting system based on depth of gaze and scene verification; Figure 8 This is another structural diagram of a shooting system based on depth of gaze and scene verification. Detailed Implementation
[0018] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.
[0019] The developers of this technical solution discovered that existing technologies treat "human eye / head movements" merely as switch signals or directional signals, failing to fully explore the deep intent contained in human eye gaze information, and further failing to fuse and verify human eye information with the semantic content of the in-vehicle and out-of-vehicle scenes. When observing objects, the human visual system includes not only directional information but also distance information (achieved through binocular parallax); simultaneously, whether the content being observed by the user has recording value needs to be verified through external scene semantic analysis. Existing solutions neglect both gaze distance and scene semantic verification, resulting in an inability to accurately distinguish between "normal driving" and "shooting intent." Furthermore, existing solutions generally lack dynamic assessment of driving safety load, failing to properly balance the relationship with driving safety while pursuing intelligent shooting.
[0020] To address the problems of existing technologies, this technical solution provides a shooting method, system, vehicle, and medium based on gaze depth and scene verification. This method can accurately identify and capture moments when the driver has real shooting value, enabling intelligent and precise shooting of in-vehicle and external images. The following are various embodiments of this technical solution.
[0021] Please refer to Figure 1 , Figure 1 This is a flowchart of a shooting method based on depth of gaze and scene verification.
[0022] This embodiment provides a shooting method based on gaze depth and scene verification, including: S101. Acquire the driver's facial image inside the target vehicle, the interior and exterior scene images of the target vehicle, and the vehicle status data of the target vehicle.
[0023] S102. Extract the gaze direction vector and the gaze depth distance calculated based on the driver's binocular disparity from the facial image.
[0024] S103. Perform semantic segmentation and target recognition on the images of the inside and outside of the vehicle to generate a set of scene semantic labels containing spatial location information.
[0025] S104. Using the gaze direction vector and gaze depth distance, locate the specific semantic target corresponding to the driver's gaze point in a three-dimensional spatial reference system constructed using spatial location information. When the specific semantic target belongs to a preset high-value record category, determine that the first triggering condition is met.
[0026] S105. Based on vehicle status data, determine the cognitive load index and operational load index of the target vehicle's current driving task. When both the cognitive load index and operational load index are lower than the preset safety threshold, determine that the second triggering condition is met.
[0027] S106. When the first triggering condition and the second triggering condition are met, a shooting instruction for a specific semantic target is generated, and the imager is controlled to shoot according to the position and motion state of the specific semantic target through the shooting instruction.
[0028] Understandably, firstly, at the intent recognition level, it not only analyzes the driver's gaze direction but, more importantly, introduces gaze depth distance calculated based on binocular parallax. Combined with 3D spatial labels generated by scene semantic segmentation, it accurately maps 2D gaze to specific targets in 3D space. Only when the target belongs to a high-value record category is the first trigger condition met, effectively filtering meaningless random glances and ensuring that only valuable content truly of interest to the driver is captured. Secondly, at the safety decision-making level, vehicle status data is used to dynamically assess the cognitive and operational load of the driving task. Only when both load indices are below a safety threshold is the second trigger condition met. This mechanism forces the system to prioritize driving safety, avoiding automatic shooting interference when the driver is under stress or operating complex tasks. Finally, only when both a valuable gaze target and a safe driving state are simultaneously satisfied does the system generate a shooting command and control the imager. This solution achieves intelligent triggering based on genuine intent while ensuring that the shooting behavior does not endanger driving safety, resulting in accurate, intelligent, and safe image capture.
[0029] Please refer to Figure 2 , Figure 2 This is a flowchart of step S102 in a shooting method based on depth of gaze and scene verification.
[0030] In some embodiments, step S102 includes: S201. Detect the coordinate positions of the centers of the left and right pupils in the facial image, and calculate the pixel position deviation values of the centers of the left and right pupils in the horizontal direction.
[0031] S202. Based on the preset actual distance parameters between the driver's binoculars and the equivalent focal length parameters of the camera that acquires facial images, the inverse relationship between pixel position deviation value and gaze depth distance, determine the gaze depth distance.
[0032] S203. In a sequence of facial image frames composed of multiple temporally consecutive facial images, if the spatial position difference between the calculated gaze depth distances in two adjacent facial images is less than a preset proximity distance threshold, then the gaze actions in the two adjacent facial images are classified as the same gaze event, and the duration of the same gaze event is accumulated to determine the gaze duration.
[0033] Specifically, when calculating the pixel position deviation value, the binocular parallax method is used: detect the center coordinates of the left and right pupils, calculate the parallax value d, and substitute it into the formula D = (B × f) / d, where B is the binocular baseline distance (obtained through calibration), f is the equivalent focal length of the camera, and d is the parallax (pixels).
[0034] Specifically, when calculating the duration of gaze, a gaze tracking queue is established, and gaze points in consecutive frames that are less than a threshold (such as 0.5 meters) are grouped into the same gaze event, and the duration is accumulated.
[0035] Understandably, this embodiment significantly improves the accuracy of intent recognition by accurately calculating gaze depth through pupil deviation and geometric parameters; it also introduces temporal frame sequence analysis to filter short-term flickering into continuous gaze events and accumulate their duration, effectively eliminating meaningless momentary glances. Combined with dual verification of existing scene semantics and safety load, this solution not only accurately identifies high-value targets that the driver is truly focused on, but also ensures that shooting is triggered only under safe and continuous gaze conditions. This achieves intelligent and accurate capture of in-vehicle and out-of-vehicle images while ensuring driving safety, significantly reducing invalid recordings.
[0036] In some embodiments, semantic segmentation processing is performed on the external scene image of the target vehicle to assign a corresponding semantic category label to each pixel in the external scene image.
[0037] In some embodiments, facial expression recognition and human posture estimation processing are performed on the interior scene image of the target vehicle to detect the facial expression category and body movement posture category of the occupants, and the combination of expressions and postures that represent warm interaction is determined as emotional signals.
[0038] In some embodiments, abnormal traffic event detection processing is uniformly performed on images of in-vehicle and out-of-vehicle scenes to identify whether there are traffic accident patterns, road obstacles, or pedestrians crossing the road.
[0039] In some embodiments, a deep learning model is used to directly regress a depth map from a single image.
[0040] In some embodiments, distance is measured directly using a dedicated ToF sensor or lidar.
[0041] In some embodiments, a shooting method based on gaze depth and scene verification further includes: If an explicit trigger signal is received from manual operation, voice command, or remote terminal, a shooting command is generated; if no explicit trigger signal is received and the current mode is detected to be automatic, when determining the first trigger condition, if it is detected that the driver's gaze duration toward a specific semantic target exceeds a preset first duration threshold, the step of determining whether the second trigger condition is met is executed.
[0042] Understandably, in this embodiment, explicit signals provide immediate control, satisfying the user's active recording needs. In automatic mode, a gaze duration threshold is introduced as a pre-judgment to ensure that safety load verification is only initiated when the driver maintains continuous and stable attention to a high-value target. This mechanism effectively avoids false triggering caused by momentary glances, further improving the reliability and accuracy of shooting intent recognition, achieving an upgrade from passive response to intelligent prediction, and balancing operational flexibility with system rigor.
[0043] Please refer to Figure 3 , Figure 3 This is a flowchart of step S104 in a shooting method based on depth of gaze and scene verification.
[0044] In some embodiments, step S104 includes: S301. Transform the gaze point position determined based on the gaze direction vector from the facial image coordinate system to the vehicle space coordinate system with the target vehicle as the reference.
[0045] S302. Determine the vehicle space partition where the gaze point is located based on the converted gaze point position. The vehicle space partition includes at least the windshield viewing area, the window viewing area, and the rear passenger area.
[0046] S303. When the gaze point is located in the perspective area of the vehicle window, the gaze behavior type is determined as a distant view gaze type, a medium view gaze type, or a near view gaze type based on the gaze depth distance.
[0047] S304. Check whether the proportion of pixel area occupied by a specific semantic target belonging to the high-value record category in the vehicle exterior scene image area corresponding to the vehicle space partition where the gaze point is located exceeds the preset proportion threshold.
[0048] S305. Check whether the gaze type corresponding to the gaze depth distance is consistent with the spatial attributes of the actual semantic labels in the external scene image. Among them, the distant gaze type corresponds to the distant semantic label, the mid-range gaze type corresponds to the mid-range semantic label, and the near gaze type corresponds to the near semantic label.
[0049] S306. When the pixel area ratio exceeds the proportion threshold, the gaze type is consistent with the spatial attributes of the actual semantic label, and the confidence of the scene semantic label set is higher than the preset confidence value, the first triggering condition is determined to be met.
[0050] In this embodiment, checking the external scene image area corresponding to the vehicle space partition where the gaze falls refers to checking whether the image area corresponding to the gaze area contains semantic content with recording value. For example, if the gaze is on the left window, check whether the semantic proportion of "natural landscape" in the left camera image exceeds a threshold (such as 30%); if the gaze is on the right window, check whether "accident vehicle" or "abnormal event" is detected; if the gaze is on the back seat, check whether "smiling face" or "child / pet" is detected.
[0051] Checking whether the gaze type corresponding to the gaze depth distance is consistent with the spatial attributes of the actual semantic labels in the external scene image means checking whether the gaze distance D matches the scene semantics. For example, if D > 50 meters, check whether the scene contains distant semantics (mountains, sky, distant buildings); if D < 10 meters, check whether the scene contains near semantics (roadside flowers, pedestrians).
[0052] Understandably, this embodiment accurately locates the gaze area through coordinate system transformation and spatial partitioning; it determines the far / mid / near shot type based on depth distance, and introduces triple verification of "target pixel ratio," "spatial attribute consistency," and "semantic confidence." This effectively filters out background interference, misjudgments, and low-confidence scenes, ensuring that shooting is triggered only when the driver is focused on a high-value and distinctive target. This mechanism significantly improves the robustness and accuracy of intent recognition, eliminates invalid captures, and achieves truly high-quality intelligent shooting based on genuine points of interest.
[0053] In some embodiments, when eye tracking fails (e.g., when wearing dark sunglasses), the method degenerates to using head posture as an approximate estimate of the gaze direction.
[0054] In some embodiments, fuzzy voice commands are combined with gaze information to further improve the confidence level of intent recognition.
[0055] Please refer to Figure 4 , Figure 4 This is a flowchart of step S105 in a shooting method based on depth of gaze and scene verification.
[0056] In some embodiments, step S105 includes: S401. Based on the cognitive load index and operational load index of the current driving task of the target vehicle, determine whether the driver's driving load status is in the low load range, medium load range, or high load range.
[0057] S402. When the driving load is in the low load range, if the driver's continuous gaze duration on a specific semantic target exceeds the preset second duration threshold, the second triggering condition is determined to be met.
[0058] S403. When the driving load is in the medium load range, if the driver's continuous gaze duration on a specific semantic target exceeds a preset third duration threshold, it is determined that the second triggering condition is met, wherein the third duration threshold is greater than the second duration threshold.
[0059] S404. When the driving load is in the high load range, it is determined that the second trigger condition is not met.
[0060] Specifically, the method for driver load assessment is to calculate the driver load level L in real time based on a multi-parameter fusion algorithm. The calculation formula is: L = w1·f1(v, a) + w2·f2(PERCLOS) + w3·f3(θ, ω) Wherein, f1(v, a) is the driving complexity function based on vehicle speed and acceleration signals from the vehicle state perception module; f2(PERCLOS) is the fatigue function based on the proportion of eye closure time, which is calculated by analyzing the proportion of eye closure frames in consecutive frames; f3(θ, ω) is the operation intensity function based on steering wheel angular velocity and yaw rate; w1, w2, and w3 are weighting coefficients obtained through real vehicle calibration.
[0061] The load level L ranges from [0,1] and is divided into three intervals: Low load: L<0.3 (such as high-speed cruising, straight driving); Medium load: 0.3 ≤ L ≤ 0.6 (e.g., following other vehicles on urban roads); High load: L>0.6 (such as turning at complex intersections, emergency avoidance).
[0062] Read the current driver load level L. If L < 0.3 (low load) and the gaze duration T exceeds the second threshold T2 (default 1.5 seconds), confirm the triggering of the shooting event; if 0.3 ≤ L ≤ 0.6 (medium load) and the gaze duration T exceeds the third threshold T3 (default 2.5 seconds), confirm the triggering of the shooting event; if L > 0.6 (high load), stop shooting.
[0063] Understandably, the system categorizes states into low, medium, and high zones based on cognitive and operational load, and dynamically adjusts trigger thresholds: rapid response under low load, extended gaze duration under medium load to prevent accidental triggering, and direct prohibition of recording under high load. This tiered strategy achieves "safety-first" adaptive control, ensuring driver focus during busy periods while also accommodating recording needs during idle periods, significantly improving the system's safety and intelligent decision-making capabilities in complex road conditions.
[0064] In some embodiments, a large visual language model is introduced to gain a deeper understanding and description of the image and directly determine whether the scene is worth recording (such as "mountains under the setting sun").
[0065] In some embodiments, navigation map information or cloud-based crowdsourced data are combined to predict whether the current road segment is a scenic route or whether there are frequently occurring accident points, thus assisting in scenario judgment.
[0066] Please refer to Figure 5 , Figure 5 This is a flowchart of step S106 in a shooting method based on depth of gaze and scene verification.
[0067] In some embodiments, step S106 includes: S501. Based on the vehicle space partition where the specific semantic target is located and the distance attribute of the specific semantic target, select the target imager to perform this shooting task from multiple imagers on the target vehicle.
[0068] S502. Convert the coordinate system used to acquire facial images to the device coordinate system where the target imager is located, and calculate the horizontal rotation angle parameter and vertical pitch angle parameter that the target imager needs to adjust to capture a specific semantic target in the device coordinate system.
[0069] S503. Calculate the focusing distance parameter required for the target imager to achieve focusing based on the depth of gaze distance.
[0070] S504. Send control commands, including horizontal rotation angle parameters, vertical pitch angle parameters, and focus distance parameters, to the drive device associated with the target imager to drive the target imager to complete attitude adjustment and focus setting.
[0071] S505. After the target imager's attitude stabilizes, send an image acquisition command to the target imager to capture static images or record dynamic videos through the target imager.
[0072] In this embodiment, the target imager for performing this shooting task is selected as follows: the most suitable active camera device is selected based on the shooting scene and target area. For example, if shooting the left side of the vehicle, the left rearview mirror gimbal camera or the roof gimbal camera is activated to turn left; if shooting the right side of the vehicle, the right rearview mirror gimbal camera or the roof gimbal camera is activated to turn right; if shooting the front of the vehicle, the roof gimbal camera or the front-view camera is activated; if shooting the interior of the vehicle, the gimbal camera at the interior rearview mirror is activated.
[0073] Understandably, this embodiment intelligently schedules the best acquisition device through spatial partitioning and multi-imager optimization; combined with coordinate system transformation and depth calculation, it automatically generates composite control commands including rotation, pitch, and focus, driving the imager to quickly lock onto the target and capture a clear image. This solution solves the problem that traditional fixed-viewpoint systems cannot cover dynamic targets, ensuring that the system can automatically adjust to the optimal shooting posture and focal length the moment the driver looks at it, significantly improving the success rate of capture, image quality, and response speed.
[0074] In some embodiments, the azimuth angle is calculated by transforming the human eye gaze direction vector (dx, dy, dz) from the camera coordinate system to the vehicle coordinate system, and then to the PTZ camera coordinate system to obtain the target horizontal angle α and pitch angle β.
[0075] Angle commands (α, β) are sent to the gimbal motor driver, and focal length commands (f) are sent to the lens driver chip to control the active camera device to accurately align and focus, and execute the shooting action.
[0076] Once the gimbal is stable, the image sensor is triggered to take a photo or record video. Recording mode can be started when the gaze begins and will stop when the gaze leaves or is triggered again.
[0077] In some embodiments, multiple fixed-focus cameras with different focal lengths (such as wide-angle and telephoto) are arranged around the vehicle. The system automatically switches to the camera view with the appropriate focal length based on the viewing distance.
[0078] In some embodiments, a wide-angle camera is used for shooting, but the area being viewed by the user is precisely cropped and digitally zoomed using a high-pixel sensor and electronic image stabilization.
[0079] In some embodiments, step S503 includes: (1) When the viewing depth distance is greater than 50 meters, set the optical magnification of the target imager to the maximum telephoto magnification.
[0080] (2) When the depth of gaze is greater than 10 meters and less than or equal to 50 meters, the optical magnification of the target imager is set to the standard focal length magnification.
[0081] (3) When the fixation depth distance is less than or equal to ten meters, set the optical magnification of the target imager to the minimum wide-angle magnification, where the maximum telephoto magnification is greater than the standard focal length magnification, and the standard focal length magnification is greater than the minimum wide-angle magnification.
[0082] Specifically, determine the focal length value f according to the fixation distance D and the preset imaging rule: If D > 50 meters, f = f_max (telephoto end, such as 10x optical zoom); if 10 meters < D ≤ 50 meters, f = f_mid (standard focal length, such as 3x optical zoom); if D ≤ 10 meters, f = f_min (wide-angle end, such as 1x optical zoom).
[0083] Please refer to Figure 6 , Figure 6 which is another step flowchart of a shooting method based on fixation depth and scene verification.
[0084] The shooting method based on fixation depth and scene verification in this embodiment includes the following steps: Step S100: System initialization and real-time monitoring.
[0085] After the system is powered on, the human eye intention perception module collects driver face images at a frame rate of 60fps and runs a face key point detection algorithm; the in-vehicle and out-of-vehicle scene perception module collects video streams of each camera at a frame rate of 30fps; the vehicle state perception module reads CAN bus data at a frequency of 100Hz.
[0086] Step S200: Multimodal information analysis.
[0087] The storage and comprehensive processing module processes the human eye data: human eye intention analysis and scene semantic analysis.
[0088] Step S300: Judging shooting trigger conditions.
[0089] The system continuously listens for trigger signals and judges whether a manual / voice / remote trigger signal is received. If a clear trigger signal is received, it has the highest priority and directly enters step S400.
[0090] If the system is in automatic mode, detect whether the fixation duration T exceeds a preset first threshold T1 (default value 0.5 seconds). If T < T1, return to step S200; if T ≥ T1, enter the preliminary interest judgment.
[0091] The system continuously monitors the vehicle state. If a limit state (such as skidding, sharp turning, collision, etc.) is detected, force all cameras to perform panoramic recording. For example, lateral acceleration |ay| > 0.5g (about 4.9 m / s²), yaw rate |ω| > 30° / s, vehicle body roll angle |φ| > 10°, airbag trigger signal.
[0092] Step S400: Shooting execution and optimization control.
[0093] Step S500: Data storage and post-processing.
[0094] Specifically, the captured images / videos are encapsulated and stored along with metadata. The image / video files are saved together with a .json metadata file of the same name. The metadata includes timestamps, GPS coordinates, vehicle speed, vehicle attitude, driver workload level, and gaze point coordinates.
[0095] Important events (such as those triggered by extreme conditions) are stored in the protected "Events" partition, which is protected by a loop; ordinary shots are stored in the "Highlights" partition, which can be manually deleted by the user.
[0096] This technical solution also provides implementation examples for specific scenarios.
[0097] In one embodiment of intelligent recording of the scenery outside the vehicle, the driver gazes at distant mountains through the windshield. The human eye intention perception module calculates that the gaze direction is 15° to the left and forward, the gaze distance is D≈650 meters, and the gaze duration begins to accumulate.
[0098] Meanwhile, the scene semantic analysis engine analyzes the front-view camera footage: the semantic segmentation results show that the top 35% of the image is "sky", the middle 40% is "mountains", and the rest is "road"; natural landscapes account for more than 60%, the scene is classified as "natural scenery" with a confidence level of 0.92; the scene has recording value (it is scenery).
[0099] At 1.8 seconds, T = 1.8 seconds. Load level L = 0.25 (low load).
[0100] Two-way verification: The gaze point is located in the left window area, corresponding to a high semantic proportion of "mountain range" in the image, and the region matching is successful; D=650 meters matches the semantic meaning of "distant view", with a confidence level of 0.92>0.7, indicating that the driver is paying attention to a valuable scene, and the distance matching is successful.
[0101] The system confirms the shooting intention, controls the pan-tilt camera to turn in the direction of the view, adjusts the focal length to the telephoto end, and takes a photo.
[0102] Save the photos and metadata, with the metadata recording the scene semantic tag "natural scenery".
[0103] In another specific embodiment, the driver gazes at the road ahead. The human eye intention perception module calculates that the gaze direction is directly forward, the gaze distance D ≈ 25 meters (following distance), and the gaze duration is continuous.
[0104] The scene semantic analysis engine analyzes the image from the forward-facing camera: the semantic segmentation results show that the main components in the image are "road" (60%), "vehicle in front" (20%), and "building" (15%); the proportion of natural landscape is <5%, the scene is classified as "urban road" with a confidence level of 0.95; the scene is determined to have no recording value (ordinary road). At this time, the driver's gaze lasts for 2.0 seconds, and the load level L=0.5 (medium load).
[0105] Based on the above analysis, if the triggering conditions are not met, the system will not trigger shooting and will continue monitoring.
[0106] In one specific embodiment of recording the in-vehicle environment, the driver turns their head to look at the children in the back seat. The human eye intention perception module detects the gaze point coordinates in the back seat area, D=2.5 meters.
[0107] The scene semantic analysis engine analyzed the rear camera footage and detected a child's face with a "smiling" expression (confidence 0.88). The scene was classified as "warm moment" (confidence 0.85). At this time, the driver's gaze lasted for 3.2 seconds, and the load level L=0.35.
[0108] When determining the trigger condition, since the gaze point corresponds to the child's face, the region matching is successful. D=2.5 meters matches the close-up view inside the car, and the distance matching is successful. Confidence level: 0.85>0.7.
[0109] The system confirms the shooting intention and controls the rearview mirror pan-tilt camera to shoot in a wide angle.
[0110] In a specific embodiment of extreme state recording, when the vehicle skids on a wet and slippery road surface, the system detects via the CAN bus that the lateral acceleration ay = 0.6g (5.88 m / s²) > 0.5g threshold, the yaw rate ω = 35° / s > 30° / s threshold, and the vehicle roll angle φ = 8° (close to the threshold).
[0111] Upon detecting an extreme state, regardless of scene semantics, current driver load level, and gaze status, a forced recording mode is immediately triggered.
[0112] The system sends synchronous recording commands to all cameras (front view, rear view, left and right side view, surround view, and pan / tilt).
[0113] All cameras start recording simultaneously, with the frame rate increased to 60fps to capture details.
[0114] The recording lasts for 30 seconds, including 5 seconds before the slide (pre-recorded in the loop buffer) and 25 seconds after the slide.
[0115] Recorded files are automatically encrypted and stored, marked as "extreme events," and cannot be deleted in any way (they can only be cleared during maintenance via diagnostic equipment).
[0116] Compared with the prior art, the above embodiments have the following beneficial effects: (1) Improved ease of operation and accuracy of intent recognition: By integrating three-dimensional information of gaze direction, duration and distance, an interactive method that conforms to the human intuition of "gazing is taking pictures" is realized, avoiding the need to learn complex instructions. At the same time, the introduction of gaze duration and distance dimensions can effectively distinguish between a brief glance and a conscious gaze, reducing the false trigger rate from the source.
[0117] (2) Improved shooting quality and accuracy: By accurately estimating the viewing distance, the active camera device is dynamically driven to adjust the focal length (optical zoom) and accurately align, achieving high-quality imaging that is "what you see is what you get" for the user, and avoiding image quality loss caused by electronic cropping.
[0118] (3) Enhanced driving safety: A safety arbitration mechanism was constructed by introducing the driver's load level as a weighting factor for triggering shooting. Under high-load driving conditions, the system automatically increases the trigger threshold or suppresses non-emergency shooting, effectively balancing the relationship between intelligent services and driving safety.
[0119] (4) Reduced false trigger rate and invalid records: Through the two-way confirmation mechanism of "gazing intent analysis" and "scene semantic analysis", the shooting is triggered only when the content being gazed at by the user is confirmed by scene analysis to be of recording value (such as scenery, smiling face, accident), which greatly improves the accuracy of intent recognition and avoids a large number of invalid records.
[0120] (5) Rich recording dimensions and application scenarios: It not only covers the scenery outside the car, but also intelligently captures the warm moments inside the car, and combines vehicle status data to give each recording rich spatiotemporal information, meeting the diverse recording needs of users. At the same time, it supports multiple triggering methods such as manual, voice, automatic, and extreme events, which complement each other.
[0121] Please refer to Figure 7 , Figure 7 This is a schematic diagram of a shooting system based on depth of gaze and scene verification.
[0122] This embodiment provides a shooting system based on gaze depth and scene verification, including: The acquisition module 601 is used to acquire facial images of the driver inside the target vehicle, images of the interior and exterior scenes of the target vehicle, and vehicle status data of the target vehicle. Calculation module 602 is used to parse the gaze direction vector and the gaze depth distance calculated based on the driver's binocular disparity from the facial image; The recognition module 603 is used to perform semantic segmentation and target recognition on images of scenes inside and outside the vehicle, and generate a set of scene semantic labels containing spatial location information. The first judgment module 604 is used to locate the specific semantic target corresponding to the driver's gaze point in a three-dimensional spatial reference system composed of spatial position information by using the gaze direction vector and gaze depth distance. When the specific semantic target belongs to a preset high-value record category, it is determined that the first triggering condition is met. The second judgment module 605 is used to determine the cognitive load index and operational load index of the current driving task of the target vehicle based on vehicle status data. When both the cognitive load index and the operational load index are lower than the preset safety threshold, it is determined that the second triggering condition is met. The output module 606 is used to generate a shooting command for a specific semantic target when the first trigger condition and the second trigger condition are met, and to control the imager to shoot according to the position and motion state of the specific semantic target through the shooting command.
[0123] Please refer to Figure 8 , Figure 8 This is another structural diagram of a shooting system based on depth of gaze and scene verification.
[0124] like Figure 8 As shown, this embodiment provides a shooting system based on gaze depth and scene verification. The system also includes a human eye intention perception module, a vehicle state perception module, an in-vehicle and out-of-vehicle event scene perception module, a triggering module, a storage and comprehensive processing module, and a display and setting module.
[0125] The human eye intention perception module consists of at least one camera or camera system facing the driver (such as located behind the steering wheel, in the rearview mirror, or in the dashboard area), and may also include devices such as a camera for the passenger. For example, the camera uses a global shutter CMOS image sensor with a frame rate of not less than 60fps and a resolution of not less than 720p, and is equipped with an 850nm near-infrared fill light to ensure normal operation in low-light environments.
[0126] This module is used to collect the state of the driver's and / or co-driver's head, face and eyes in real time. It can output information such as head posture, facial expression and eye gaze, but is not limited to this information: head posture (yaw angle, pitch angle, roll angle), eye gaze information (gaze direction - three-dimensional line of sight vector (dx, dy, dz), gaze point coordinates (three-dimensional coordinates in the vehicle coordinate system), gaze distance D).
[0127] The in-vehicle and out-of-vehicle event scene perception module includes a fixed-view camera array and active camera equipment. It is primarily used to collect in-vehicle and out-of-vehicle events and adjust the movements of the active camera equipment within the storage and integrated processing module. The fixed-view camera array includes front-view, rear-view, side-view, and surround-view cameras to provide a panoramic view and basic recording without blind spots. One or more camera units with independently controllable angle and focal length, such as gimbal cameras with pitch / rotate capabilities or cameras on vehicle-mounted drones, are the core actuators for performing "what you see is what you get" (WYSIWYG) shooting. Furthermore, by analyzing the data from the in-vehicle and out-of-vehicle event scene perception module, the complexity of the current driving environment (such as the number of targets and traffic density) is directly assessed, serving as a proxy indicator of load.
[0128] The vehicle state perception module is used to acquire vehicle motion state information and driver operation state information from traditional vehicle sensor systems, including wheel speed sensors, steering wheel angle sensors, IMU inertial measurement units and other sensors and control units. It can output information such as vehicle speed, yaw rate, roll angle and angular velocity, pitch angle and angular velocity, longitudinal acceleration, lateral acceleration, vertical acceleration, steering wheel angle, accelerator pedal position and speed, and brake pedal position and speed.
[0129] The trigger module supports multiple triggering methods, including: Automatic triggering means that it is triggered by the analysis results of the human eye intention perception module or by the vehicle's extreme state. The triggering conditions include the gaze duration exceeding the threshold and the gaze distance being stable, as well as the vehicle's extreme state signal.
[0130] Manual triggering includes the multi-function buttons on the steering wheel or near the passenger side, and the virtual buttons on the central control touchscreen. The buttons use capacitive touch sensing and support different operations such as single click, double click, and long press.
[0131] Voice triggering involves picking up voice commands through a microphone array, which are then processed by a voice recognition chip to identify keywords such as "take a photo" and "record a video".
[0132] Remote triggering is achieved by receiving trigger commands sent by the user's mobile app via a wireless communication module.
[0133] The storage and integrated processing module, as the core processing unit of the system, adopts an automotive-grade system-on-a-chip (SoC), such as TI TDA4VM or NVIDIA Orin, and integrates the following functions: (1) Time series synchronization: Accurate timestamps are added to all sensor and / or camera data to achieve microsecond-level synchronization.
[0134] (2) Driver load assessment: Driver load level L is calculated in real time based on a multi-parameter fusion algorithm.
[0135] (3) Scene semantic analysis engine: Based on deep learning models, the engine performs real-time semantic understanding of images inside and outside the vehicle and outputs scene types such as natural scenery, urban roads, highways, accident scenes, heartwarming moments, and dangerous working conditions; pixel-level segmented semantic regions such as sky, mountains, buildings, roads, vehicles, pedestrians, children, and pets; and scene quantitative indicators such as the proportion of natural landscapes, sky visibility, building density, and confidence of abnormal events.
[0136] (4) Intent arbitration and decision-making: Integrate gaze information, scene semantics, trigger signals, load level and vehicle status to execute shooting decision logic.
[0137] The display and settings module, integrated into the central control display screen, displays system status icons, a heatmap of currently identified gaze points, and preview images from each camera. It allows setting shooting modes (manual / automatic / voice / remote, etc.), viewing angle preferences (driver's view / passenger's view, where the system references the passenger's gaze point in passenger's view mode), active camera parameters (including default focal length - wide-angle / standard / telephoto), shooting resolution (1080P / 4K, etc.), sensitivity, etc. Additionally, the display and settings module also features event data playback and other related functions.
[0138] This invention also provides a vehicle control system, including a memory, a processor, and a program stored in the memory and executable on the processor. When the program is executed by the processor, it implements the shooting method based on gaze depth and scene verification described in the above embodiments.
[0139] Taking the example of a processor and memory in a vehicle controller being connected via a bus, the memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs and non-transitory computer-executable programs. Furthermore, the memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some embodiments, the memory may optionally include memory remotely located relative to the control processor, and these remote memories can be connected to the control device via a network. The non-transitory software programs and instructions required to implement the control methods of the above embodiments are stored in the memory, and when executed by the processor, the control methods of the above embodiments are performed.
[0140] The system embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.
[0141] This invention also provides a vehicle, including the vehicle control system described in the above embodiments.
[0142] The vehicle can be a private car, such as a sedan, SUV, MPV, or pickup truck. It can also be a commercial vehicle, such as a van, bus, small truck, or large semi-trailer. The vehicle must have an electric motor capable of outputting power or acting as a generator to store mechanical energy. When the vehicle is a new energy vehicle, it can be a hybrid or a pure electric vehicle.
[0143] Since the vehicle applies all the technical solutions of the above-mentioned control device or vehicle controller, it has at least all the beneficial effects brought about by the technical solutions of the above embodiments, which will not be repeated here.
[0144] Furthermore, one embodiment of the present invention provides a computer-readable storage medium storing computer-executable instructions for performing the above-described shooting method based on depth of gaze and scene verification.
[0145] It is worth noting that, since the computer-readable storage medium of the present invention is capable of executing the shooting method based on gaze depth and scene verification of any of the above embodiments, the specific implementation and technical effects of the computer-readable storage medium of the present invention can be referred to the specific implementation and technical effects of the shooting method based on gaze depth and scene verification of any of the above embodiments.
[0146] Furthermore, one embodiment of the present invention also provides a computer program product, including a computer program or computer instructions, which are stored in a computer-readable storage medium. A processor of a computer device reads the computer program or computer instructions from the computer-readable storage medium and executes the computer program or computer instructions, causing the computer device to perform the above-described shooting method based on gaze depth and scene verification.
[0147] It is worth noting that, since the computer program product of this embodiment can execute the shooting method based on gaze depth and scene verification of any of the above embodiments, the specific implementation method and technical effect of the computer program product of this embodiment can refer to the specific implementation method and technical effect of the shooting method based on gaze depth and scene verification of any of the above embodiments.
[0148] It will be understood by those skilled in the art that all or some of the steps and systems in the methods disclosed above can be implemented as software, firmware, hardware, and suitable combinations thereof. Some or all of the physical components can be implemented as software executed by a processor, such as a central processing unit, digital signal processor, or microprocessor, or as hardware, or as an integrated circuit, such as an application-specific integrated circuit. Such software can be distributed on a computer-readable medium, which can include computer storage media (or non-transitory media) and communication media (or transient media). As is known to those skilled in the art, the term computer storage media includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information (such as computer-readable instructions, data structures, program modules, or other data). Computer storage media includes, but is not limited to, RAM, ROM, EEPROM, flash memory or other memory technologies, CD-ROM, digital versatile disc (DVD) or other optical disc storage, magnetic cartridges, magnetic tape, disk storage or other magnetic storage devices, or any other medium that can be used to store desired information and is accessible to a computer. Furthermore, as is known to those skilled in the art, communication media typically include computer-readable instructions, data structures, program modules, or other data in modulated data signals such as carrier waves or other transmission mechanisms, and may include any information delivery medium.
[0149] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.
Claims
1. A shooting method based on depth of gaze and scene verification, characterized in that, The method includes: Acquire the driver's facial image inside the target vehicle, the interior and exterior scene images of the target vehicle, and the vehicle status data of the target vehicle; The gaze direction vector and the gaze depth distance calculated based on the driver's binocular disparity are extracted from the facial image; Semantic segmentation and target recognition are performed on the images of the interior and exterior scenes of the vehicle to generate a set of scene semantic tags containing spatial location information; Using the gaze direction vector and the gaze depth distance, the specific semantic target corresponding to the driver's gaze point is located in a three-dimensional spatial reference system constructed using the spatial position information. When the specific semantic target belongs to a preset high-value record category, it is determined that the first triggering condition is met. Based on the vehicle status data, the cognitive load index and operational load index of the current driving task of the target vehicle are determined. When both the cognitive load index and the operational load index are lower than the preset safety threshold, it is determined that the second triggering condition is met. When the first triggering condition and the second triggering condition are met, a shooting instruction for the specific semantic target is generated, and the imager is controlled to shoot according to the position and motion state of the specific semantic target through the shooting instruction.
2. The method according to claim 1, characterized in that, The step of resolving the gaze direction vector and the gaze depth distance calculated based on the driver's binocular disparity from the facial image includes: Detect the coordinate positions of the centers of the left and right pupils in the facial image, and calculate the pixel position deviation values of the centers of the left and right pupils in the horizontal direction; The gaze depth distance is determined based on the preset actual distance parameters between the driver's binoculars and the equivalent focal length parameters of the camera that acquires the facial image, the inverse relationship between pixel position deviation value and gaze depth distance; In a sequence of facial image frames composed of multiple temporally consecutive facial images, if the spatial position difference between the calculated gaze depth distances in two adjacent facial images is less than a preset proximity distance threshold, then the gaze actions in the two adjacent facial images are classified as the same gaze event, and the duration of the same gaze event is accumulated to determine the gaze duration.
3. The method according to claim 1, characterized in that, The method further includes: If an explicit trigger signal is received from manual operation, voice command, or remote terminal, the shooting command is generated. If no explicit trigger signal is received and the current mode is detected to be automatic, then when determining the first trigger condition, if it is detected that the driver's gaze duration toward the specific semantic target exceeds a preset first duration threshold, the step of determining whether the second trigger condition is met is executed.
4. The method according to claim 3, characterized in that, The method involves using the gaze direction vector and the gaze depth distance to locate the specific semantic target corresponding to the driver's gaze point in a three-dimensional spatial reference frame constructed using the spatial position information. When the specific semantic target belongs to a preset high-value record category, a first triggering condition is determined to be met, including: The gaze point position determined based on the gaze direction vector is transformed from the facial image coordinate system to the vehicle space coordinate system with the target vehicle as the reference. The vehicle space partition where the gaze point is located is determined based on the converted gaze point location. The vehicle space partition includes at least the windshield viewing area, the window viewing area, and the rear passenger area. When the gaze point is located in the perspective area of the vehicle window, the gaze behavior is determined to be a distant view gaze, a medium view gaze, or a close view gaze based on the gaze depth distance. Check whether the proportion of pixel area occupied by a specific semantic target belonging to the high-value record category in the vehicle exterior scene image area corresponding to the vehicle space partition where the gaze point is located exceeds a preset proportion threshold. Check whether the gaze type corresponding to the gaze depth distance is consistent with the spatial attributes of the actual semantic labels in the vehicle exterior scene image. Among them, the distant gaze type corresponds to the distant semantic label, the mid-range gaze type corresponds to the mid-range semantic label, and the near gaze type corresponds to the near semantic label. When the pixel area ratio exceeds the percentage threshold, the gaze type is consistent with the spatial attributes of the actual semantic tag, and the confidence level of the scene semantic tag set is higher than the preset confidence level, the first triggering condition is determined to be met.
5. The method according to claim 3, characterized in that, Based on the vehicle status data, the cognitive load index and operational load index of the target vehicle's current driving task are determined. When both the cognitive load index and operational load index are lower than a preset safety threshold, a second triggering condition is determined to be met, including: Based on the cognitive load index and operational load index of the current driving task of the target vehicle, the driver's driving load status is determined to be in the low load range, medium load range, or high load range. When the driving load is in the low load range, if the driver's continuous gaze duration on the specific semantic target exceeds a preset second duration threshold, it is determined that the second triggering condition is met. When the driving load is in the medium load range, if the driver's continuous gaze duration on the specific semantic target exceeds a preset third duration threshold, it is determined that the second triggering condition is met, wherein the third duration threshold is greater than the second duration threshold. When the driving load is in the high load range, it is determined that the second triggering condition is not met.
6. The method according to claim 1, characterized in that, The process of generating a shooting command for the specific semantic target, and controlling the imager to take a picture based on the position and motion state of the specific semantic target using the shooting command, includes: Based on the vehicle space partition where the specific semantic target is located and the distance attribute of the specific semantic target, the target imager is selected from multiple imagers on the target vehicle to perform this shooting task; The coordinate system used to acquire facial images is converted to the device coordinate system where the target imager is located, and the horizontal rotation angle parameter and vertical pitch angle parameter required for the target imager to frame the specific semantic target are calculated in the device coordinate system. The focusing distance parameter required for the target imager to achieve focusing is calculated based on the gaze depth distance; The system sends control commands, including the horizontal rotation angle parameter, the vertical pitch angle parameter, and the focusing distance parameter, to the drive device associated with the target imager to drive the target imager to complete attitude adjustment and focus setting. After the target imager has stabilized, an image acquisition command is sent to the target imager to capture static images or record dynamic videos.
7. The method according to claim 6, characterized in that, The step of calculating the focusing distance parameter required for the target imager to focus based on the gaze depth distance includes: When the depth of gaze is greater than fifty meters, the optical magnification of the target imager is set to the maximum telephoto magnification. When the viewing depth distance is greater than 10 meters and less than or equal to 50 meters, the optical magnification of the target imager is set to the standard focal length magnification. When the viewing depth distance is less than or equal to ten meters, the optical magnification of the target imager is set to the minimum wide-angle magnification, wherein, The maximum telephoto magnification is greater than the standard focal length magnification, and the standard focal length magnification is greater than the minimum wide-angle magnification.
8. A shooting system based on gaze depth and scene verification, characterized in that, The system includes: The acquisition module is used to acquire facial images of the driver inside the target vehicle, images of the interior and exterior scenes of the target vehicle, and vehicle status data of the target vehicle. The calculation module is used to parse the gaze direction vector and the gaze depth distance calculated based on the driver's binocular disparity from the facial image; The recognition module is used to perform semantic segmentation and target recognition on the in-vehicle and out-of-vehicle scene images, and generate a scene semantic tag set containing spatial location information; The first judgment module is used to locate the specific semantic target corresponding to the driver's gaze point in a three-dimensional spatial reference system constructed using the gaze direction vector and the gaze depth distance. When the specific semantic target belongs to a preset high-value record category, it is determined that the first triggering condition is met. The second judgment module is used to determine the cognitive load index and operational load index of the current driving task of the target vehicle based on the vehicle status data. When both the cognitive load index and the operational load index are lower than the preset safety threshold, it is determined that the second triggering condition is met. The output module is used to generate a shooting instruction for the specific semantic target when the first triggering condition and the second triggering condition are met, and to control the imager to shoot according to the position and motion state of the specific semantic target through the shooting instruction.
9. A vehicle, characterized in that, The vehicle includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the shooting method based on gaze depth and scene verification as described in any one of claims 1 to 7.
10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the shooting method based on depth of gaze and scene verification as described in any one of claims 1 to 7.