Depth estimation system and method thereof
By combining pinhole and fisheye cameras, a depth estimation system is developed. This system utilizes image scaling and rotation compensation techniques to address the challenges of depth estimation under heterogeneous camera configurations. It enables flexible and accurate depth information estimation in environments such as automobiles and can be applied to the optimization of various automotive control systems.
Patent Information
- Application Number
- CN202510402296.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2025-02-11
- Filing Date
- 2025-04-01
- Publication Date
- 2025-11-14
AI Technical Summary
Existing depth estimation systems struggle to effectively address the geometric, rotational, and positional misalignment issues in non-overlapping imaging when using heterogeneous camera configurations. This makes it difficult to establish consistent pixel correspondences between images, resulting in poor flexibility, especially in spatially constrained environments.
A combination of pinhole and fisheye cameras is used, and image scaling and rotation compensation are performed through processing circuitry. Epipolar constraints and pixel mapping information are utilized to estimate the depth information of entities. The processing circuitry records pixel correspondences through a mapping table, aligns the image viewpoint using rotation compensation parameters, and calculates depth based on the camera focal length and position offset.
It enables flexible estimation of depth information under heterogeneous camera configurations, improving the accuracy and efficiency of depth estimation. It can be applied to the optimization of various functions in automotive control systems, such as adjusting steering wheel position, seat position, air conditioning system, and airbag deployment.
Smart Images

Figure CN120953343A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to image analysis, and more particularly to a depth estimation system and method thereof. Background Technology
[0002] Depth estimation is a key technology used in various fields, including autonomous driving, driver assistance systems, robotics, and augmented reality. Accurate depth information enables systems to perceive the 3D structure of a scene, supporting applications such as obstacle avoidance, spatial navigation, and object tracking. Depth estimation systems typically use cameras to extract differences or mappings between images captured from different viewpoints, and then apply optical geometry to compute the corresponding 3D information in space.
[0003] In traditional depth estimation systems using stereo cameras, the two cameras are typically of the same type and share an overlapping field of view (FoV). For example, two pinhole cameras or two fisheye cameras can be mounted together to simplify the depth calculation process. These systems rely on well-calibrated configurations where the cameras are nearly coplanar and rotationally aligned. However, strict alignment requirements often limit the flexibility of camera placement, especially in environments with limited installation space or angles.
[0004] For example, in automotive cockpit applications, depth estimation systems can be used to monitor the position of occupants (such as the driver or passengers) to improve safety and enhance driver assistance features. In this case, overlapping fields of view from cameras are used to extract depth information of objects or individuals within the cockpit. However, due to cabin design and space constraints, relying on the same cameras with a strictly aligned configuration can be challenging.
[0005] Recent implementations have used pinhole cameras and fisheye cameras to fulfill different functional purposes beyond depth estimation. For example, in automotive cockpit applications, pinhole cameras are used in Driver Monitoring Systems (DMS) to focus on capturing high-resolution, narrow-view images of the driver's facial features, while fisheye cameras are used in Occupant Monitoring Systems (OMS) to provide wide-view coverage of the entire cockpit for occupant detection or activity monitoring.
[0006] While such heterogeneous camera configurations are commonly implemented in modern systems for various applications, their use in depth estimation is not straightforward. The inherent differences between heterogeneous cameras pose significant challenges to depth estimation. Specifically, unlike traditional stereo camera setups that use identical cameras with the same overlapping field of view (FoV) and aligned mounting, heterogeneous camera systems involve cameras with different FoVs, different optical properties, and are typically mounted at different angles and locations. These differences introduce complex problems such as non-overlapping imaging geometry, rotation and positional misalignment, and the difficulty in establishing consistent pixel correspondences between images.
[0007] Therefore, there is a need for a depth estimation system and method that can flexibly estimate depth information using heterogeneous camera configurations while addressing the challenges posed by heterogeneous camera configurations. Summary of the Invention
[0008] An embodiment of the present invention provides a depth estimation system. The depth estimation system includes a pinhole camera, a fisheye camera, and processing circuitry. The pinhole camera is configured to capture a narrow field-of-view image of an entity. The fisheye camera is configured to capture a wide field-of-view image of the entity. An orientation deviation and a positional offset exist between the pinhole camera and the fisheye camera. The processing circuitry is configured to reduce the narrow field-of-view image to produce a scaled image, rotate the wide field-of-view image using a rotation compensation parameter determined based on the orientation deviation, determine pixel mapping information between the rotated and scaled images based on an epipolar constraint, and estimate the depth information of the entity relative to the pinhole camera based on the pixel mapping information and the positional offset.
[0009] In one embodiment, the epipolar constraint for each pixel of the scaled image is determined based on orientation deviation and position offset, and is stored in a mapping table that records a correspondence between each pixel of the scaled image and its epipolar constraint. The processing circuit is further configured to extract the epipolar constraint from the mapping table based on the correspondence recorded for each pixel of the scaled image.
[0010] In one embodiment, the epipolar constraint is defined by a set of coefficients of an epipolar line. The processing circuit is further configured to determine pixel mapping information between the rotated and scaled images based on the epipolar constraint, and includes the following steps: extracting feature values of a target pixel in the scaled image and multiple candidate pixels along the epipolar line in the rotated image; comparing the feature values of the target pixel and the candidate pixels to determine a similarity score for each candidate pixel; and selecting the candidate pixel with the highest similarity score as a corresponding pixel of the target pixel, thereby determining a mapping between the target pixel and the corresponding pixel. The feature values are determined based on the pixel intensity within a predefined neighborhood of each pixel.
[0011] In one embodiment, the processing circuitry is further configured to reduce a first scale of the narrow field-of-view image to produce a scaled image having a second scale that is substantially smaller than the first scale.
[0012] In one embodiment, the second scale is determined based on a comparison of the size of an entity in the wide-field-of-view image and the size of an entity in the narrow-field-of-view image. The entities in the wide-field-of-view image are approximately the same size as the entities in the scaled image.
[0013] In one embodiment, the processing circuitry is further configured to reduce the narrow field-of-view image using a scaling factor. The scaling factor is determined based on the focal length of the pinhole camera and the fisheye camera.
[0014] In one embodiment, the depth estimation system further includes a volatile memory. In an initial phase, the processing circuitry is further configured to allocate a first contiguous segment of the volatile memory and initialize the first contiguous segment of the volatile memory with a predefined value. In an online phase, the first contiguous segment of the volatile memory is used to rotate the wide-field-of-view image.
[0015] In one embodiment, in the initial stage, the processing circuit is further configured to allocate a second consecutive segment, a third consecutive segment, and a fourth consecutive segment of volatile memory; and in the online stage, the processing circuit is further configured to use the second consecutive segment of volatile memory to narrow the field of view image, use the third consecutive segment of volatile memory to determine pixel mapping information, and use the fourth consecutive segment of volatile memory to estimate the depth information of the entity relative to the pinhole camera.
[0016] In other embodiments, during the initial phase, the processing circuitry is further configured to allocate a first contiguous segment of volatile memory and initialize the first contiguous segment of volatile memory with a predefined value. During the online phase, the processing circuitry is further configured to use the first contiguous segment of volatile memory to store a wide-field-of-view image, use a first independent segment of volatile memory to rotate the wide-field-of-view image, and overwrite the wide-field-of-view image in the first contiguous segment with the rotated image.
[0017] In other embodiments, during the initial phase, the processing circuitry is further configured to allocate a second contiguous segment of volatile memory. During the online phase, the processing circuitry is further configured to use the second contiguous segment of volatile memory to store a narrow field-of-view image, use a second independent segment of volatile memory to reduce the narrow field-of-view image to produce a scaled image, and overwrite the narrow field-of-view image in the second contiguous segment with the scaled image.
[0018] In one embodiment, the processing circuitry is further configured to correct distortion in the narrow field-of-view image before shrinking it using a first distortion coefficient, and to correct distortion in the wide field-of-view image before rotating it using a second distortion coefficient.
[0019] In one embodiment, the processing circuitry is further configured to align the tones of the scaled and rotated images by performing grayscale conversion on the scaled and rotated images.
[0020] In one embodiment, the directional deviation is represented in either an Euler angle format or a quaternion format.
[0021] In one embodiment, the processing circuitry is further configured to identify a target region in the rotated image and to determine pixel mapping information between the target region and the scaled image based on epipolar constraints. The target region is a region encompassing a facial area of an entity within the rotated image.
[0022] In one embodiment, the processing circuitry is further configured to use the estimated depth information of the entity relative to the pinhole camera to adjust an operating parameter of a vehicle control system. In a further embodiment, the vehicle control system includes at least one of an eye-tracking-based instrument panel display system, a driver attention warning system, a steering wheel adjustment system, a seat position adjustment system, an air conditioning system, a head-up display system, or an airbag deployment system.
[0023] An embodiment of the present invention provides a depth estimation method. The depth estimation method is executed by a processing circuit to estimate depth information of an entity relative to a pinhole camera based on a narrow-field-of-view image and a wide-field-of-view image of an entity. The narrow-field-of-view image and the wide-field-of-view image are captured by a pinhole camera and a fisheye camera, respectively, and there is an orientation deviation and a positional offset between the pinhole camera and the fisheye camera. The depth estimation method includes a step of reducing the narrow-field-of-view image to produce a scaled image; a step of rotating the wide-field-of-view image using a rotation compensation parameter to produce a rotated image, the rotation compensation parameter being determined based on the orientation deviation; a step of determining pixel mapping information between the rotated image and the scaled image based on a pair of polar constraints; and a step of estimating the depth information of the entity relative to the pinhole camera based on the pixel mapping information and the positional offset.
[0024] In one embodiment, the epipolar constraint for each pixel of the scaled image is determined based on orientation deviation and position offset, and is stored in a mapping table. The mapping table records a correspondence between each pixel of the scaled image and its epipolar constraint. The step of determining pixel mapping information further includes extracting the epipolar constraint from the mapping table based on the correspondence recorded in the mapping table for each pixel of the scaled image.
[0025] In one embodiment, the epipolar constraint is defined by a set of coefficients of an epipolar line. The step of determining pixel mapping information between the rotated image and the scaled image further includes extracting feature values of a target pixel in the scaled image and multiple candidate pixels along the epipolar line in the rotated image; comparing the feature values of the target pixel and the candidate pixels to determine a similarity score for each candidate pixel; and selecting one of the candidate pixels with the highest similarity score as a corresponding pixel of the target pixel, thereby determining a mapping between the target pixel and the corresponding pixel.
[0026] In one embodiment, the step of reducing the narrow field of view image further includes reducing a first scale of the narrow field of view image to produce a scaled image having a second scale that is substantially smaller than the first scale.
[0027] In one embodiment, the second scale is determined based on a comparison of the size of an entity in the wide-field-of-view image with the size of an entity in the narrow-field-of-view image. The entities in the wide-field-of-view image are approximately the same size as the entities in the scaled image.
[0028] In one embodiment, the step of reducing the narrow field of view image further includes using a scaling factor to reduce the narrow field of view image, the scaling factor being determined based on the focal length of the pinhole camera and the fisheye camera.
[0029] In one embodiment, the depth estimation method further includes an initialization phase and an online phase. In the initialization phase, a first contiguous segment of volatile memory is allocated and initialized with a predefined value. In the online phase, the first contiguous segment of volatile memory is used to rotate a wide-field-of-view image.
[0030] In one embodiment, during the initial phase, a second consecutive segment, a third consecutive segment, and a fourth consecutive segment of volatile memory are allocated. During the online phase, the second consecutive segment of volatile memory is used to narrow the field of view image, the third consecutive segment of volatile memory is used to determine pixel mapping information, and the fourth consecutive segment of volatile memory is used to estimate the depth information of entities relative to the pinhole camera.
[0031] In other embodiments, during the initial phase, a first contiguous segment of volatile memory is allocated and initialized with a predefined value. During the online phase, the first contiguous segment of volatile memory is used to store a wide-field-of-view image, a first independent segment of volatile memory is used to rotate the wide-field-of-view image, and the rotated image is used to overwrite the wide-field-of-view image in the first contiguous segment.
[0032] In other embodiments, during the initial phase, a second contiguous segment of volatile memory is allocated. During the online phase, the second contiguous segment of volatile memory is used to store a narrow field-of-view image, a second independent segment of volatile memory is used to reduce the narrow field-of-view image to produce a scaled image, and the scaled image is used to overwrite the narrow field-of-view image in the second contiguous segment.
[0033] In one embodiment, the depth estimation method further includes correcting distortion in the narrow field-of-view image using a set of first distortion coefficients before scaling down the narrow field-of-view image, and correcting distortion in the wide field-of-view image using a set of second distortion coefficients before rotating the wide field-of-view image.
[0034] In one embodiment, the depth estimation method further includes aligning the tones of the scaled and rotated images by performing grayscale conversion on the scaled and rotated images.
[0035] In one embodiment, the directional deviation is represented in either an Euler angle format or a quaternion format.
[0036] In one embodiment, the depth estimation method further includes identifying a target region in the rotated image and determining pixel mapping information between the target region and the scaled image based on epipolar constraints. The target region is a region containing the facial extent of an entity within the rotated image.
[0037] In one embodiment, the depth estimation method further includes using the estimated depth information of the entity relative to the pinhole camera to adjust an operating parameter of a vehicle control system. In a further embodiment, the vehicle control system includes at least one of an eye-tracking-based instrument panel display system, a driver attention warning system, a steering wheel adjustment system, a seat position adjustment system, an air conditioning system, a head-up display system, or an airbag deployment system. Attached Figure Description
[0038] The invention can be more fully understood by reading the following detailed description and embodiments and by referring to the accompanying drawings, wherein:
[0039] Figure 1 This is a schematic diagram of a depth estimation system according to an embodiment of the present disclosure;
[0040] Figure 2 This is a schematic diagram illustrating the positional offset and directional deviation between a pinhole camera and a fisheye camera according to an embodiment of the present disclosure;
[0041] Figure 3 According to one embodiment of this disclosure, by Figure 1 The flowchart shows the depth estimation method performed by the depth estimation system shown.
[0042] Figure 4 This is a schematic diagram illustrating the operation of the adaptive projection step according to an embodiment of the present disclosure;
[0043] Figure 5 This is a schematic diagram illustrating the operation of the viewpoint rotation compensation step according to an embodiment of the present disclosure;
[0044] Figure 6A and Figure 6B This describes the process of deriving rotation compensation parameters during the offline calibration stage, based on an embodiment of the present disclosure.
[0045] Figure 7A This is a flowchart of the scene correlation calculation steps according to an embodiment of the present disclosure;
[0046] Figure 7B This is a schematic diagram of scene correlation calculation steps according to an embodiment of the present disclosure;
[0047] Figure 8 Additional steps are described according to various embodiments of this disclosure, any one or a combination thereof may be included. Figure 3 The depth estimation method shown; and
[0048] Figure 9 This is a schematic diagram of a mapping table according to an embodiment of the present disclosure.
[0049] The reference numerals in the attached figures are explained as follows:
[0050] 10: Depth Estimation System
[0051] 101: Entity
[0052] 13: Processing Circuit
[0053] 11: Pinhole Camera
[0054] 12: Fisheye camera
[0055] 150: In-depth information
[0056] 111, 121: Field of view
[0057] 112: Narrow field of view image
[0058] 122: Wide field of view image
[0059] 201: Scene
[0060] 21: Position Offset
[0061] 22: Directional deviation
[0062] 30: Depth Estimation Methods
[0063] S31: Adaptive Projection Steps
[0064] S32: View Rotation Compensation Steps
[0065] S33: Scene Correlation Calculation Steps
[0066] S34: Depth Estimation Steps
[0067] 301: Scaled image
[0068] 302: Rotated image
[0069] 303: Pixel mapping information
[0070] S701, S702, S703: Steps
[0071] 71: Target pixel
[0072] 72, 92: Polar lines
[0073] 721, 722, 723: Candidate pixels
[0074] 730, 731, 732, 733: Eigenvalues
[0075] 750: Target Area
[0076] S81: Color Correction Steps
[0077] S82: Distortion removal steps
[0078] S83: Target Area Restriction Steps
[0079] 91: pixels
[0080] 900: Mapping Table Detailed Implementation
[0081] The following description is for illustrative purposes only and should not be construed as limiting. The scope of the invention is best determined by referring to the appended claims.
[0082] In each of the following embodiments, the same reference numerals denote the same or similar elements or components.
[0083] The ordinal terms used in the claims, such as "first," "second," "third," etc., are for ease of interpretation only and do not imply any priority relationship between them.
[0084] The descriptions provided for embodiments of the apparatus or system below also apply to embodiments of the method, and vice versa.
[0085] Figure 1 This is a schematic diagram of a depth estimation system 10 according to an embodiment of the present disclosure. Figure 1As shown, the depth estimation system 10 includes a pinhole camera 11, a fisheye camera 12, and a processing circuit 13.
[0086] A pinhole camera is a camera designed to capture high-resolution images with a narrow field of view (typically 50° to 70°), often used to capture detailed and focused images of a specific area or object.
[0087] In contrast, the fisheye camera 12 is designed to capture images with an extremely wide field of view 121, typically ranging from 180° to 220°. While the sensor resolution of a fisheye camera image may be the same as that of a pinhole camera image, the spatial resolution, or pixel density, of a fisheye camera is typically lower than that of a pinhole camera. This is because the wider field of view (FoV) of a fisheye camera distributes the same number of pixels over a larger area, thus reducing the amount of detail captured per unit area. However, the narrower field of view of a pinhole camera provides higher spatial resolution, allowing each pixel to capture more detail from a smaller portion of the scene. Therefore, the wide field of view of a fisheye camera enables the significant coverage of a larger area in a single frame. This makes the fisheye camera 12 complementary to the pinhole camera 11, as the former captures a wider background, while the latter provides richer detail in a more focused area of the same scene.
[0088] According to embodiments of this disclosure, a pinhole camera 11 and a fisheye camera 12 are used to capture a narrow field-of-view image 112 and a wide field-of-view image 122 of an entity 101, respectively. The narrow field-of-view image 112 provides high-detail information about the entity 101, while the wide field-of-view image 122 contains a broader scene, which may include information about the entity 101 and its additional background.
[0089] The processing circuit 13 may be implemented by a general-purpose processor or a dedicated hardware circuit. In an embodiment where the processing circuit 13 is implemented by a general-purpose processor (e.g., a CPU), the processing circuit 13 receives data from a storage medium (although...). Figure 1 (Not shown) A program or instruction set is loaded to execute a depth estimation method. In another embodiment, the processing circuit 13 is implemented by dedicated hardware circuitry, such as an application-specific integrated circuit (ASIC) or a field-programmable gate array (FPGA), and is configured or programmed to perform the corresponding steps of the depth estimation method.
[0090] The depth estimation method executed by the processing circuit 13 generally involves estimating depth information 150 of entity 101 relative to pinhole camera 11 based on narrow field-of-view image 112 and wide field-of-view image 122 of entity 101. Further details regarding the depth estimation method will be provided below.
[0091] Depth information 150 represents the spatial distance between entity 101 and pinhole camera 11. Depth information 150 may take various forms, such as a depth map providing detailed distance information over the overall image, the distance to a specific feature of entity 101 (e.g., eyes, nose, or other key points), and / or the closest or shortest distance between the entity and the camera, depending on various application requirements, but this disclosure is not limited thereto.
[0092] although Figure 1 The illustration describes an in-vehicle application scenario where entity 101 is driven, but it should be understood that this is merely an example and not a limitation. Embodiments of this disclosure are not limited to automotive applications and can be used in various environments and scenarios, including but not limited to robotics, surveillance, augmented reality, and medical imaging.
[0093] In one embodiment, the processing circuitry 13 is further configured to use the estimated depth information 150 of entity 101 relative to pinhole camera 11 to adjust an operating parameter of a vehicle control system. The vehicle control system includes at least one of an eye-tracking-based dashboard display system, a driver attention alert (DAA) system, a steering wheel adjustment system, a seat position adjustment system, an air conditioning system, a heads-up display system, or an airbag deployment system.
[0094] Specifically, the estimated depth information 150 can be used to adjust the position or angle of the steering wheel to align with the driver's seat position, thereby improving ergonomics and comfort. The estimated depth information 150 can also be used to modify the seat position, such as moving it forward or backward, to ensure the safety and convenience of the driver or passenger. Furthermore, the depth information 150 can be used to regulate the air conditioning system by more effectively guiding airflow based on the occupant's position. For head-up display systems, the depth information 150 can be used to calibrate the display position or adjust the content size to align with the driver's position and / or line of sight. Additionally, in the event of a collision, the depth information 150 can be used to determine the appropriate airbag deployment force based on the proximity of the driver or passenger, thereby enhancing overall safety.
[0095] Furthermore, the estimated depth information 150 can also be used to adjust operating parameters of eye-tracking-based dashboard display systems and / or driver attention alert systems. In eye-tracking-based dashboard display systems, depth information 150 is used to determine the driver's head position and the distance between their eyes and the dashboard. Based on this information, the system can dynamically adjust the brightness, focus, and content scaling of the dashboard display. For example, if the driver is far away, the system may increase the font size and brightness to maintain readability. Conversely, if the driver is closer, the system may reduce the brightness to prevent glare and adjust the focus for optimal clarity. In driver attention alert (DAA) systems, depth information 150 is used to monitor the driver's head position and movement patterns over time. If depth information 150 indicates that the driver's head is consistently tilted downwards or away from the road, the system can activate auditory warnings, visual warnings on the dashboard, or haptic feedback via the steering wheel or seat.
[0096] In practice, besides the differences in optical characteristics, the pinhole camera 11 and the fisheye camera 12 are deployed at different locations and capture images from different viewpoints. In other words, there is a positional offset and an orientation deviation between the pinhole camera 11 and the fisheye camera 12. The positional offset and orientation deviation represent the relative extrinsic parameters of the pinhole camera 11 and the fisheye camera 12. The existence of the positional offset and orientation deviation introduces additional challenges to depth estimation.
[0097] Figure 2 This is a schematic diagram of the positional offset 21 and orientation deviation 22 between a pinhole camera 11 and a fisheye camera 12 according to an embodiment of the present disclosure. Figure 2As shown, position offset 21 represents the spatial displacement between pinhole camera 11 and fisheye camera 12. This displacement indicates that the two cameras are mounted in different physical positions, which may result in different perspectives of the scene 201 (containing entity 101) captured by them. On the other hand, orientation offset 22 represents the angular misalignment between the optical axes of pinhole camera 11 and fisheye camera 12. This angular difference reflects a slight difference in the direction the cameras are pointing, resulting in changes in the field of view and the relative positions of objects in their respective images.
[0098] Although for the sake of simplicity, Figure 2 Position offset 21 and orientation deviation 22 are described in two dimensions, but it should be noted that these differences also occur in three-dimensional space. Position offset 21 may involve displacement along the X, Y, and Z axes, representing differences in the camera's physical configuration, such as lateral movement, vertical height, or forward / backward spacing. Similarly, orientation deviation 22 may involve rotation about the X, Y, and Z axes, such as pitch, yaw, or roll, reflecting angular misalignment between the camera's optical axes.
[0099] In one embodiment, the directional deviation 22 is represented as either an Euler angles format or a quaternion format. The Euler angles format provides an intuitive way to independently describe rotations about three axes, thus directly visualizing each component of the angular misalignment. On the other hand, the quaternion format, widely used in 3D spatial computation, offers advantages such as avoiding gimbal lock and enabling efficient mathematical transformations to align with camera-captured images.
[0100] Figure 3 According to embodiments of this disclosure, by Figure 1 The flowchart shows the depth estimation method 30 performed by the depth estimation system 10. Figure 3 As shown, the depth estimation method 30 includes an adaptive projection step S31, a viewpoint rotation compensation step S32, a scene correlation calculation step S33, and a depth estimation step S34.
[0101] In the adaptive projection step S31, the narrow field-of-view image 112 is reduced to produce a scaled image 301. The adaptive projection step S31 aims to harmonize the ratios of the narrow field-of-view image 112 and the wide field-of-view image 122 to facilitate subsequent correspondence calculations.
[0102] In the viewpoint rotation compensation step S32, a rotation compensation parameter is used to rotate the wide-field-of-view image 122 to produce a rotated image 302. The rotation compensation parameter is determined based on the directional deviation 22 between the pinhole camera 11 and the fisheye camera 12. The purpose of the viewpoint rotation compensation step S32 is to align the viewpoint of the wide-field-of-view image 122 with the viewpoint of the narrow-field-of-view image 112, thereby facilitating subsequent corresponding calculations.
[0103] In the scene correlation calculation step S33, pixel mapping information between the rotated image 302 and the scaled image 301 is determined based on an epipolar constraint. An epipolar constraint is a geometric property used to establish relationships between corresponding points in two images captured by cameras with different viewpoints. The epipolar constraint reduces the search space for pixel matching by imposing a restriction that corresponding points must lie on a specific path defined by the relative positioning of the cameras. By utilizing the epipolar constraint, the scene correlation calculation step S33 identifies the correspondence between pixels in the scaled image 301 and the rotated image 302, thereby enabling the extraction of pixel mapping information 303 as the basis for depth estimation.
[0104] In the depth estimation step S34, depth information 150 of entity 101 relative to pinhole camera 11 is estimated based on pixel mapping information 303 and position offset 21. Specifically, pixel mapping information identifies corresponding points between two images, such as pixels (u) in scaled image 301. k ,v k ) and the corresponding point (x) of that pixel in the rotated image 302 k ,y k Position offset 21 provides the spatial distance between the two cameras. Utilizing the principle of similar triangles, the depth of entity 101 relative to pinhole camera 11 can be calculated by analyzing the geometric relationship between the corresponding pixel coordinates and the relative positions of the multiple cameras, as well as the intrinsic parameters of these cameras (such as focal length).
[0105] Figure 4 This is a schematic diagram of the adaptive projection step S31 according to an embodiment of the present disclosure. Figure 4 As shown, after the adaptive projection step S31, the scale of the narrow field-of-view image 112 (hereinafter referred to as the "first scale") is significantly reduced to produce a scaled image 301 with a scale much smaller than the first scale (hereinafter referred to as the "second scale").
[0106] In a further embodiment, the second scale is determined based on a comparison between the size of entity 101 in the wide-field-of-view image 122 and the size of entity 101 in the narrow-field-of-view image 112. Therefore, entity 101 in the wide-field-of-view image 122 and entity 101 in the scaled image 301 are substantially the same size. For example, the difference in pixel area occupied by entity 101 in the scaled image 301 and entity 101 in the wide-field-of-view image 122 is within a predetermined threshold (e.g., 5%), ensuring that the relative sizes of entities 101 in the two images are sufficiently consistent to facilitate accurate correspondence calculations. This adjustment ensures that subsequent correspondence calculations between the scaled image 301 and the rotated image 302 can be performed more efficiently because geometric differences caused by scale inconsistencies have been minimized.
[0107] In another embodiment, a scaling factor determined based on the focal lengths of the pinhole camera 11 and the fisheye camera 12 is used to reduce the narrow field-of-view image. Specifically, the scaling factor can be proportional to the ratio between the focal length of the fisheye camera 12 and the focal length of the pinhole camera 11. This ensures that the scaled image 301 reflects the same relative proportions as the wide field-of-view image 122 captured by the fisheye camera 12, thereby promoting consistent geometrical alignment between the two images.
[0108] Figure 5 This is a schematic diagram illustrating the operation of the perspective rotation compensation step S32 according to an embodiment of the present disclosure. Figure 5 As shown, due to the orientation deviation 22 between the pinhole camera 11 and the fisheye camera 12, entity 101 in the wide-field-of-view image 122 exhibits different orientation and angular alignment compared to its representation in the scaled image 301. This misalignment leads to inconsistent geometric relationships of identical features of entity 101, making feature matching more challenging. This problem is particularly evident in automotive applications, where objects (including entity 101) are typically captured at close range, resulting in features occupying a large number of pixels. Any change in orientation introduces significant inconsistencies, which can substantially reduce the accuracy of feature correspondence.
[0109] After the viewpoint rotation compensation step S32, entity 101 in the rotated image 302 is corrected to a viewpoint aligned with that of the scaled image 301 (or the narrow field-of-view image 112). This correction minimizes rotational differences and ensures that the features of entity 101 are geometrically consistent in both images, thereby significantly improving the accuracy and efficiency of subsequent correspondence calculations.
[0110] As previously mentioned, the rotation compensation parameters used in the viewpoint rotation compensation step S32 are determined based on the orientation deviation 22, which represents the relative external parameters of the pinhole camera 11 and the fisheye camera 12. The external parameters of the pinhole camera 11 and the fisheye camera 12 can be calibrated offline using standard calibration techniques. This calibration process allows for the pre-calculation and storage of the rotation compensation parameters, thereby eliminating the need for the processing circuit 13 to perform complex calculations to derive the parameters during the online phase.
[0111] Figure 6A and Figure 6B The process of deriving rotation compensation parameters in an offline calibration stage is described according to embodiments of this disclosure. These rotation compensation parameters are later used in the online view rotation compensation step S32 to correct the wide-field-of-view image 122.
[0112] like Figure 6A As shown, a target plane S in the world coordinate system w (For example, a calibration board) is projected onto the respective camera coordinate systems of the two cameras. These projections result in two camera-fixed geometric planes, S1 and S2, reflecting the relative positions and orientations of the cameras. These planes are further transformed using the corresponding camera intrinsic parameters to produce image planes I1 and I2. However, due to directional deviations between the two cameras, S1 and S2, and their respective image planes I1 and I2, are not geometrically aligned.
[0113] Figure 6B This demonstrates the geometric alignment achieved through a rotation compensation process. As shown, plane S1 is calculated and adjusted during the offline calibration phase to produce a plane S1' that is aligned with plane S2 in the second camera coordinate system after correction. This adjustment ensures that when S1' is transformed using the intrinsic parameters of the first camera, the resulting image plane I1' is parallel to I2, thus solving the problem of... Figure 6A The misalignment shown.
[0114] The derived relation is I1′=K1RK1 -1 I1, where R represents the directional deviation, this relationship establishes the rotation compensation parameter K1RK1 -1 The derived rotation compensation parameters are pre-calculated during the calibration process and then applied during the online phase to align multiple images captured by the two cameras with directional deviations.
[0115] In traditional pixel matching methods for stereo cameras, optical properties are used to correct the corresponding epipolar lines of all pixels, ensuring that identical pixels from different viewpoints lie on the same horizontal line. This facilitates an algorithmic search for pixel correspondences. However, in the embodiments of this disclosure, due to the orientation deviation 22 between the pinhole camera 11 and the fisheye camera 12, applying traditional epipolar line correction methods would result in significant feature corruption, thus hindering pixel matching. Therefore, this paper proposes a solution that utilizes pre-calibrated intrinsic and extrinsic parameters of the camera, obtainable offline, to determine the relationship between pixels at different viewpoints and their corresponding tilted epipolar lines. Matching pixels between scaled image 301 and rotated image 302 are then identified along these tilted epipolar lines. This method simplifies the traditional epipolar line correction process while avoiding the feature corruption problem caused by applying traditional correction algorithms to heterogeneous camera configurations.
[0116] Figure 7A This is a flowchart of the scene correlation calculation step S33 according to an embodiment of the present disclosure. In this embodiment, the polar constraint is defined by the coefficient set of an epipolar line, for example, (a,b,c) for the linear equation ax+by+c=0. Figure 7A As shown, the scene correlation calculation step S33 may also include more detailed steps S701-S703. Figure 7B This is a schematic diagram of the scene correlation calculation step S33. Figure 7A and Figure 7B This can be referenced together to better understand this embodiment.
[0117] In step S701, feature values of a target pixel are extracted from the scaled image, and feature values of candidate pixels along the epipolar line are extracted from the rotated image. The feature values are determined based on the pixel intensity within a predefined neighborhood of each pixel, such as the average intensity, variance, or gradient magnitude calculated in a 3×3 or 5×5 pixel window. Figure 7B In the example shown, feature value 730 is extracted from target pixel 71 in the scaled image 301. Meanwhile, in the rotated image 302, feature values 731, 732, and 733 are extracted along the epipolar line 72 from candidate pixels 721, 722, and 723, respectively.
[0118] In step S702, the feature values of the target pixel and the candidate pixels are compared to determine similarity. The comparison involves calculating a similarity score using predefined metrics, such as normalized cross-correlation (NCC), mean squared error (MSE), or cosine similarity. Figure 7B As shown, the feature value 730 of the target pixel 71 is compared with the feature values 731, 732 and 733 of the candidate pixels 721, 722 and 723 along the epipolar line 72, respectively, to generate the corresponding similarity scores.
[0119] In step S703, the candidate pixel with the highest similarity score is selected as the corresponding pixel of the target pixel, thus determining the mapping between the target pixel and the corresponding pixel. This mapping establishes a correspondence between the two images, allowing for further depth estimation calculations. Figure 7B As shown, among candidate pixels 721, 722 and 723, candidate pixel 722 is identified as the best match for target pixel 71 because it has the highest similarity score.
[0120] Figure 8 According to embodiments of this disclosure, additional color correction step S81, view rotation compensation step S82, and target area limitation step S83 may be included, any one or a combination thereof. Figure 3 The depth estimation method 30 is shown below. These steps will be described in detail below.
[0121] In one embodiment, the depth estimation method 30 may further include a color correction step S81 prior to the adaptive projection step S31 and a color correction step S81 prior to the viewpoint rotation compensation step S32. The color correction step S81 involves aligning the hues of the narrow-field image 112 and the wide-field image 122 to reduce inconsistencies between the two images. This alignment can be achieved through various methods, such as histogram matching, color balance adjustment, or white balance correction. By minimizing color differences, the computational complexity of subsequent steps (such as the scene relevance calculation step S33) can be significantly reduced. In one embodiment, alignment can be achieved by converting the hues of the narrow-field image 112 and the wide-field image 122 to grayscale. This method not only simplifies the color correction process but also standardizes the intensity values, making feature extraction and matching more efficient.
[0122] In a preferred alternative embodiment, the color correction step S81 is performed after the adaptive projection step S31 and the view rotation compensation step S32, instead of as... Figure 8The process is performed before the adaptive projection step S31 and the viewpoint rotation compensation step S32. In other words, instead of aligning the tones of the narrow-field image 112 and the wide-field image 122, the color correction step S81 is applied to the scaled image 301 and the rotated image 302. By delaying the color adjustment to a later stage, more detail is preserved for subsequent processing, thus avoiding damage from subtle changes in intensity and contrast, which helps with feature extraction and matching in the scene relevance calculation step S33.
[0123] In one embodiment, the depth estimation method 30 may further include a distortion correction step S82 prior to the adaptive projection step S31 and the viewpoint rotation compensation step S32. The distortion correction step S82 involves using a first set of distortion coefficients and a second set of distortion coefficients to correct distortions in the narrow field-of-view image 112 and the wide field-of-view image 122, respectively. The first set of distortion coefficients corresponds to the pinhole camera 11 and typically includes parameters describing radial and tangential distortions specific to the lens characteristics of the pinhole camera 11. The second set of distortion coefficients corresponds to the fisheye camera 12 and takes into account the extreme radial distortion introduced by the fisheye lens. Correcting these distortions prior to the adaptive projection step S31 and the viewpoint rotation compensation step S32 ensures that the images are geometrically corrected, thereby facilitating more accurate calculation of scene correspondences in subsequent steps.
[0124] In one embodiment, the depth estimation method 30 may further include a target region limitation step S83. The target region limitation step S83 involves identifying a target region 750 in the rotated image 302, the target region 750 being a region encompassing the facial extent of an entity within the rotated image 302. For example... Figure 7B As shown, a target region 750 containing the facial area of entity 101 is identified in the rotated image 302. The target region limitation step S83 enables the subsequent scene correlation calculation step S33 to focus only on determining the pixel mapping information between the target region 750 and the scaled image 301, without processing the complete rotated image 302, thereby improving computational efficiency.
[0125] In the online phase of the target region restriction step S83, the target region 750 can be identified by calculating a bounding box around the entity using an object recognition or face recognition algorithm. Alternatively, the target region can be predefined as a fixed region, such as the overlapping field of view between the pinhole camera 11 and the fisheye camera 12, in the offline phase. Furthermore, the two methods for identifying the target region 750 can be implemented simultaneously, allowing the system to dynamically identify the target region in the online phase using an on-the-fly object recognition or face recognition algorithm, while simultaneously utilizing a predefined fixed region from the offline phase, such as the overlapping field of view between the pinhole camera and the fisheye camera. This dual approach ensures greater flexibility and accuracy in target region identification. In in-vehicle applications, entities are typically driven, and their positions tend to remain relatively stable. Therefore, predefining the target region in the offline phase may be particularly effective, reducing the need for on-the-fly computation and further optimizing overall system performance.
[0126] In some embodiments, epipolar constraints are calculated during the online phase as part of the scene relevance calculation step S33. However, since the epipolar constraints for each pixel are determined entirely based on camera parameters (e.g., directional deviation and positional offset between pinhole camera 11 and fisheye camera 12) rather than on the actual image content, these constraints can be pre-calculated during the offline phase. Therefore, in some other embodiments, the epipolar constraints for each pixel of the scaled image 301 are pre-calculated and stored in a mapping table during the offline phase. During the online phase, the scene relevance calculation step S33 directly extracts the pre-calculated epipolar constraints from the mapping table, thereby significantly reducing the computational burden. Specifically, the epipolar constraints for each pixel of the scaled image are determined based on known directional deviations and positional offsets, and this correspondence is recorded in the mapping table. By storing this mapping, the system avoids on-the-fly recalculation of epipolar constraints, thereby simplifying the online phase.
[0127] Figure 9 This is a schematic diagram of a mapping table 900 according to an embodiment of the present disclosure. Figure 9 As shown, mapping table 900 records the correspondence between each pixel of the scaled image 301 and its associated epipolar constraint. For example, pixel 91 at coordinates (u,v) in the scaled image 301 corresponds to an epipolar line 92 in the target region 750 of the rotated image. This correspondence is pre-calculated and stored in mapping table 900.
[0128] During the deployment phase, the scene relevance calculation step S33 extracts the epipolar constraint of pixel 91 from the mapping table 900. The mapping table provides a set of coefficients for the epipolar line 92, enabling the system to immediately use this information for feature extraction and pixel matching without performing additional geometric calculations.
[0129] It should be noted that, although Figure 9The mapping of a single pixel 91 is shown, but the mapping table 900 is designed to cover all pixels in the scaled image 301, providing a complete lookup resource for epipolar constraints.
[0130] The pre-computed mapping table 900 not only accelerates real-time processing but also enhances the consistency and reliability of the depth estimation process. This is particularly useful in applications where computational efficiency is critical, such as in-vehicle monitoring systems used for depth estimation of entities like a driver.
[0131] like Figure 8 As shown, the depth estimation method may include several steps, such as color correction step S81, distortion correction step S82, adaptive projection step S31, viewpoint rotation compensation step S32, target region limitation step S83, scene correlation calculation step S33, and depth estimation step S34. These steps are repeatedly executed by sequentially inputting narrow-field-of-view image sequences and wide-field-of-view image sequences. Furthermore, each of these steps involves processing intermediate data that occupies memory blocks. In a depth estimation system employing a conventional memory allocation method, the memory block location occupied by the intermediate data generated in each step is not fixed, and the time for releasing the memory block after each step is unpredictable. When the rate of memory block release cannot keep up with the processing demands of the depth estimation process, memory shortages may occur.
[0132] For example, when the depth estimation process reaches the view rotation compensation step S32 of the k-th frame of the wide-field-of-view image 122 (a step requiring a relatively large amount of memory), a memory shortage may occur because the memory block occupied by the adaptive projection step S31 of the (k-1)-th frame of the narrow-field-of-view image sequence has not yet been released. This memory shortage will subsequently interrupt the progress of the process, potentially leading to processing delays or even system failures. To prevent such problems and ensure the smooth execution of the depth estimation process, this paper proposes an optimized memory allocation strategy.
[0133] In one embodiment, the depth estimation system 10 further includes volatile memory. Furthermore, in an initial phase, the viewpoint rotation compensation step S32 pre-allocates a contiguous segment of the volatile memory. This contiguous segment, hereinafter referred to as the first contiguous segment of the volatile memory, consists of multiple contiguous memory blocks. The viewpoint rotation compensation step S32 assigns an index to indicate the starting position of the first contiguous segment of the volatile memory. Subsequently, the processing circuit 13 initializes the first contiguous segment of the volatile memory with a predefined value (e.g., 0 or -1) to clear any residual data from previous operations. Therefore, in the initial phase of the viewpoint rotation compensation step S32, the pre-allocated first contiguous segment of the volatile memory can be used to perform rotation of the wide-field-of-view image 122 without worrying about insufficient memory.
[0134] In a further embodiment, in addition to the viewpoint rotation compensation step S32, other steps in the depth estimation process can also pre-allocate corresponding memory space in the initial stage. Specifically, in the initial stage, the second, third, and fourth consecutive segments of the volatile memory can be allocated to the adaptive projection step S31, the scene correlation calculation step S33, and the depth estimation step S34, respectively. Therefore, in the online stage, the second consecutive segment of the volatile memory can be used to reduce the narrow field-of-view image 112, the third consecutive segment can be used to determine pixel mapping information, and the fourth consecutive segment can be used to estimate depth information 150, without worrying about insufficient memory.
[0135] Similarly, other steps (such as color correction step S81, distortion correction step S82, and target area limitation step S83) can also allocate dedicated memory portions in the initial stage. By pre-allocating memory space for each step, the system can minimize runtime latency caused by dynamic memory allocation and ensure the smooth progress of the depth estimation process. This method not only improves the overall efficiency of memory management but also reduces the risk of memory contention during the online phase.
[0136] It should be understood that the terms "first contiguous segment," "second contiguous segment," "third contiguous segment," and "fourth contiguous segment" are used only to distinguish different allocated memory spaces and do not imply any particular order in the physical arrangement or allocation sequence. These names are only used for reference and to clearly describe how different steps in the depth estimation process utilize different portions of volatile memory.
[0137] In an alternative embodiment, the depth estimation system 10 further includes a volatile memory. To optimize memory management and prevent allocation delays during runtime, the processing circuitry 13 is configured to pre-allocate a first contiguous segment of the volatile memory in an initial phase. This first contiguous segment, consisting of multiple contiguous memory blocks, serves as a buffer to temporarily store the wide-field-of-view image 122. Furthermore, to ensure data integrity and avoid unintended residual effects from previous operations, the processing circuitry 13 initializes the first contiguous segment of the volatile memory with predefined values (e.g., 0 or -1). During the online phase, when a new wide-field-of-view image 122 is captured, the processing circuitry 13 stores the wide-field-of-view image 122 in the pre-allocated first contiguous segment of the volatile memory. To perform image rotation, the processing circuitry 13 further uses a first independent segment of the volatile memory to process the rotation operation independently. Once the rotation is complete, the rotated image is written back to the first contiguous segment of the volatile memory, efficiently overwriting the initially stored wide-field-of-view image 122.
[0138] In a further embodiment, in an initial stage, the processing circuit 13 is configured to allocate a second contiguous segment of volatile memory. This second contiguous segment, consisting of multiple contiguous memory blocks, serves as a buffer to temporarily store the narrow field-of-view image 112. In the online stage, the processing circuit 13 uses the allocated second contiguous segment of volatile memory to store the narrow field-of-view image 112 during capture. To perform a scaling down operation, the processing circuit 13 further uses a second independent segment of volatile memory to generate a scaled image 301 from the stored narrow field-of-view image 112. After the resizing operation is completed, the scaled image 301 is written back to the second contiguous segment of volatile memory, effectively overwriting the original narrow field-of-view image 112.
[0139] The preceding paragraphs describe the subject from multiple perspectives. Clearly, the teachings of this specification can be performed in various ways. Any particular structure or function disclosed in the examples is merely representative. Those skilled in the art should note, based on the teachings of this specification, that any aspect disclosed can be performed alone, or in combination with other aspects.
[0140] While the invention has been described by way of example and preferred embodiments, it should be understood that the invention is not limited to the disclosed embodiments. Rather, the invention is intended to cover various modifications and similar arrangements (which will be apparent to those skilled in the art). Therefore, the appended claims should be given the broadest interpretation to cover all such modifications and similar arrangements.
Claims
1. A depth estimation system, comprising: A pinhole camera configured to capture a narrow field-of-view image of an entity; A fisheye camera is configured to capture a wide field-of-view image of the entity, wherein there is an orientation deviation and a positional offset between the pinhole camera and the fisheye camera; A processing circuit is configured as follows: Reduce the narrow field-of-view image to produce a scaled image; The wide-field-of-view image is rotated using a rotation compensation parameter to produce a rotated image, wherein the rotation compensation parameter is determined based on the orientation deviation; The one-pixel mapping information between the rotated image and the scaled image is determined based on a pair of polar constraints; as well as Based on the pixel mapping information and the position offset, a depth information of the entity relative to the pinhole camera is estimated.
2. The depth estimation system of claim 1, wherein the epipolar constraint for each pixel of the scaled image is determined based on the orientation deviation and the positional offset, and is stored in a mapping table, wherein the mapping table records a correspondence between each pixel of the scaled image and the epipolar constraint; and The processing circuit is further configured to extract the epipolar constraint from the mapping table based on the correspondence of each pixel of the scaled image recorded in the mapping table.
3. The depth estimation system of claim 1, wherein the epipolar constraint is defined by a set of coefficients of an epipolar line; and The processing circuit is further configured to perform the following steps to determine the pixel mapping information between the rotated image and the scaled image based on the epipolar constraint: Extract feature values of a target pixel in the scaled image and multiple candidate pixels along the epipolar line in the rotated image, wherein the feature values are determined based on the pixel intensity in a predefined neighborhood of each pixel; The feature value of the target pixel is compared with that of the candidate pixel to determine a similarity score for each candidate pixel; as well as The candidate pixel with the highest similarity score is selected as the corresponding pixel of the target pixel, thereby determining a mapping between the target pixel and the corresponding pixel.
4. The depth estimation system of claim 1, wherein the processing circuitry is further configured to reduce a first scale of the narrow field-of-view image to produce the scaled image having a second scale substantially smaller than the first scale.
5. The depth estimation system of claim 4, wherein the second scale is determined based on a comparison of a size of the entity in the wide-field-of-view image and a size of the entity in the narrow-field-of-view image; and The entity in the wide-view image is approximately the same size as the entity in the scaled image.
6. The depth estimation system of claim 1, wherein the processing circuitry is further configured to reduce the narrow field-of-view image using a scaling factor, wherein the scaling factor is determined based on the focal lengths of the pinhole camera and the fisheye camera.
7. The depth estimation system of claim 1, further comprising a volatile memory, wherein the processing circuitry is further configured as follows: In an initial phase, a first contiguous segment of the volatile memory is allocated, and the first contiguous segment of the volatile memory is initialized using a predefined value; and During the initial online phase, the first contiguous segment of the volatile memory is used to rotate the wide-view image.
8. The depth estimation system of claim 7, wherein in the initial stage, the processing circuitry is further configured to allocate a second contiguous segment, a third contiguous segment, and a fourth contiguous segment of the volatile memory; and During this online phase, the processing circuit is further configured as follows: The second contiguous segment of the volatile memory is used to narrow the field of view image; The third contiguous segment of the volatile memory is used to determine the pixel mapping information; and The fourth contiguous segment of the volatile memory is used to estimate the depth information of the entity relative to the pinhole camera.
9. The depth estimation system of claim 1, further comprising a volatile memory, wherein the processing circuitry is further configured as follows: In an initial phase, a first contiguous segment of the volatile memory is allocated, and the first contiguous segment of the volatile memory is initialized with a predefined value; as well as During the initial launch phase: The first contiguous segment of the volatile memory is used to store the wide-field-of-view image; A first independent segment of the volatile memory is used to rotate the wide-view image; The rotated image is used to overwrite the wide-field image in the first continuous segment.
10. The depth estimation system of claim 9, wherein in the initial stage, the processing circuitry is further configured to allocate a second contiguous segment of the volatile memory; and During this online phase, the processing circuit is further configured as follows: The second contiguous segment of the volatile memory is used to store the narrow field-of-view image; A second independent segment of the volatile memory is used to reduce the narrow field-of-view image to produce the scaled image; as well as The scaled image is used to overwrite the narrow field-of-view image in the second consecutive segment.
11. The depth estimation system of claim 1, wherein the processing circuitry is further configured as follows: Using a first set of distortion coefficients, the distortion in the narrow field-of-view image is corrected before the narrow field-of-view image is reduced; and A second set of distortion coefficients is used to correct distortion in the wide field of view image before rotating it.
12. The depth estimation system of claim 1, wherein the processing circuitry is further configured to align the tones of the scaled image and the rotated image by performing grayscale conversion on the scaled image and the rotated image.
13. The depth estimation system of claim 1, wherein the orientation deviation is represented in one of an Euler angle format and a quaternion format.
14. The depth estimation system of claim 1, wherein the processing circuitry is further configured as follows: Identify a target region in the rotated image, wherein the target region is a region encompassing a facial portion of the entity within the rotated image; and The pixel mapping information between the target region and the scaled image is determined based on the epipolar constraint.
15. The depth estimation system of claim 1, wherein the processing circuitry is further configured as follows: The estimated depth information relative to the pinhole camera is used to adjust an operating parameter of an automotive control system.
16. The depth estimation system of claim 15, wherein the vehicle control system includes at least one of an eye-tracking based instrument panel display system, a driver attention warning system, a steering wheel adjustment system, a seat position adjustment system, an air conditioning system, a head-up display system, or an airbag deployment system.
17. A depth estimation method, executed by a processing circuit, for estimating depth information of an entity relative to a pinhole camera based on a narrow field-of-view image and a wide field-of-view image of an entity, wherein the narrow field-of-view image and the wide field-of-view image are captured by the pinhole camera and a fisheye camera, respectively, and wherein there is an orientation deviation and a positional offset between the pinhole camera and the fisheye camera, the depth estimation method comprising the following steps: Reduce the narrow field-of-view image to produce a scaled image; The wide-field-of-view image is rotated using a rotation compensation parameter to produce a rotated image, wherein the rotation compensation parameter is determined based on the orientation deviation; The one-pixel mapping information between the rotated image and the scaled image is determined based on a pair of polar constraints; as well as The depth information of the entity relative to the pinhole camera is estimated based on the pixel mapping information and the position offset.
18. The depth estimation method of claim 17, wherein the epipolar constraint for each pixel of the scaled image is determined based on the orientation deviation and the positional offset, and is stored in a mapping table, wherein the mapping table records a correspondence between each pixel of the scaled image and the epipolar constraint; and The step of determining the pixel mapping information further includes extracting the epipolar constraint from the mapping table based on the correspondence recorded in the mapping table for each pixel of the scaled image.
19. The depth estimation method of claim 17, wherein the epipolar constraint is defined by a set of coefficients of an epipolar line; and The step of determining the pixel mapping information between the rotated image and the scaled image further includes: Extract feature values of a target pixel in the scaled image and multiple candidate pixels along the epipolar line in the rotated image, wherein the feature values are determined based on the pixel intensity in a predefined neighborhood of each pixel; The feature value of the target pixel is compared with that of the candidate pixel to determine a similarity score for each candidate pixel; as well as The candidate pixel with the highest similarity score is selected as the corresponding pixel of the target pixel, thereby determining a mapping between the target pixel and the corresponding pixel.
20. The depth estimation method of claim 17, wherein the step of reducing the narrow field-of-view image further comprises reducing a first scale of the narrow field-of-view image to produce the scaled image having a second scale substantially smaller than the first scale.
21. The depth estimation method of claim 20, wherein the second scale is determined based on a comparison of a size of the entity in the wide-field-of-view image and a size of the entity in the narrow-field-of-view image; and The entity in the wide-view image is approximately the same size as the entity in the scaled image.
22. The depth estimation method of claim 17, wherein the step of reducing the narrow field-of-view image further comprises using a scaling factor to reduce the narrow field-of-view image, wherein the scaling factor is determined based on the focal lengths of the pinhole camera and the fisheye camera.
23. The depth estimation method as described in claim 17, further comprising: In an initial stage, a first contiguous segment of a volatile memory is allocated, and the first contiguous segment of the volatile memory is initialized with a predefined value; as well as During the initial online phase, the first contiguous segment of the volatile memory is used to rotate the wide-view image.
24. The depth estimation method as described in claim 23, further comprising: In this initial stage, a second consecutive segment, a third consecutive segment, and a fourth consecutive segment of the volatile memory are allocated; as well as During this online phase, the second consecutive segment of the volatile memory is used to narrow the field-of-view image, the third consecutive segment of the volatile memory is used to determine the pixel mapping information, and the fourth consecutive segment of the volatile memory is used to estimate the depth information of the entity relative to the pinhole camera.
25. The depth estimation method as described in claim 17, further comprising: In an initial stage, a first contiguous segment of a volatile memory is allocated, and the first contiguous segment of the volatile memory is initialized with a predefined value; as well as During the initial launch phase, the following are included: The first contiguous segment of the volatile memory is used to store the wide-field-of-view image; A first independent segment of the volatile memory is used to rotate the wide-view image; The rotated image is used to overwrite the wide-field image in the first consecutive segment.
26. The depth estimation method as described in claim 25, further comprising: In this initial stage, a second contiguous segment of the volatile memory is allocated; During this launch phase: The second contiguous segment of the volatile memory is used to store the narrow field-of-view image; A second independent segment of the volatile memory is used to reduce the narrow field-of-view image to produce the scaled image; as well as The scaled image is used to overwrite the narrow field-of-view image in the second consecutive segment.
27. The depth estimation method as described in claim 17, further comprising: The distortion in the narrow field-of-view image is corrected before the narrow field-of-view image is reduced using a first set of distortion coefficients. as well as The distortion in the wide-field image is corrected before rotation using a second set of distortion coefficients.
28. The depth estimation method as described in claim 17, further comprising: The grayscale values of the scaled and rotated images are aligned by performing grayscale conversion on both images.
29. The depth estimation method of claim 17, wherein the orientation deviation is represented in either an Euler angle format or a quaternion format.
30. The depth estimation method as described in claim 17, further comprising: Identify a target region in the rotated image, wherein the target region is a region that includes a facial area of the entity within the rotated image; as well as The pixel mapping information between the target region and the scaled image is determined based on the epipolar constraint.
31. The depth estimation method as described in claim 17, further comprising: The estimated depth information relative to the pinhole camera is used to adjust an operating parameter of an automotive control system.
32. The depth estimation method of claim 31, wherein the vehicle control system includes at least one of an eye-tracking-based instrument panel display system, a driver attention warning system, a steering wheel adjustment system, a seat position adjustment system, an air conditioning system, a head-up display system, or an airbag deployment system.