Vehicle control method, device, storage medium and program product
By recognizing and converting the driver's eye position into three-dimensional spatial coordinates using a binocular camera, and combining this with a vehicle model to calculate the safe passage area, the problem of visual disconnect in traditional systems is solved. This enables the display of a safe passage area that matches the driver's visual perspective, reducing the risk of misoperation.
Patent Information
- Application Number
- CN202511990473.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-26
- Publication Date
- 2026-02-06
AI Technical Summary
Traditional driver assistance systems provide auxiliary display information that is out of sync with the driver's actual perspective, resulting in delays in the driver's judgment of road conditions, deviations in spatial perception, and an increased probability of misoperation.
The system captures facial images of the driver using binocular cameras, identifies the two-dimensional position of the midpoint between the two eyes, converts it into three-dimensional spatial coordinates in the vehicle coordinate system, combines the outer contour data of the vehicle's three-dimensional model, calculates the safe passage area through perspective projection, and updates it in real time to match the driver's perspective.
Ensure that the safe passage area is consistent with the driver's subjective perspective to reduce the possibility of misoperation and avoid misjudgment caused by perspective deviation.
Smart Images

Figure CN121469618A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of vehicle control technology, and in particular to a vehicle control method, device, storage medium, and program product. Background Technology
[0002] With the rapid development of intelligent driving technology, drivers' need to perceive the surrounding environment of the vehicle is increasing, especially in complex road scenarios, where drivers need to quickly determine whether the vehicle can safely pass through obstacles or width-restricted facilities.
[0003] Traditional driver assistance systems mostly rely on panoramic imaging, ultrasonic radar, or augmented reality head-up displays (AR-HUD) to provide auxiliary information; however, the auxiliary display information provided by existing methods is disconnected from the driver's real perspective, resulting in delays in the driver's judgment of road conditions and deviations in spatial perception, which increases the probability of driver error. Summary of the Invention
[0004] This application provides a vehicle control method, device, storage medium, and program product to solve the technical problem that the auxiliary display information provided by traditional driver assistance systems is disconnected from the driver's real perspective, resulting in delays in the driver's judgment of road conditions and deviations in spatial perception, which increases the probability of driver error.
[0005] In a first aspect, embodiments of this application provide a vehicle control method, the vehicle including a binocular camera, the method comprising:
[0006] Multiple facial images of the driver were captured using a binocular camera;
[0007] Two-dimensional position of the midpoint between the two eyes is identified based on multiple driver facial images;
[0008] Convert the two-dimensional position of the midpoint between the two eyes into three-dimensional spatial coordinates in the vehicle coordinate system;
[0009] Using three-dimensional spatial coordinates as the origin and combining the outer contour data of the vehicle's three-dimensional model, the safe passage area from the current perspective is calculated through perspective projection.
[0010] Displays the safe passage area.
[0011] In one possible implementation, the binocular camera includes two DMS cameras, which are respectively positioned on both sides of the driver's seat; converting the two-dimensional position of the midpoint between the two eyes into three-dimensional spatial coordinates in the vehicle coordinate system includes:
[0012] Parallax information is obtained based on the two-dimensional position of the midpoint between the two eyes in multiple driver facial images;
[0013] Based on parallax information, the baseline distance and focal length between the two DMS cameras, the three-dimensional spatial coordinates of the midpoint between the two eyes are calculated using triangulation.
[0014] In one possible implementation, before calculating the safe passage area from the current viewpoint using perspective projection and taking the three-dimensional spatial coordinates as the origin, combined with the outer contour data of the vehicle's three-dimensional model, the following steps are also included:
[0015] Real-time detection of driver head posture changes to obtain head posture parameters;
[0016] Dynamically adjust the three-dimensional spatial coordinates based on head pose parameters;
[0017] Real-time monitoring of vehicle attitude information;
[0018] Based on the vehicle's attitude information, the attitude parameters in the outer contour data are dynamically adjusted.
[0019] In one possible implementation, using three-dimensional spatial coordinates as the origin and combining the outer contour data of the vehicle's three-dimensional model, the safe passage area from the current viewpoint is calculated through perspective projection, including:
[0020] Using the adjusted three-dimensional spatial coordinates as the origin, combined with the attitude parameters in the corrected outer contour data and the vehicle's surrounding environment data, the safe passage area is calculated through perspective projection.
[0021] In one possible implementation, the vehicle further includes an elongated display screen disposed on the lower inner side of the windshield; displaying a safe passage area, including:
[0022] The safe passage area is displayed on the long display screen;
[0023] The process of displaying the safe passage area on the long display screen also includes:
[0024] The driver's line of sight is detected based on images acquired by a binocular camera, and the line of sight parameters are obtained.
[0025] Based on the line-of-sight parameter, the display range and visual priority of the elongated display screen are dynamically adjusted; the visual priority is achieved by differentiating key information through color contrast and / or flashing prompts.
[0026] In one possible implementation, the method further includes:
[0027] The image acquisition parameters of the binocular camera are dynamically adjusted according to environmental conditions, including light intensity and rain / fog density, and the image acquisition parameters include camera exposure time and white balance parameters.
[0028] Multiple facial images of the driver were captured using a binocular camera, including:
[0029] Multiple facial images of the driver were acquired based on the adjusted image acquisition parameters of the binocular camera.
[0030] In one possible implementation, the vehicle includes a first computing unit and a second computing unit deployed locally within the vehicle, wherein...
[0031] The first computing unit is used to calculate the parallax information of multiple driver facial images;
[0032] The second computing unit is used to perform steps related to 3D model rendering and perspective projection calculations.
[0033] Secondly, embodiments of this application provide a vehicle control device, the device comprising:
[0034] The acquisition module is used to capture multiple facial images of the driver using a binocular camera.
[0035] The processing module is used to identify the two-dimensional position of the midpoint between the two eyes based on multiple driver facial images.
[0036] The processing module is also used to convert the two-dimensional position of the midpoint between the two eyes into three-dimensional spatial coordinates in the vehicle coordinate system.
[0037] The processing module is also used to calculate the safe passage area from the current perspective by using the three-dimensional spatial coordinates as the origin and combining the outer contour data of the vehicle's three-dimensional model through perspective projection.
[0038] The processing module is also used to display the safe passage area.
[0039] In one possible implementation, the processing module is further configured to obtain parallax information based on the two-dimensional position of the midpoint between the two eyes in multiple driver facial images.
[0040] The processing module is also used to calculate the three-dimensional spatial coordinates of the midpoint between the two eyes using triangulation based on parallax information, the baseline distance between the two DMS cameras, and the focal length.
[0041] In one possible implementation, the acquisition module is also used to detect changes in the driver's head posture in real time and obtain head posture parameters.
[0042] The processing module is also used to dynamically adjust the three-dimensional spatial coordinates based on head posture parameters.
[0043] The acquisition module is also used to monitor the vehicle's attitude information in real time.
[0044] The processing module is also used to dynamically adjust the attitude parameters in the outer contour data based on the vehicle's attitude information.
[0045] In one possible implementation, the processing module is further configured to calculate the safe passage area by perspective projection, using the adjusted three-dimensional spatial coordinates as the origin, combining the attitude parameters in the corrected outer contour data, and the vehicle's surrounding environment data.
[0046] In one possible implementation, the processing module is also configured to display a safe passage area on a long display screen.
[0047] The processing module is also used to detect the driver's line of sight based on images captured by the binocular camera and obtain line of sight parameters.
[0048] The processing module is also used to dynamically adjust the display range and visual priority of the elongated display screen based on the line-of-sight parameter; the visual priority is achieved by differentiating key information through color contrast and / or flashing prompts.
[0049] In one possible implementation, the processing module is further configured to dynamically adjust the image acquisition parameters of the binocular camera according to environmental conditions; the environmental conditions include light intensity and rain / fog density, and the image acquisition parameters include camera exposure time and white balance parameters.
[0050] The acquisition module is also used to acquire multiple driver facial images based on the image acquisition parameters adjusted by the binocular camera.
[0051] Thirdly, embodiments of this application provide an electronic device, including: a memory and a processor; the memory stores computer-executable instructions; the processor executes the computer-executable instructions stored in the memory, causing the processor to perform the first aspect and / or various possible implementations of the first aspect as described above.
[0052] Fourthly, embodiments of this application provide a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, are used to implement the first aspect and / or various possible implementations of the first aspect.
[0053] Fifthly, embodiments of this application provide a computer program product, including a computer program that, when executed by a processor, implements the first aspect and / or various possible implementations of the first aspect.
[0054] The vehicle control method, device, storage medium, and program product provided in this application have the following features: The method acquires multiple facial images of the driver using a binocular camera and identifies the two-dimensional position of the midpoint between the driver's eyes based on these images. It then converts the two-dimensional position of the midpoint into three-dimensional spatial coordinates in the vehicle coordinate system. Using these three-dimensional spatial coordinates as the origin and combining the outer contour data of the vehicle's three-dimensional model, it calculates the safe passage area from the current perspective through perspective projection and displays this safe passage area. This method uses the driver's eye position as the reference point for three-dimensional spatial calculation and achieves accurate reconstruction of the eye position using triangulation. This overcomes the limitations of traditional systems that use the vehicle coordinate system or camera perspective as a reference, avoiding the perspective deviation caused by traditional systems using the vehicle coordinate system or camera perspective as a reference, and reducing the cognitive load on the operator and the possibility of misoperation. Furthermore, by updating the driver's eye position and vehicle posture in real time, it ensures that the safe passage area is always consistent with the driver's subjective perspective, avoiding misjudgments caused by perspective deviation. Attached Figure Description
[0055] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0056] Figure 1 A schematic diagram of a scenario for the vehicle control method provided in this application;
[0057] Figure 2 Flowchart of the vehicle control method provided in this application Figure 1 ;
[0058] Figure 3 Flowchart of the vehicle control method provided in this application Figure 2 ;
[0059] Figure 4 A schematic diagram of the vehicle control device provided in this application;
[0060] Figure 5 This is a schematic diagram of the structure of the electronic device provided in this application.
[0061] The accompanying drawings illustrate specific embodiments of this application, which will be described in more detail below. These drawings and descriptions are not intended to limit the scope of the concept in any way, but rather to illustrate the concept of this application to those skilled in the art through reference to particular embodiments. Detailed Implementation
[0062] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims.
[0063] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, storage, use, processing, transmission, provision, disclosure, and application of the relevant data all comply with the relevant laws, regulations, and standards of the relevant countries and regions, have taken necessary confidentiality measures, do not violate public order and good morals, and provide corresponding operation access points for users to choose to authorize or refuse.
[0064] Furthermore, the technical solution involved in this application, which involves big data analysis of user information (including but not limited to personal biometrics, identity data, consumption data, asset data, electronic terminal operation data, etc.) and the use of artificial intelligence technology for automated decision-making, and makes decisions that have a significant impact on personal rights based on the results of automated decision-making, provides users with corresponding operation entry points for users to choose to agree to or reject the results of automated decision-making; if the user chooses to reject, the process will proceed to the expert decision-making process.
[0065] In this application, the terms "exemplary" or "for example" are used to indicate that something is an example, illustration, or description. Any embodiment or design described as "exemplary" or "for example" in this application should not be construed as being more preferred or advantageous than other embodiments or designs. Specifically, the use of terms such as "exemplary" or "for example" is intended to present the relevant concepts in a specific manner.
[0066] In the embodiments of this application, the use of terms such as "first" and "second" is to distinguish between identical or similar items that have essentially the same function and effect. For example, "first electronic device" and "second electronic device" are merely used to distinguish different electronic devices and do not limit their order of execution. Those skilled in the art will understand that the terms "first" and "second" do not limit the quantity or execution order, and that "first" and "second" do not necessarily imply that they are different.
[0067] In this application embodiment, "at least one" refers to one or more, and "more than one" refers to two or more. "And / or" describes the relationship between associated objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following associated objects have an "or" relationship.
[0068] The following is an explanation of some terms used in the embodiments of this application:
[0069] AR-HUD: Augmented Reality Head-Up Display, or AR-HUD for short, is an in-vehicle interactive system that precisely overlays augmented reality virtual information onto the real road scene and projects it onto the vehicle's windshield, allowing the driver to obtain driving data and intelligent prompts without looking down.
[0070] DMS: Driver Monitoring System, is one of the core subsystems of driver assistance systems. It mainly uses onboard sensors to monitor and analyze the driver's physiological state and behavior in real time to determine whether the driver is in a safe driving state.
[0071] Epipolar constraint matching is a core feature matching constraint rule in binocular stereo vision. It is used to limit the search range of corresponding feature points in the two images captured by the left and right cameras, thereby reducing matching complexity and mismatch rate.
[0072] SIFT Feature Matching: Scale-Invariant Feature Transform, or SIFT for short, is a local feature extraction and matching algorithm that is scale-, rotation-, and illumination-invariant. It can accurately match feature points of the same target in images with different viewpoints, lighting, and scaling.
[0073] Triangulation is the core algorithm in binocular stereo vision for recovering three-dimensional spatial coordinates from the coordinates of feature points in a two-dimensional image. It utilizes the known positional relationship between two cameras and combines the projected coordinates of the same spatial point in the two images to solve the three-dimensional coordinates of the point through geometric calculations.
[0074] GPU: Graphics Processing Unit, has a massive number of parallel computing cores and is good at handling large-scale data parallel tasks (such as image rendering and feature extraction).
[0075] FPGA: Field-Programmable Gate Array, based on programmable logic gate arrays, allows for customizable hardware circuits and excels at low-latency, highly deterministic real-time computing tasks.
[0076] With the rapid development of intelligent driving technology, the complexity of vehicle driving scenarios is constantly increasing, and drivers' need for real-time perception and accurate judgment of the vehicle's surrounding environment is becoming more and more urgent. In particular, when vehicles pass through complex road scenarios such as obstacles, width-restricted facilities, and construction sections, drivers need to quickly and accurately judge the feasibility of the vehicle's passage. This need is directly related to the driver's driving safety and the vehicle's traffic efficiency, and has become one of the core focuses of the field of intelligent driving assistance technology.
[0077] In existing driver assistance systems, most of the environmental perception assistance methods for the above-mentioned vehicle driving scenarios use technologies such as panoramic imaging systems, ultrasonic radar detection, or augmented reality head-up displays (AR-HUD) to output assistance information to the driver. Among them, panoramic imaging systems generate a bird's-eye view of the vehicle's surroundings by stitching together multiple cameras, ultrasonic radar provides feedback on obstacle distance information through distance detection, and AR-HUD overlays assistance signs onto the image in the driver's forward field of vision.
[0078] However, the auxiliary display information provided by existing methods is disconnected from the driver's actual perspective, failing to fully match the driver's natural visual perception habits and the actual field of vision corresponding to the eye position. This leads to delays in the driver's judgment of road conditions and spatial perception deviations, increasing the probability of driver errors. For example, the bird's-eye view of panoramic images differs significantly from the driver's eye-level view, requiring the driver to perform additional perspective conversion and spatial scale calculations. If the auxiliary signs overlaid on the AR-HUD are not accurately calibrated based on the driver's eye position, they are prone to offset or misalignment with the actual scene.
[0079] To address the aforementioned problems, this application provides a vehicle control method.
[0080] First, the implementation scenario of this application will be explained.
[0081] Figure 1 The scenario diagram of the vehicle control method provided in this application is shown below. Figure 1As shown, driver 101 can drive vehicle 102, which is equipped with a control unit (not shown), an image acquisition device 103, and an image acquisition device 104. The image acquisition devices 103 and 104 simultaneously acquire facial images of driver 101, and the control unit determines the two-dimensional coordinates of the midpoint between the driver 101's eyes in the image coordinate system. The two-dimensional coordinates are then converted into three-dimensional spatial coordinates. Using the three-dimensional spatial coordinates as the origin of the viewpoint, a perspective projection is performed on the three-dimensional model of vehicle 102 to calculate the spatial area that vehicle 102 can safely pass through in the current posture. The safe passage area is then converted into the pixel coordinate system of the elongated screen and projected in real time onto the elongated screen on the lower inside of the windshield.
[0082] For example, the vehicle 102 is a left-hand drive vehicle. The image acquisition device 103 can be DMS-1, and DMS-1 is located inside the left A-pillar of the vehicle 102. The image acquisition device 104 can be DMS-2, and DMS-2 is integrated into the center of the interior rearview mirror. DMS-1 and DMS-2 together form a binocular field of view to obtain facial images of the driver 101 from different angles.
[0083] In this embodiment of the application, the image acquisition device installed in the vehicle 102 is used to capture the facial image of the driver 101. The image acquisition device can be, for example, a DMS camera module. The control unit can identify and calculate the two-dimensional coordinates of the midpoint between the two eyes based on the captured facial image of the driver 101, and then use the eye point space mapping module to convert the eye point position (two-dimensional coordinates) to the three-dimensional spatial position in the coordinate system of the vehicle 102.
[0084] Understandably, the control unit can perform comprehensive analysis and processing on the acquired facial images to achieve real-time calculation and display updates of the safe passage area; the control unit also includes a vehicle 3D model library and a long projection screen; the vehicle 3D model library is used to store three-dimensional models of different vehicle outlines, track width, wheelbase, rearview mirrors and other key dimensions; the long projection screen is installed on the lower inside of the windshield to display the safe passage area.
[0085] In view of this, this application provides a vehicle control method that utilizes a binocular DMS camera to acquire real-time images of the driver's face and employs triangulation to calculate the three-dimensional coordinates of the eye point. Then, using these eye point coordinates as a reference point, and combining the vehicle's three-dimensional model and real-time environmental data, a perspective projection algorithm is used to dynamically generate a safe passage area centered on the driver's viewpoint, which is then displayed naturally via a long screen. This method combines DMS eye point detection with a vehicle 3D model to achieve the calculation and display of a safe passage area based on the driver's viewpoint. Through triangulation and perspective projection algorithms, it ensures that the displayed content is highly consistent with the position of obstacles in the driver's actual field of vision, eliminating subjective judgment bias caused by viewpoint discrepancies in traditional systems.
[0086] The technical solutions of this application will be described in detail below with reference to specific embodiments. The specific embodiments described below can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of this application will be described below with reference to the accompanying drawings.
[0087] Figure 2 Flowchart of the vehicle control method provided in the embodiments of this application Figure 1 The execution entity of this method can be a control unit installed on a smart car, such as an onboard central computing platform. Figure 2 As shown, the method includes:
[0088] S201: Capture multiple facial images of the driver using a binocular camera.
[0089] Among them, the facial image refers to the digital image data containing the key facial feature areas of the driver 101, which is obtained by the image acquisition device installed on the vehicle 102.
[0090] In this embodiment of the application, the binocular camera is a stereo vision acquisition system constructed based on the parallax principle, consisting of image acquisition device 103 and image acquisition device 104. Specifically, image acquisition device 103 and image acquisition device 104 synchronously acquire facial images of driver 101, and there is a difference in the image acquisition angle between image acquisition device 103 and image acquisition device 104, and depth information is calculated through parallax.
[0091] For example, the binocular cameras (including DMS-1 and DMS-2) in the DMS system are deployed at the A-pillar and the interior rearview mirror of the vehicle 102, respectively, to capture the driver's facial image in real time.
[0092] S202: Recognize the two-dimensional position of the midpoint between the two eyes based on multiple driver facial images.
[0093] The midpoint between the two eyes refers to the geometric midpoint of the line connecting the key feature points of the driver's eyes 101, which is the core reference point used to determine the head posture and gaze direction of the driver 101.
[0094] Understandably, two-dimensional position refers to the coordinate information required to determine the position of a point or object in a two-dimensional plane space, which is usually described by two values (i.e., two dimensions); the two-dimensional position of the midpoint between the two eyes refers to the two-dimensional coordinate information of the geometric midpoint of the line connecting the key feature points of the two eyes in the facial image.
[0095] Therefore, the midpoint between the two eyes is used to detect eye feature points on the facial image of the driver 101, locate the pupil center or geometric center of the eyes, determine the corresponding two-dimensional coordinates, and calculate the geometric center of the line connecting the key feature points of the two eyes, that is, obtain the two-dimensional coordinates of the midpoint between the two eyes.
[0096] For example, in the facial image acquired by the image acquisition device 103 for DMS-1, the two-dimensional coordinates of the center of the left pupil are: The two-dimensional coordinates of the center of the right pupil are Calculate the average of the horizontal and vertical coordinates of the centers of the left and right pupils respectively to obtain the two-dimensional position of the midpoint between the two eyes. The details are as follows:
[0097]
[0098] S203: Convert the two-dimensional position of the midpoint between the two eyes into three-dimensional spatial coordinates in the vehicle coordinate system.
[0099] Understandably, three-dimensional spatial coordinates refer to the spatial position of the midpoint between the two eyes in the vehicle coordinate system, specifically including three dimensions: X, Y, and Z. For example, the position of the midpoint between the driver 101's two eyes in the vehicle coordinate system is (X=0.5m, Y=0.2m, Z=1.2m).
[0100] In this embodiment, computer vision algorithms (such as epipolar constraint matching and SIFT feature matching) are used to calculate the horizontal position difference (i.e., parallax) of the midpoint between the two eyes in multiple facial images of the driver 101. Then, using triangulation, the two-dimensional coordinates of the midpoint between the two eyes are mapped to three-dimensional spatial coordinates in the camera coordinate system using the known camera baseline distance, focal length, and parallax. Finally, the three-dimensional coordinates of the camera coordinate system are mapped to three-dimensional spatial coordinates in the vehicle coordinate system using the homogeneous transformation matrix between the camera coordinate system and the vehicle coordinate system.
[0101] For example, the epipolar constraint algorithm is used to locate the midpoint between the driver 101's eyes in the facial images acquired by DMS-1 and DMS-2. The pixel coordinates are: "DMS-1: "DMS-2: " The vertical coordinates of the two are consistent. and This represents DMS-1 and DMS-2 installed horizontally with no vertical parallax. Based on this, the calculation is performed. and The difference between the two eyes can be used to obtain the disparity at the midpoint between the two eyes. And the parallax is for Using triangulation, the two-dimensional position of the midpoint between the two eyes was determined. Mapping to the left camera coordinate system's 3D coordinates, the resulting 3D coordinates in the left camera coordinate system are: Then, through homogeneous transformation matrices (including rotation matrices) Translation vector ), the corresponding three-dimensional coordinates in the left camera coordinate system Transforming to the vehicle coordinate system, the resulting three-dimensional spatial coordinates in the vehicle coordinate system are: This completes the mapping from the two-dimensional position of the facial image to the three-dimensional spatial coordinates in the vehicle coordinate system. The vehicle coordinate system is defined as O-XYZ, with its origin O being the projection of the intersection of the rear wheel axle centerline of vehicle 102 and the longitudinal plane of symmetry onto the horizontal ground. The X-axis points along the longitudinal plane of symmetry towards the forward direction of vehicle 102, the Y-axis is perpendicular to the longitudinal plane of symmetry and points to the left side of vehicle 102, and the Z-axis is perpendicular to the ground and points upwards. Then, the point... In the vehicle coordinate system, this means: X-axis is 6.5m ahead, Y-axis is in the center, and Z-axis is 0.8m above ground.
[0102] S204: Using three-dimensional spatial coordinates as the origin and combining the outer contour data of the vehicle's three-dimensional model, the safe passage area from the current perspective is calculated through perspective projection.
[0103] S205: Displays the safe passage area.
[0104] The safe passage area refers to the spatial range in which vehicle 102 can safely pass.
[0105] Understandably, the outer contour data of vehicle 102 refers to the set of core dimensional parameters and feature point coordinates that characterize the spatial geometric boundary of the vehicle body and key components, and is used to define the physical boundary of vehicle 102 in three-dimensional space; such outer contour data includes, for example, track width, wheelbase, and rearview mirror size.
[0106] Perspective projection calculation is a geometric calculation method that uses the observation point as the origin to map three-dimensional space onto a two-dimensional display plane. For example, through perspective projection with the midpoint between the driver's two eyes as the origin, the outline of vehicle 102 and the position of obstacles are mapped onto the two-dimensional plane of the windshield of vehicle 102.
[0107] In this embodiment, the three-dimensional spatial coordinates in the vehicle coordinate system corresponding to the midpoint between the driver's two eyes are used as the origin of perspective projection because these three-dimensional spatial coordinates coincide with the actual observation angle of the driver 101, which enables the calculated safe passage area to be completely matched with the driver 101's real field of vision, thereby eliminating the perspective deviation of the traditional fixed origin scheme. In addition, this origin provides a precise virtual image projection reference for vehicle-mounted display devices such as AR-HUD, ensuring that the boundary of the safe area is accurately matched with the real road scene.
[0108] The vehicle 102 is also equipped with a linear display area for projecting safe passage area information; the linear display area may be located, for example, on the lower inside of the windshield, specifically a 10cm × 2cm rectangular area along the lower edge of the windshield.
[0109] Using the three-dimensional spatial coordinates of the vehicle coordinate system at the midpoint between the driver's two eyes as the origin of perspective projection (i.e., simulating the real viewpoint of the driver 101), and combining the spatial coordinate data of the three-dimensional outer contour of the vehicle 102, the boundary of the area that the vehicle 102 can safely pass through in actual complex driving scenarios (such as width-restricted road sections and obstacle areas) is calculated through perspective projection transformation, and this safe area is superimposed on the linear display area of the vehicle.
[0110] In this embodiment, the control unit of the intelligent vehicle pre-stores the outer contour data of the vehicle 102. For example, the key outer contour feature points of the vehicle 102 are pre-calibrated using vehicle 3D modeling software. All feature points are defined based on the vehicle coordinate system, specifically including basic size parameters and key component parameters: the basic size parameters are a wheelbase of 2800mm, a front track of 1650mm, a rear track of 1670mm, and an overall vehicle width of 1850mm; the key component parameters are that the rearview mirror protrudes 150mm from the outer side of the vehicle body, and the Y-axis coordinate of the outer end of the rearview mirror in the vehicle coordinate system is ±1000mm; the minimum spatial bounding box of the vehicle 102 is fitted based on the above outer contour data, which is used for subsequent perspective projection calculation of the safe passage area.
[0111] For example, the set of feature points of the three-dimensional outer contour of vehicle 102 Fit to minimum spatial bounding box The six sides of the enclosure box correspond to the top, bottom, left, right, front, and rear limits of the vehicle body, respectively:
[0112]
[0113] in, and Corresponding to the leftmost / rightmost limits of the vehicle body; and Corresponding to the foremost / rearmost limit of the vehicle body.
[0114] When vehicle 102 is in a width-restricted scenario, the coordinates of obstacle boundary points in the road scene ahead are collected. The bounding box feature points of vehicle 102 and the obstacle boundary points are then substituted into the perspective projection formula to convert them into two-dimensional coordinates of the imaging plane. In other words, the outer contour points of vehicle 102 and the obstacle points in the scene are projected onto a two-dimensional imaging plane (such as the virtual image plane of an AR-HUD or the camera imaging plane) from the driver 101's perspective. The core formula for perspective projection is (based on the vehicle coordinate system, with the projection origin...). Based on this, specifically as shown in Formula 1:
[0115]
[0116] in, The two-dimensional coordinates of the imaging plane, For the focal length of the viewpoint, The coordinates are the center coordinates of the imaging plane.
[0117] By comparing the imaging plane coordinates of the vehicle 102 boundary (i.e., the bounding box feature points of the vehicle 102) with the obstacle boundary, the safe passage area under the current view is calculated and superimposed onto the 10cm×2cm rectangular area at the bottom edge of the windshield (i.e., the vehicle display device).
[0118] The vehicle control method provided in this embodiment acquires multiple facial images of the driver using a binocular camera and identifies the two-dimensional position of the midpoint between the driver's eyes based on these images. The two-dimensional position of the midpoint is then converted into three-dimensional spatial coordinates in the vehicle coordinate system. Using these three-dimensional spatial coordinates as the origin and combining the outer contour data of the vehicle's three-dimensional model, the safe passage area from the current perspective is calculated through perspective projection and displayed. This method uses the driver's eye position as the reference point for three-dimensional spatial calculation and achieves accurate reconstruction of the eye position using triangulation. This overcomes the limitations of traditional methods that rely on the vehicle coordinate system or camera perspective, avoiding the perspective deviation caused by these methods. This reduces the cognitive load on the operator and the possibility of misoperation. Furthermore, by updating the driver's eye position and vehicle posture in real time, it ensures that the safe passage area is always consistent with the driver's subjective perspective, avoiding misjudgments caused by perspective deviation.
[0119] Figure 3 Flowchart of the vehicle control method provided in the embodiments of this application Figure 2 .like Figure 3 As shown, in this embodiment... Figure 2 Based on the embodiments, the vehicle control method is described in detail, which includes:
[0120] S301: Captures multiple facial images of the driver using a binocular camera.
[0121] S302: Recognize the two-dimensional position of the midpoint between the two eyes based on multiple driver facial images.
[0122] Multiple facial images of the driver are captured by a binocular camera. For each facial image of the driver 101, the geometric center of each eye is identified, and the two-dimensional position of the midpoint between the two eyes is calculated based on this.
[0123] In one possible implementation, the two-dimensional position of the midpoint between the two eyes is achieved based on deep learning feature detection and the principle of binocular epipolar constraints. Specifically, firstly, a convolutional neural network model is used to locate face and eye feature points in facial images captured by binocular cameras, and the coordinates of the core feature points of the left and right eyes are obtained through image segmentation and geometric fitting. Secondly, based on the epipolar constraint relationship of the binocular camera calibration parameters, the accurate matching of eye feature points in the left and right images is completed. Finally, through outlier removal and weighted fusion of multi-frame coordinates, single-frame positioning errors are eliminated to obtain the final two-dimensional pixel coordinates of the midpoint between the two eyes.
[0124] In one possible implementation, the image acquisition parameters of the binocular camera are dynamically adjusted according to environmental conditions; based on the adjusted image acquisition parameters of the binocular camera, multiple facial images of the driver are acquired.
[0125] Among them, environmental conditions include light intensity and rain / fog density, and image acquisition parameters include camera exposure time and white balance parameters.
[0126] It is understandable that dynamically adjusting the image acquisition parameters of the binocular camera according to environmental conditions is to ensure that the facial images acquired under different environments have high definition and high recognition, thereby improving the accuracy of the midpoint recognition between the two eyes.
[0127] For example, in strong light environments (such as direct midday sunlight or oncoming vehicle headlights), if the camera exposure time is too long, the image sensor will receive an excessive number of photons, resulting in overexposure of the facial image. In this case, facial details (such as eye contours and pupil changes) will be obscured by highlights and cannot be effectively identified. If the white balance parameters do not match the color temperature of the strong light, the image will show color cast (such as a bluish-white cast under strong light), further interfering with facial feature extraction. Therefore, the exposure time can be shortened to reduce the number of photons received by the image sensor and avoid overexposure. Adjusting the white balance parameters (such as increasing the color temperature compensation value) can restore the true colors of the facial area and ensure the clarity of skin tone and facial features.
[0128] In low-light environments (such as at night or inside tunnels), if the exposure time is too short, the image sensor cannot obtain enough light, resulting in increased noise, a dark image, and blurred facial contours. If the white balance parameters are not adapted to the low-light environment (such as the warm yellow light of streetlights), the image will appear yellowish or reddish, making it difficult to distinguish facial features (such as the degree of eyelid closure). Therefore, the exposure time can be appropriately extended to increase the amount of light entering the image and improve its brightness. Adjusting the white balance parameters to calibrate the color temperature can eliminate color interference from ambient light and highlight the contrast of facial features.
[0129] Because rain and fog reduce air transparency, light is scattered and refracted during propagation, resulting in blurry facial images captured by the camera, reduced contrast, and blurred boundaries between the face and background, directly affecting the accuracy of feature recognition. In this case, the exposure time can be appropriately extended to increase the overall brightness of the image and offset the light attenuation caused by rain and fog; at the same time, avoid excessive extension that causes image ghosting (due to the movement of vehicle 102 or slight head movement of driver 101).
[0130] In addition, since the scattered light on rainy or foggy days is mostly cool-toned (such as grayish-white), white balance parameter calibration can correct the cool color cast of the image, enhance the color differentiation of facial skin, hair, and clothing, and improve the accuracy of feature extraction.
[0131] S303: Parallax information is obtained based on the two-dimensional position of the midpoint between the two eyes in multiple driver facial images.
[0132] S304: Based on parallax information, the baseline distance and focal length between the two DMS cameras, the three-dimensional spatial coordinates of the midpoint between the two eyes are calculated using triangulation.
[0133] Parallax information refers to the difference in the horizontal position of the same target point in a binocular image.
[0134] Understandably, the binocular camera system includes two DMS cameras, which are positioned on either side of the driver's seat.
[0135] Triangulation is based on geometric principles, using known baseline distance and parallax to estimate the depth of a target point (i.e., the midpoint between the two eyes); for example, baseline distance... It is 5cm, parallax It is 20, focal length The depth of the target point is 30mm. The vertical distance from the target point to the imaging plane of the binocular camera, i.e., the depth of the target point, is calculated using the following formula 2. :
[0136]
[0137] In this embodiment, when calculating the three-dimensional spatial coordinates of the midpoint between the two eyes in the camera coordinate system using triangulation, the calculation process relies on the calibration parameters of the binocular camera and the imaging pixel coordinates of the target point. The calibration parameters of the binocular camera are pre-calibrated and include binocular intrinsic and extrinsic parameters. The imaging pixel coordinates of the target point refer to the two-dimensional pixel position of the midpoint between the driver's 101's two eyes on the corresponding imaging plane of the binocular camera, serving as direct input data for triangulation. This application does not impose any special restrictions on the calibration method of the binocular camera's calibration parameters.
[0138] Specifically, binocular intrinsic parameters are the imaging geometric parameters of a single camera, and are calibrated separately for the left and right individual cameras. They are used to correct optical distortions during camera manufacturing and installation, thereby establishing the projection relationship between the three-dimensional spatial points and two-dimensional pixel points of a single camera. The core parameters of these binocular intrinsic parameters include: focal length. and principal point coordinates .
[0139] Among them, focal length (unit: The physical essence of the principal point coordinates is the distance from the optical center of the camera to the imaging plane, converted to pixel size, which determines the calculation accuracy of the triangulation method. (unit: The principal point coordinates match the camera's resolution and correspond to the center pixel position of the imaging plane. The principal point is the optical center of the camera's imaging, and all projected rays from spatial points will pass through this point. When calculating the pixel offset of the target point, the principal point coordinates should be used as a reference to eliminate projection errors caused by the offset of the imaging plane.
[0140] Binocular extrinsic parameters are the relative pose parameters of the binocular cameras. These parameters describe the spatial position and orientation relationship between the binocular cameras and are a prerequisite for binocular image matching and disparity calculation. The core parameters of these binocular extrinsic parameters include: baseline distance. Rotation matrix Translation vector .
[0141] Among them, baseline distance (Unit: mm) refers to the horizontal straight-line distance between the optical centers of the left and right cameras, which is a known fixed geometric quantity measured by triangulation; baseline distance. The longer the baseline, the greater the parallax of targets at the same depth, and the higher the accuracy of depth calculation; conversely, if the baseline is too short, the parallax will be overwhelmed by noise, leading to calculation failure.
[0142] Rotation matrix Translation vector Together, they defined the pose transformation relationship between the right camera and the left camera; this rotation matrix The identity matrix indicates that the left and right cameras are horizontally parallel with no relative rotation, and the imaging planes are completely parallel, which simplifies parallax calculation (only the horizontal pixel difference needs to be calculated); for example, the translation vector , which indicates the positional offset of the right camera in the coordinate system of the left camera: there is an offset of -200mm only in the X-axis direction (consistent with the baseline distance), and no offset in the Y and Z-axis directions; this parameter is used to transform the pixel coordinates of the right camera to the coordinate system of the left camera to achieve spatial alignment of the binocular images.
[0143] For the two-dimensional position of the midpoint between the eyes in multiple facial images of driver 101 acquired by two DMS cameras, the difference in the horizontal position of the midpoint between the eyes in different facial images is calculated to obtain the disparity information of the midpoint between the eyes. Based on the disparity information acquired in this instance, the baseline distance and focal length between the two DMS cameras, the three-dimensional spatial coordinates of the midpoint between the eyes in the camera coordinate system are calculated using triangulation. Then, a rotation matrix is used... Translation vector The three-dimensional spatial coordinates in the camera coordinate system are converted to three-dimensional spatial coordinates in the vehicle coordinate system.
[0144] In one possible implementation, the average two-dimensional position in multiple facial images of the driver 101 can be calculated for a single DMS camera to obtain a unique two-dimensional position corresponding to the single DMS camera. Based on the two-dimensional positions of the two DMS cameras, the disparity information of the midpoint between the two eyes can be determined. Alternatively, for two facial images acquired at the same acquisition time from two DMS cameras, the disparity information of the midpoint between the two eyes at the same acquisition time can be calculated, and then the average value of the disparity information from multiple acquisition times can be calculated to obtain the disparity information of the current parameter three-dimensional spatial coordinate transformation. This application does not impose any special limitations on the acquisition of disparity information.
[0145] For example, if the binocular intrinsic parameters include: the focal length of the left camera 1200 Principal point coordinates (Corresponding image center resolution is 1280×720); Binocular extrinsic parameters include: baseline distance The rotation matrix of DMS-2 relative to DMS-2 is 200mm (horizontal distance between the optical centers of the left and right cameras). The identity matrix and the translation vector (Horizontal parallel installation, no rotation) In this case, the homogeneous transformation matrix for converting the camera coordinate system to the vehicle coordinate system includes the rotation matrix. Translation vector , where the rotation matrix for:
[0146]
[0147] Translation vector for:
[0148]
[0149] Understandable, rotation matrix Translation vector , and , rotation matrix Translation vector It describes the pose relationship between different coordinate systems; among them, the rotation matrix Translation vector It describes the pose relationship between the two cameras and is used to transform the coordinate system of the right camera corresponding to DMS-2 to the coordinate system of the left camera corresponding to DMS-1; rotation matrix Translation vector These are the external pose parameters of the camera and the vehicle, used to convert the three-dimensional coordinates under the camera's view (such as the midpoint between the driver's eyes) into coordinates under the vehicle coordinate system. In other words, the coordinate system of the left camera corresponding to DMS-1 is converted into the vehicle coordinate system, which is the core of the safe passage area calculation.
[0150] Using the epipolar constraint algorithm, the pixel coordinates of the midpoint between the driver's two eyes are located in the left and right images, and the disparity information of the midpoint between the two eyes is calculated. The obtained disparity information is 40px. The depth of the midpoint between the two eyes is calculated using Formula 2. Then, pixel normalization is performed to transform the two-dimensional position to the camera normalization plane, thus obtaining the corresponding camera coordinate system. , This allows us to obtain the three-dimensional coordinates corresponding to the camera coordinate system. Then, through homogeneous transformation matrices (including rotation matrices) Translation vector ), and the three-dimensional coordinates in the coordinate system of the left camera corresponding to DMS-2 Transforming to the vehicle coordinate system, the resulting three-dimensional spatial coordinates in the vehicle coordinate system are: .
[0151] S305: Real-time detection of driver head posture changes, obtaining head posture parameters, and dynamically adjusting three-dimensional spatial coordinates based on head posture parameters.
[0152] S306: Monitors vehicle attitude information in real time and dynamically adjusts attitude parameters in the outer contour data based on the vehicle attitude information.
[0153] Among them, the head posture parameters are used to describe the driver's head tilt angle, pitch angle and other posture data; the head posture data includes, for example, a head tilt angle of 15° and a pitch angle of 5°; the vehicle 102 posture information is used to describe the dynamic parameters of the vehicle 102 such as pitch angle and roll angle; the vehicle 102 posture information includes, for example, a pitch angle of -5° when the vehicle 102 is driving on a slope.
[0154] In this embodiment, since the driver 101 may perform head movements such as turning, tilting, and rolling during driving (e.g., turning the head to observe the rearview mirror or looking down to adjust the air conditioning), the perspective of the facial image captured by the binocular camera will change. Therefore, by detecting the changes in the driver's head posture in real time, updating the three-dimensional spatial coordinates of the midpoint between the driver's two eyes in real time, and dynamically adjusting the perspective projection matrix, the real-time performance and accuracy of the safe area calculation are ensured.
[0155] As vehicle 102 undergoes attitude changes due to road conditions (such as climbing, descending, turning, and bumping) during its journey, parameters such as pitch and roll angles of vehicle 102 will change accordingly. Since the virtual information overlay of AR-HUD is based on the relative attitude calculation between vehicle 102 and the road, the attitude changes of vehicle 102 itself will directly affect the accuracy of perspective projection. Therefore, by loading the vehicle's 3D model and combining it with real-time sensor data, the attitude parameters of the model are dynamically corrected, solving the problem that traditional static models cannot adapt to the actual motion state of vehicle 102.
[0156] For example, when the driver 101 turns his head to the left, the control unit automatically corrects the calculation origin of the three-dimensional spatial coordinates to ensure that the generated safe passage area is always consistent with the driver's current viewpoint. This process achieves dynamic adaptation through head pose estimation algorithms (such as 3D face key point detection) to provide a stable benchmark for subsequent projection calculations.
[0157] When driving on a slope, the control unit automatically adjusts the model's pitch angle parameters based on altitude sensor data to ensure that the calculation of the safe passage area is based on the actual attitude of vehicle 102; this process provides a more accurate benchmark for subsequent projection calculations through real-time fusion of sensor data and model parameters.
[0158] In one possible implementation, the head attitude parameters include yaw angle, pitch angle, and roll angle, wherein the yaw angle is used to describe the angle of left and right rotation of the head (rotation about Z), the pitch angle is used to describe the angle of up and down rotation of the head (rotation about X), and the roll angle is used to describe the angle of left and right tilt of the head (rotation about Y).
[0159] The three-dimensional motion of the head can be broken down into two actions: rotation and translation. The corresponding transformation matrices include the rotation matrix. Translation vector Together, they form the head pose transformation matrix; through the head pose transformation matrix, the initial three-dimensional coordinates of the midpoint between the two eyes in the vehicle coordinate system are mapped to new three-dimensional spatial coordinates that match the real-time head pose.
[0160] For example, the rotation matrices corresponding to the three angles of yaw, pitch, and roll are respectively , , The combined total rotation matrix is ,and Translation vector The Z-axis coordinate is used to characterize the displacement change of the head relative to the initial position and can be calculated by binocular parallax (e.g., when the head is extended forward, the Z-axis coordinate decreases, and when it is tilted backward, the Z-axis coordinate increases).
[0161] In one possible implementation, the attitude information of vehicle 102 includes pitch angle and roll angle; by establishing a linkage mapping relationship between the attitude of vehicle 102 and the outer contour data, and by compensating for the spatial attitude deviation caused by the pitch and roll of vehicle 102, the outer contour of virtual information (such as safe passage area, navigation guidance) is always matched with the real road / vehicle attitude.
[0162] For example, the pitch angle compensation formula is shown in Formula 3 below:
[0163]
[0164] in, The adjusted pitch angle of the outer contour. The initial pitch angle of the outer contour. The pitch angle collected in real time for vehicle 102 (positive for climbing and negative for descending).
[0165] The roll angle compensation formula is shown in Formula 4 below:
[0166]
[0167] in, The adjusted outer contour camber angle. The initial side tilt angle of the outer contour. The roll angle (positive for left roll, negative for right roll) is collected in real time for vehicle 102.
[0168] S307: Using the adjusted three-dimensional spatial coordinates as the origin, combined with the attitude parameters in the corrected outer contour data and the vehicle's surrounding environment data, the safe passage area is calculated through perspective projection.
[0169] S308: Displays the safe passage area.
[0170] Using the three-dimensional spatial coordinates of the midpoint between the driver's two eyes in the vehicle coordinate system as the origin of perspective projection, combined with the spatial coordinate data of the corrected three-dimensional outer contour of the vehicle 102, and the surrounding environmental data of the vehicle (such as obstacle data and weather data), the boundary of the area that the vehicle 102 can safely pass through in the actual driving scenario (such as a width-restricted road section or an obstacle area) is calculated through perspective projection transformation, and this safe area is superimposed on the linear display area of the vehicle 102.
[0171] In one possible implementation, the vehicle also includes a long display screen positioned on the lower inner side of the windshield; the long display screen shows a safe passage area.
[0172] During the process of displaying the safe passage area on the long display screen, the driver's line of sight is detected based on the images captured by the binocular camera to obtain the line of sight parameters; based on the line of sight parameters, the display range and visual priority of the long display screen are dynamically adjusted.
[0173] Among them, the field of vision parameter is used to describe parameters such as the driver's gaze point and head turning angle; visual priority is achieved by differentiating key information through color contrast and / or flashing cues.
[0174] As is understandable, a long strip display screen refers to a long strip-shaped in-vehicle display terminal installed inside the vehicle 102 (such as on the dashboard or in front of the passenger seat). It has the ability to display in sections and dynamically adjust the resolution, and is used to display driving-related information such as safe passage areas, lane guidance, and warning prompts. The display range refers to the area of the long strip display screen that is currently used to display effective information, which can be adjusted through operations such as offsetting, zooming, and cropping.
[0175] The safety zone is displayed on the long strip screen in the form of boundary markers and highlighted areas, and is suitable for scenarios such as passing on narrow roads, reversing into parking spaces, and merging into ramps.
[0176] For example, when driver 101 observes an obstacle on the left, the control unit automatically enhances the display ratio of the left area and strengthens visual priority through high-contrast colors, thereby automatically enhancing the display of the area that driver 101 is looking at and reducing the risk of misjudgment due to obstructed vision. Specifically, before the display ratio is adjusted, the elongated display screen adopts a symmetrical 1:1 display ratio, with the obstacle area on the left and the oncoming traffic area on the right each occupying 50%. After the adjustment, the display ratio of the obstacle area on the left is increased to 70%, the oncoming traffic area on the right is compressed to 30%, and the guardrail boundary and ditch depth markings on the left are magnified by 2 times to clearly present the distance gap between the obstacle and the vehicle body.
[0177] In one possible implementation, the vehicle 102 includes a first computing unit and a second computing unit deployed locally on the vehicle 102, wherein the first computing unit is used to calculate the parallax information of multiple driver facial images; and the second computing unit is used to perform three-dimensional model rendering related steps and perspective projection calculation.
[0178] Understandably, due to the latency in the existing processing chain, the displayed content cannot adapt to changes in the driver's actions in real time. Therefore, edge computing devices (such as on-board GPUs or FPGAs) are deployed in vehicle 102 to parallelize high-load tasks (such as eye detection, 3D mapping, and passability calculations) and reduce end-to-end latency through hardware acceleration.
[0179] By optimizing edge computing and parallel processing, the latency problem caused by traditional serial processing links is solved. For example, FPGA is used to perform real-time parallax calculation on binocular images, while GPU is used to execute 3D model rendering and perspective projection in parallel, which significantly shortens the response time from image acquisition to display update. Through hardware acceleration and task parallelization, it is ensured that the displayed content is always synchronized with the driver's real-time actions, avoiding the risk of misjudgment due to latency and significantly improving the real-time interactive experience in complex scenarios.
[0180] The vehicle control method provided in this embodiment acquires multiple driver facial images using a binocular camera and identifies the two-dimensional position of the midpoint between the two eyes based on these images. Based on the two-dimensional position of the midpoint between the eyes in the multiple driver facial images, disparity information is obtained. Then, based on the disparity information, the baseline distance and focal length between the two DMS cameras, triangulation is used to calculate the three-dimensional spatial coordinates of the midpoint between the eyes. Real-time detection of driver head posture changes yields head posture parameters, and the three-dimensional spatial coordinates are dynamically adjusted based on these parameters. Real-time monitoring of vehicle posture information is also performed, and posture parameters in the outer contour data are dynamically adjusted based on this information. Using the adjusted three-dimensional spatial coordinates as the origin, combined with the posture parameters in the corrected outer contour data and the vehicle's surrounding environment data, a safe passage area is calculated through perspective projection and displayed. This method uses the driver's eye position as the reference point for three-dimensional spatial calculation and achieves accurate reconstruction of the eye position using triangulation. This overcomes the limitations of traditional methods that rely on vehicle coordinate systems or camera perspectives, avoiding the perspective deviations caused by these methods and reducing the cognitive load and possibility of operator errors.
[0181] Figure 4 This is a schematic diagram of the structure of the vehicle control device provided in the embodiments of this application, such as... Figure 4 As shown, this application embodiment provides a vehicle control device 400, which includes:
[0182] The acquisition module 401 is used to acquire multiple facial images of the driver through a binocular camera.
[0183] Processing module 402 is used to identify the two-dimensional position of the midpoint between the two eyes based on multiple driver facial images.
[0184] The processing module 402 is also used to convert the two-dimensional position of the midpoint between the two eyes into three-dimensional spatial coordinates in the vehicle coordinate system.
[0185] The processing module 402 is also used to calculate the safe passage area from the current perspective by using the three-dimensional spatial coordinates as the origin and combining the outer contour data of the vehicle's three-dimensional model through perspective projection.
[0186] The processing module 402 is also used to display the safe passage area.
[0187] In one possible implementation, the processing module 402 is further configured to obtain parallax information based on the two-dimensional position of the midpoint between the two eyes in multiple driver facial images.
[0188] The processing module 402 is also used to calculate the three-dimensional spatial coordinates of the midpoint between the two eyes using triangulation based on parallax information, the baseline distance between the two DMS cameras and the focal length.
[0189] In one possible implementation, the acquisition module 401 is also used to detect changes in the driver's head posture in real time and obtain head posture parameters.
[0190] The processing module 402 is also used to dynamically adjust the three-dimensional spatial coordinates based on the head posture parameters.
[0191] The acquisition module 401 is also used to monitor the vehicle's attitude information in real time.
[0192] The processing module 402 is also used to dynamically adjust the attitude parameters in the outer contour data based on the vehicle's attitude information.
[0193] In one possible implementation, the processing module 402 is further configured to calculate the safe passage area by perspective projection, using the adjusted three-dimensional spatial coordinates as the origin, combining the attitude parameters in the corrected outer contour data, and the vehicle's surrounding environment data.
[0194] In one possible implementation, the processing module 402 is also configured to display a safe passage area on a long strip display screen.
[0195] The processing module 402 is also used to detect the driver's line of sight based on the images acquired by the binocular camera and obtain the line of sight parameters.
[0196] The processing module 402 is also used to dynamically adjust the display range and visual priority of the elongated display screen based on the line-of-sight parameter; the visual priority is achieved by differentiating key information through color contrast and / or flashing prompts.
[0197] In one possible implementation, the processing module 402 is further configured to dynamically adjust the image acquisition parameters of the binocular camera according to environmental conditions; the environmental conditions include light intensity and rain / fog density, and the image acquisition parameters include camera exposure time and white balance parameters.
[0198] The acquisition module 401 is also used to acquire multiple driver facial images based on the image acquisition parameters adjusted by the binocular camera.
[0199] The vehicle control device provided in this application embodiment can be used to execute the technical solution of the vehicle control method in any of the above embodiments of this application. Its implementation principle and technical effect are similar, and will not be described again here.
[0200] Figure 5 A schematic diagram of the structure of the electronic device provided in this application. Figure 5 As shown, the electronic device 500 provided in this embodiment includes at least one processor 501 and a memory 502. Optionally, the device 500 further includes a communication component 503. The processor 501, memory 502, and communication component 503 are connected via a bus.
[0201] In a specific implementation, at least one processor 501 executes computer execution instructions stored in memory 502, causing at least one processor 501 to perform the above-described method.
[0202] The specific implementation process of processor 501 can be found in the above method embodiments, and its implementation principle and technical effect are similar. It will not be repeated here.
[0203] In the above embodiments, it should be understood that the processor can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), etc. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the method disclosed in this invention can be directly implemented by a hardware processor, or implemented by a combination of hardware and software modules within the processor.
[0204] The memory may include random access memory (RAM) and may also include non-volatile memory (NVM), such as at least one disk storage device.
[0205] The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. Buses can be categorized as address buses, data buses, control buses, etc. For ease of illustration, the buses shown in the accompanying drawings are not limited to a single bus or a single type of bus.
[0206] This application also provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implement the method of any of the foregoing embodiments.
[0207] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the method of any of the foregoing embodiments.
[0208] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For instance, the division of modules is only a logical functional division, and there may be other division methods in actual implementation. For example, multiple modules may be combined or integrated into another system, or some features may be ignored or not executed.
[0209] The integrated modules described above, implemented as software functional modules, can be stored in a computer-readable storage medium. These software functional modules, stored in a storage medium, include several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) or processor to execute some steps of the methods of the various embodiments of this application.
[0210] The aforementioned storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof. Examples of storage media include Static Random-Access Memory (SRAM) or Electrically Erasable Programmable Read Only Memory (EEPROM).
[0211] Storage media can be, for example, erasable programmable read-only memory (EPROM) or programmable read-only memory (PROM). Storage media can also be read-only memory (ROM), magnetic storage, flash memory, magnetic disks, or optical disks. Storage media can be any available medium accessible to general-purpose or special-purpose computers.
[0212] An exemplary storage medium is coupled to a processor, enabling the processor to read information from and write information to the storage medium. Alternatively, the storage medium can be an integral part of the processor. The processor and storage medium can reside within an application-specific integrated circuit (ASIC). Alternatively, the processor and storage medium can exist as discrete components within an electronic device or host device.
[0213] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.
[0214] The sequence numbers of the embodiments in this application are merely for description and do not represent the superiority or inferiority of the embodiments. Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method.
[0215] Based on this understanding, the technical solution of this application, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods of the various embodiments of this application.
[0216] The above are merely preferred embodiments of this application and do not limit the patent scope of this application. Any equivalent structural or procedural transformations made using the content of this application's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of this application.
[0217] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, as some steps may be performed in other orders or simultaneously according to this application. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are all optional embodiments, and the actions and modules involved are not necessarily essential to this application.
[0218] It should be further noted that although the steps in the flowchart are shown sequentially as indicated by the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise explicitly stated in this document, there is no strict order requirement for the execution of these steps, and they can be executed in other orders.
[0219] Furthermore, at least some steps in the flowchart may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but may be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but may be executed in turn or alternately with other steps or at least some of the sub-steps or stages of other steps.
[0220] In the above embodiments, the descriptions of each embodiment have their own emphasis. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments. The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as the combination of these technical features does not contradict each other, it should be considered within the scope of this specification.
[0221] Other embodiments of this application will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of this application that follow the general principles of this application and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this application are indicated by the following claims.
[0222] It should be understood that this application is not limited to the precise structure described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this application is limited only by the appended claims.
Claims
1. A vehicle control method, characterized in that, The vehicle includes a binocular camera, and the method includes: Multiple facial images of the driver were captured using a binocular camera; Based on the multiple driver facial images, the two-dimensional position of the midpoint between the two eyes is identified; The two-dimensional position of the midpoint between the two eyes is converted into three-dimensional spatial coordinates in the vehicle coordinate system; Using the three-dimensional spatial coordinates as the origin and combining the outer contour data of the vehicle's three-dimensional model, the safe passage area from the current perspective is calculated through perspective projection. The safe passage area is displayed.
2. The method according to claim 1, characterized in that, The binocular camera includes two DMS cameras, which are respectively located on both sides of the driver's seat; The step of converting the two-dimensional position of the midpoint between the two eyes into three-dimensional spatial coordinates in the vehicle coordinate system includes: Based on the midpoint between the two eyes, parallax information is obtained from the two-dimensional position of the multiple driver facial images; Based on the parallax information, the baseline distance and focal length between the two DMS cameras, the three-dimensional spatial coordinates of the midpoint between the two eyes are calculated using triangulation.
3. The method according to claim 2, characterized in that, Before calculating the safe passage area from the current viewpoint using perspective projection with the three-dimensional spatial coordinates as the origin and the outer contour data of the vehicle's three-dimensional model as the origin, the method further includes: Real-time detection of driver head posture changes to obtain head posture parameters; Based on the head posture parameters, the three-dimensional spatial coordinates are dynamically adjusted; Real-time monitoring of the vehicle's attitude information; Based on the vehicle's attitude information, the attitude parameters in the outer contour data are dynamically adjusted.
4. The method according to claim 3, characterized in that, The step of calculating the safe passage area from the current viewpoint using the three-dimensional spatial coordinates as the origin and combining the outer contour data of the vehicle's three-dimensional model through perspective projection includes: Using the adjusted three-dimensional spatial coordinates as the origin, and combining the attitude parameters in the corrected outer contour data with the surrounding environment data of the vehicle, the safe passage area is calculated through perspective projection.
5. The method according to any one of claims 1-4, characterized in that, The vehicle also includes a long strip display screen disposed on the lower inner side of the windshield; the display of the safe passage area includes: The safe passage area is displayed on the elongated display screen. The process of displaying the safe passage area on the elongated display screen also includes: Based on the images captured by the binocular camera, the driver's line of sight is detected, and the line of sight parameters are obtained; Based on the aforementioned line-of-sight parameters, the display range and visual priority of the elongated display screen are dynamically adjusted; the visual priority is achieved by differentiating key information through color contrast and / or flashing prompts.
6. The method according to any one of claims 1-4, characterized in that, The method further includes: The image acquisition parameters of the binocular camera are dynamically adjusted according to environmental conditions; the environmental conditions include light intensity and rain / fog density; the image acquisition parameters include camera exposure time and white balance parameters. The process of acquiring multiple facial images of the driver using a binocular camera includes: Based on the adjusted image acquisition parameters of the binocular camera, multiple facial images of the driver are acquired.
7. The method according to any one of claims 1-4, characterized in that, The vehicle includes a first computing unit and a second computing unit deployed locally on the vehicle, wherein, The first calculation unit is used to calculate the disparity information of the multiple driver facial images; The second computing unit is used to perform steps related to 3D model rendering and perspective projection calculation.
8. An electronic device, characterized in that, include: Memory and processor; The memory stores computer-executed instructions; The processor executes computer execution instructions stored in the memory, causing the processor to perform the method as described in any one of claims 1-7.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions, which, when executed by a processor, are used to implement the method as described in any one of claims 1-7.
10. A computer program product, characterized in that, Includes a computer program that, when executed by a processor, implements the method according to any one of claims 1-7.