Vehicle cabin control method and device, controller and vehicle

By collecting image detection feature points in the vehicle cabin and determining the size information of the target object, the cabin equipment is automatically adjusted, solving the problem of low cabin control efficiency and achieving rapid personalized adaptation.

CN121893841APending Publication Date: 2026-04-21CHONGQING LANDIAN AUTOMOBILE TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CHONGQING LANDIAN AUTOMOBILE TECHNOLOGY CO LTD
Filing Date
2026-01-04
Publication Date
2026-04-21

AI Technical Summary

Technical Problem

The equipment control efficiency in the vehicle cabin is low, requiring a lot of manual operation, and cannot achieve rapid personalization.

Method used

By acquiring images of target objects in the cockpit, detecting the planar position and depth information of feature points, and combining the spatial position to determine the size information of the target object, the cockpit equipment is automatically adjusted to achieve personalized adaptation.

Benefits of technology

It improves cockpit control efficiency, reduces user manual operation time and complexity, and enables rapid and personalized equipment adjustments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121893841A_ABST
    Figure CN121893841A_ABST
Patent Text Reader

Abstract

The invention relates to a vehicle cabin control method and device, a controller and a vehicle. The method comprises the following steps: acquiring an image acquired for a target object in a cabin, and detecting plane positions of a plurality of preset feature points of the target object from the image; determining the depth information of each feature point, and for each feature point, obtaining the spatial position of the feature point according to the plane position and the depth information of the feature point; determining size information of the target object based on the spatial position of each feature point; and controlling the equipment in the cabin based on the size information. By adopting the method, the control efficiency of the cabin can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of vehicle control technology, and in particular to a vehicle cockpit control method, device, controller, and vehicle. Background Technology

[0002] In the field of vehicle cabin technology, seats and HUDs (Head-Up Displays) are key components in enhancing the driving and riding experience, and their adjustment functions are constantly evolving. Related technologies typically offer seats with diverse functions such as multi-directional electric adjustment, massage, ventilation, and heating to meet the comfort needs of different users; HUDs can also project important driving information onto the windshield, reducing the driver's need to look away. However, many devices in the vehicle cabin still rely on manual operation by vehicle occupants, resulting in relatively low efficiency in cabin control. Summary of the Invention

[0003] Therefore, it is necessary to provide a cockpit control method, device, controller, and vehicle that can improve cockpit control efficiency in response to the above-mentioned technical problems.

[0004] In a first aspect, this application provides a vehicle cockpit control method, comprising:

[0005] Acquire images of target objects in the cockpit, and detect the planar positions of multiple preset feature points of the target objects from the images;

[0006] Determine the depth information of each feature point, and for each feature point, obtain its spatial position based on its planar position and depth information;

[0007] Based on the spatial location of each feature point, determine the size information of the target object;

[0008] Control the equipment in the cockpit based on size information.

[0009] Secondly, this application also provides a vehicle cockpit control device, comprising:

[0010] The planar position determination module is used to acquire images of target objects in the cockpit and detect the planar positions of multiple preset feature points of the target objects from the images;

[0011] The spatial location determination module is used to determine the depth information of each feature point, and for each feature point, obtain the spatial location of the feature point based on its planar position and depth information;

[0012] The size information determination module is used to determine the size information of the target object based on the spatial position of each feature point;

[0013] The equipment control module is used to control the equipment in the cockpit based on dimensional information.

[0014] Thirdly, this application also provides a controller, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in the method provided in the first aspect above.

[0015] Fourthly, this application also provides a vehicle, including a cockpit and a controller provided by the third party.

[0016] Fifthly, this application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the method provided in the first aspect above.

[0017] Sixthly, this application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps in the method provided in the first aspect above.

[0018] The aforementioned cockpit control method, device, controller, and vehicle acquire images of a target object in the cockpit, detect the planar positions of multiple feature points of the target object from the images, determine the spatial position of each feature point by combining its planar position and depth information, and then determine the size information of the target object based on the spatial positions of multiple feature points. Furthermore, the cockpit equipment is controlled according to the size information of the target object to achieve personalized cockpit equipment adjustments for the target object. In this cockpit control method, the fusion of planar position and depth information enables precise three-dimensional positioning of feature points, the spatial position enhances the reliability of size information recognition, and the automatic adjustment of cockpit equipment based on size information achieves rapid personalized adaptation, reducing the time and complexity of manual operation for users, thereby improving the control efficiency of the cockpit. Attached Figure Description

[0019] To more clearly illustrate the technical solutions in the embodiments of this application or related technologies, the drawings used in the description of the embodiments of this application or related technologies will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0020] Figure 1 This is a flowchart illustrating a vehicle cockpit control method in one embodiment;

[0021] Figure 2 This is a schematic diagram of the process for determining depth information in one embodiment;

[0022] Figure 3This is a schematic diagram illustrating the mapping operation to a second image in one embodiment;

[0023] Figure 4 This is a schematic diagram illustrating the mapping operation to a first image in one embodiment;

[0024] Figure 5 This is a flowchart illustrating the cockpit control method in another embodiment;

[0025] Figure 6 This is a schematic diagram of various feature points of the human body in one embodiment;

[0026] Figure 7 This is a schematic diagram of the feature point regression network in one embodiment;

[0027] Figure 8 This is a structural block diagram of the vehicle's cockpit control device in one embodiment;

[0028] Figure 9 This is a diagram of the internal structure of the controller in one embodiment. Detailed Implementation

[0029] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0030] It should be noted that the terms "first," "second," etc., used in this application can be used to describe various elements, but these elements are not limited by these terms. These terms are only used to distinguish the first element from the second element. The terms "comprising" and "having," and any variations thereof, used in this application, are intended to cover non-exclusive inclusion. The term "multiple" used in this application refers to two or more. The term "and / or" used in this application refers to one of the embodiments, or any combination of multiple embodiments.

[0031] In one exemplary embodiment, such as Figure 1 As shown, a vehicle cockpit control method is provided, which is executed by a controller, such as a cockpit domain controller in the vehicle. In this embodiment, the method is described using a controller in a vehicle as an example, and includes the following steps S101 to S104. Wherein:

[0032] Step S101: Acquire an image of the target object in the cockpit, and detect the planar position of multiple preset feature points of the target object from the image.

[0033] The cabin is the space inside a vehicle for the driver and passengers. It may include, but is not limited to, interior structures (such as the dashboard, seats, and steering wheel), and is equipped with cameras and display devices for sensing. For example, the cabin in a vehicle may include the driver's cabin (the operating space for the front-seat driver) and the passenger cabin (the area for rear-seat passengers). The target object can be an individual located within the vehicle cabin who requires personalized services; such as the driver or passengers. Images can be captured from the target objects within the cabin, for example, through cameras deployed within the cabin.

[0034] Feature points are representative location points with clear physiological and spatial significance selected from body parts of a target object. These may include, but are not limited to, key points selected from at least one part of the target object, such as the head, neck, shoulder, upper arm, leg, and torso. In some embodiments, feature points may include upper body feature points, which are location points selected from the upper half of the target object's body, such as key points selected from at least one part of the target object, such as the head, neck, shoulder, and upper arm. Planar position is the location of a feature point in the two-dimensional coordinate system of the image. Planar position may include two-dimensional coordinate pairs, such as (u, v), where u represents the horizontal direction (column) and v represents the vertical direction (row). Planar position can be used to describe the projected position of the corresponding feature point on the image plane. Planar position itself does not contain depth information. When combined with depth information, the planar position can be back-projected into three-dimensional space through a camera model, thereby reconstructing the three-dimensional spatial position of the feature point.

[0035] For example, the controller can acquire images of a target object within the cockpit. For instance, the controller can acquire images from cameras deployed within the cockpit via the vehicle's internal bus or interface. These cameras may include, but are not limited to, pre-existing driver monitoring system cameras or occupant monitoring system cameras. The controller can perform feature point detection based on the acquired images to detect multiple preset feature points of the target object, thereby obtaining the planar position of each feature point. For example, the controller can use a pre-trained neural network model to perform detection on the image to determine the planar positions of multiple preset feature points of the target object.

[0036] Step S102: Determine the depth information of each feature point, and for each feature point, obtain the spatial position of the feature point based on its planar position and depth information.

[0037] Depth information is used to characterize the distance between a feature point and the camera along the optical axis. It determines how far a feature point is from the camera. Each detected feature point has a unique depth information, and by combining this depth information with its planar position, the spatial position of the feature point in three-dimensional space can be reconstructed. Depth information can include scalar values; for example, the depth information for a neck feature point could be 800 mm, indicating that the feature point is approximately 80 cm from the camera. Spatial position describes the location of a feature point in three-dimensional space. This can include the coordinates of the feature point in a three-dimensional coordinate system (i.e., the camera coordinate system) with the camera as the origin. For example, spatial position can be a three-dimensional coordinate vector containing X, Y, and Z components. The X and Y coordinates describe the left-right and up-down positions of the feature point in the imaging plane perpendicular to the optical axis, while the Z coordinate represents the depth information, describing the front-back position along the optical axis. The X, Y, and Z components together uniquely determine the precise position of a feature point in three-dimensional space.

[0038] Optionally, for each feature point of the target object, the controller can determine the depth information corresponding to each feature point. For example, the controller can determine the depth information of each feature point based on a binocular camera or a Time-of-Flight (ToF) camera. In some embodiments, for a binocular camera, the controller can use a stereo matching algorithm to find the same feature points in another image captured by the second camera, calculate the disparity, and then convert the disparity to obtain the depth information of the feature points. For a ToF camera, the controller can directly determine the depth information corresponding to the feature points based on their planar positions. For each feature point, the controller can reconstruct its spatial position in three-dimensional space based on its planar coordinates and depth information.

[0039] Step S103: Determine the size information of the target object based on the spatial location of each feature point.

[0040] Size information can be used to characterize the size of the target object. For example, size information may include, but is not limited to, at least one of the following: the target object's height, head and neck height, shoulder width, etc.

[0041] For example, the controller can perform size estimation based on the spatial location of each feature point of the target object to determine the size information of the target object, such as the height information of the target object. For instance, the controller can use a pre-trained size estimation model to perform size estimation based on the spatial location of each feature point in the target object to obtain the size information of the target object, such as the height information of the target object.

[0042] In some embodiments, the controller can perform geometric calculations based on the spatial position of each feature point to determine at least one body feature of the target object, and perform size estimation based on the at least one body feature of the target object to determine the size information of the target object, such as determining the height information of the target object. For example, the controller can determine at least one upper body feature of the target object, and perform size estimation based on the at least one upper body feature of the target object to determine the size information of the target object. The body feature can be a feature of the target object that can be used to determine the body size, such as including but not limited to at least one of the following: head and neck height (three-dimensional spatial straight-line distance between the top of the head feature point and the neck feature point), shoulder width (three-dimensional spatial straight-line distance between the left shoulder feature point and the right shoulder feature point), torso height (vertical distance from the neck feature point to the theoretical contact surface between the occupant's torso and the seat back), leg length, or arm length (distance from the shoulder point to the elbow point). The body feature can be determined based on the feature points of the target object, such as by geometric calculation based on the spatial position between the various feature points. Different body features can be calculated from the spatial positions of different feature points. In some embodiments, for each body feature, a corresponding calculation rule can be preset, which can define the feature points required for the calculation of each body feature. The controller can calculate the corresponding body feature based on the spatial location of the required feature points of the body feature by means of distance calculation, such as Euclidean distance.

[0043] In some embodiments, when the controller estimates the size of a target object based on its body characteristics to determine the target object's size information, the controller can use a pre-trained size estimation model to estimate the size based on the target object's body characteristics to obtain the target object's size information, such as estimating the target object's height. In some embodiments, the controller can also obtain a preset size feature mapping relationship, which may include the correspondence between different body characteristics and corresponding size information. The controller can directly determine the target object's size information based on the target object's body characteristics and the size feature mapping relationship. The size feature mapping relationship can be pre-constructed by analyzing the size information and body characteristics of each object. In some embodiments, different objects may also correspond to different size feature mapping relationships, such as objects of different ages or genders, for which corresponding size feature mapping relationships can be preset. The controller can determine a size feature mapping relationship applicable to the target object based on at least one of the target object's age or gender, and determine the size information in combination with the target object's body characteristics.

[0044] Step S104: Control the equipment in the cockpit based on the size information.

[0045] The equipment in the cockpit may include devices that allow for adjustments to the cockpit's operating status or output content for the occupants, such as at least one of the following: seats, head-up displays (HUDs).

[0046] Optionally, the controller can refer to the size information of the target object to control the equipment in the cockpit. For example, it can adjust the seats in the cockpit (adjust the height and / or fore-aft position of the seats) or adjust the HUD in the cockpit (adjust the height and / or angle of the HUD) so that the equipment in the cockpit can be adapted to the target object, thereby achieving personalized cockpit adaptation for the target object.

[0047] In the aforementioned vehicle cockpit control method, images of target objects in the cockpit are acquired, and the planar positions of multiple feature points of the target object are detected from the images. The spatial position of the feature points is determined by combining the depth information of each feature point, thereby determining the size information of the target object for controlling the equipment in the cockpit. The fusion of planar position and depth information enables precise three-dimensional positioning of feature points. The reliability of size information recognition is enhanced based on spatial position. The equipment in the cockpit is automatically adjusted according to the size information, achieving rapid personalized adaptation, reducing the time and tediousness of manual operation by users, and thus improving the control efficiency of the cockpit.

[0048] In one exemplary embodiment, the images include a first image and a second image captured by a binocular camera mounted in the cockpit; such as Figure 2 As shown, the process for determining depth information, i.e., determining the depth information of each feature point, includes steps S201 to S203. Wherein:

[0049] Step S201: Obtain the positional mapping relationship between the first image and the second image. The positional mapping relationship is pre-calibrated for the binocular camera.

[0050] A binocular camera is an imaging device that simulates the principle of human binocular stereoscopic vision. It uses two cameras positioned a certain distance apart (baseline) in the horizontal direction to simultaneously capture two images of the same scene from slightly different perspectives to obtain depth information. The binocular camera can be installed in the vehicle's cabin, such as on the A-pillar, in the rearview mirror area, or in the roof. The binocular camera can include a left and a right camera with similar or identical optical parameters to simultaneously capture images of target objects such as the driver or passengers. The first and second images can be two related images captured simultaneously by the two cameras in the same binocular camera, but with slightly different perspectives. The main subject in both the first and second images can be the target object in the cabin. Due to the difference in perspective, the horizontal pixel position of the same feature point on the target object will differ between the first and second images.

[0051] Position mapping relationships can be used to describe the geometric correspondence between the imaging planes of two cameras in a stereo camera system. These relationships can be obtained through pre-calibration of the stereo cameras. For example, the position mapping relationship can include the intrinsic parameter matrix, distortion coefficients, and extrinsic parameter matrix describing the relative positional relationship between the two cameras for each camera. The intrinsic parameter matrix can include, but is not limited to, focal length and principal point coordinates, used to describe the imaging geometry of the camera itself; the distortion coefficients can be used to correct image distortion introduced by the lens; and the extrinsic parameter matrix can include, but is not limited to, rotation matrix R and translation vector T.

[0052] Optionally, a binocular camera can be installed in the vehicle's cabin to capture images of target objects within the cabin, obtaining a first image and a second image. The first image can be captured by the first camera in the binocular camera system, and the second image can be captured by the second camera in the binocular camera system. The controller can determine the positional mapping relationship between the first image and the second image, thereby determining the geometric correspondence between the first image and the second image. In some embodiments, the positional mapping relationship may include parameters such as intrinsic parameters, distortion coefficients, and extrinsic parameters of the binocular camera.

[0053] Step S202: For each feature point, determine the disparity between the feature point and the first image and the second image based on the position mapping relationship.

[0054] Parallax can include the difference in horizontal coordinates between the corresponding projection points of the same feature point in space, in the corrected first and second images. Parallax directly reflects the distance of the feature point from the camera. For example, for a feature point in an image (such as a detected key point on the top of the head), after determining its coordinates in the first image, its corresponding feature point is searched on the same horizontal row in the second image according to the row alignment characteristics in the position mapping relationship. The difference in the horizontal coordinates of these two corresponding feature points is the parallax of that feature point.

[0055] For example, the controller can determine the disparity of each feature point between the first and second images based on the position mapping relationship. In some embodiments, the controller can perform stereo correction on the first and second images using the position mapping relationship, such as distortion correction (eliminating barrel / pincushion distortion) and epipolar correction, so that the first and second images are transformed onto the same virtual, coplanar, and row-aligned plane. In this case, any point in the first image has a corresponding point in the second image located on the same image row (i.e., the same vertical coordinate v). The controller can determine the planar coordinates of key points detected in the first image. For each feature point (u_left, v) in the first image, the controller can perform stereo matching in the stereo-corrected second image. For example, it can slide a small image patch centered on the feature point within a preset disparity range on the v-th row of the second image to find the most similar image patch. The controller can calculate the similarity between the small image patch and candidate image patches in the second image and determine the image patch with the highest similarity. The controller can obtain the disparity of the feature point between the first and second images based on the difference between the horizontal coordinates u_right and u_left of the most similar image patch. After traversing each feature point, the disparity of each feature point between the first and second images can be determined.

[0056] Step S203: Obtain the depth information of the feature points based on their planar position and disparity.

[0057] Optionally, for each feature point, the controller can calculate the depth information of that feature point based on its planar position and disparity. For example, the controller can calculate the depth information of each feature point based on the depth Z = (focal length f × baseline length B) / disparity d. Here, focal length f and baseline B are known inherent parameters of the binocular camera.

[0058] In some embodiments, the image includes a first image and a second image captured by a binocular camera installed in the cockpit. The controller can detect the same feature point in both the first and second images. When determining the planar position of the same feature point A, the controller can select a target image from the first and second images and detect the planar position of feature point A from the target image. For example, the controller can use the first image captured by the driver monitoring system camera as the target image to determine the planar positions of multiple preset feature points of the target object, and determine the depth information of the feature point based on the planar position and parallax. In some embodiments, the controller can also perform feature point detection on the first and second images separately to detect the planar position of each feature point from both images. That is, each feature point can detect two planar positions, and the controller can fuse the two planar positions to obtain the final planar position of the feature point. For example, the controller can detect planar position 1 of feature point A from the first image and planar position 2 of feature point A from the second image. The controller can perform weighted fusion of planar position 1 and planar position 2 to obtain the planar position of feature point A. In some embodiments, the weights for weighted fusion can be pre-calibrated and determined for the binocular camera.

[0059] In this embodiment, the controller can determine the parallax of each feature point between the first and second images based on the positional mapping relationship between the first and second images captured by the binocular camera, and determine the corresponding depth information by combining the planar position of the feature point and the parallax. It can accurately determine the depth information of the feature point without adding a dedicated depth sensor, and can achieve high-precision, non-contact depth perception.

[0060] In an exemplary embodiment, determining the disparity of a feature point between a first image and a second image based on a position mapping relationship includes: obtaining a first planar position of the feature point in the first image and a second planar position of the feature point in the second image; mapping the first planar position to the second image based on the position mapping relationship to obtain a first disparity of the feature point; mapping the second planar position to the first image based on the position mapping relationship to obtain a second disparity of the feature point; and obtaining the disparity of the feature point between the first image and the second image based on the first disparity and the second disparity.

[0061] In this context, the first and second planar positions refer to the same feature point, specifically its corresponding planar positions in the two images captured by the binocular camera. These positions can be two-dimensional coordinates in pixels, such as (u, v), where u represents the horizontal (column) coordinate and v represents the vertical (row) coordinate. In an ideal binocular camera with stereo correction, the first planar position of the same feature point in the first image and the second planar position in the second image should have the same v coordinate (i.e., be in the same row), while the u coordinate should differ. This difference represents the disparity between the feature point in the first and second images. The first disparity is the pixel displacement in the horizontal direction between two corresponding positions calculated using a stereo matching algorithm to search for the corresponding position in the second image, based on the feature point's position in the first image. The second disparity is the pixel displacement calculated by searching for the corresponding position in the first image, based on the feature point's position in the second image. The first disparity represents how much the matched feature point has shifted in the second image from the perspective of the first image; the second disparity represents how much the matched feature point has shifted in the first image from the perspective of the second image. The disparity of each feature point can be determined using the first and second disparities. For example, consistency verification can be performed using the first and second disparities to determine the disparity of the feature point. This disparity characterizes the horizontal shift in the imaging position of the same feature point in the first and second images due to the different horizontal positions of the binocular cameras.

[0062] Optionally, for each feature point, the controller can determine the feature's first planar position in the first image and its second planar position in the second image. For example, the controller can perform feature point detection on the first and second images respectively, to detect the first planar position of the feature point in the first image and the second planar position of the feature point in the second image. The controller can then map the first planar position of the feature point to the second image based on a position mapping relationship to perform forward stereo matching and obtain the first disparity of the feature point.

[0063] For example, for feature point A, its first plane position in the first image is P1(u_left,v), and its second plane position in the second image is P2(u_right,v). Based on the position mapping relationship, the controller knows that the point corresponding to the first plane position P1 must be located in row v of the second image. The controller can take a small image block (template) centered on P1 and perform sliding matching within a preset disparity range on row v of the second image. The controller calculates the similarity between the template and each candidate image block in the second image. After finding the candidate image block P1' with the highest similarity, the difference between its horizontal coordinate and u_left is the first disparity d_lr of feature point A. The controller can also perform reverse matching based on the position mapping relationship to obtain the second disparity of feature point A. For example, the controller can map the second plane position P2(u_right,v) of feature point A to the first image based on the position mapping relationship to obtain the second disparity of feature point A.

[0064] In some embodiments, such as Figure 3 As shown, for feature point M, its first planar position in the first image is L1. The controller can map the first planar position L1 to the second image based on the position mapping relationship. Specifically, it can perform stereo matching in the second image based on the first planar position L1 of feature point M to determine the planar position L1' corresponding to L1 in the second image. The controller can obtain the first disparity of feature point M based on the difference in horizontal coordinates between L1 and L1'. In some embodiments, such as... Figure 4 As shown, for feature point M, its second plane position in the second image is L2. The controller can map the second plane position L2 to the first image based on the position mapping relationship. Specifically, it can perform stereo matching in the first image based on the second plane position L2 of feature point M to determine the plane position L2' corresponding to L2 in the first image. The controller can obtain the second disparity of feature point M based on the difference in horizontal coordinates between L2 and L2'.

[0065] The controller can synthesize the first parallax and the second parallax of the feature points to obtain the parallax of the feature points between the first image and the second image. For example, the controller compares the first parallax d_lr and the second parallax d_rl of the same feature point. In an ideal situation, d_lr≈-d_rl should hold. The controller can set a consistency threshold T. If |d_lr + d_rl| < T, it is considered that the two matching results are consistent, and this match is reliable. The controller can take the average of d_lr and -d_rl, or directly use d_lr as the final parallax d_final of this feature point. If the two results are inconsistent and exceed the threshold, it indicates that the match may fail at this feature point (such as due to occlusion or weak texture). The controller can mark this feature point as an invalid matching point to redetect this feature point, such as reshooting an image for feature point detection to redetermine the parallax.

[0066] In this embodiment, for each feature point, the controller can perform forward matching and backward matching through the first image and the second image respectively, and determine the parallax by synthesizing the obtained first parallax and second parallax. By performing forward and backward matching twice and cross-verifying, it can efficiently screen out those points with high matching quality and consistency, while identifying and processing unreliable matching points, thus ensuring the accuracy of the parallax.

[0067] In an exemplary embodiment, the depth information of the feature points is obtained based on the planar position and parallax of the feature points, including: obtaining the original depth information of the feature points according to the planar position and parallax of the feature points; acquiring the historical depth information of the feature points, where the historical depth information of the feature points is detected based on historical images collected for the target object; obtaining the depth information of the feature points according to the historical depth information and the original depth information of the feature points.

[0068] Among them, the original depth information is the depth information of the feature points determined based on the current frame image, and the original depth information represents the straight-line distance between the feature point and the camera optical center from the perspective of the current frame image. The historical depth information is the depth information detected by the feature points from a series of consecutive or non-consecutive historical images before the current moment, that is, the historical depth information is the depth information of the feature points determined based on historical images, and the historical depth information represents the straight-line distance between the feature point and the camera optical center from the perspective of the historical frame image.

[0069] Optionally, the controller can determine the planar position and disparity of feature points based on the current images (first image and second image), and obtain the original depth information of the feature points based on the planar position and disparity. The controller can acquire historical depth information of the feature points, which can be obtained by performing depth information determination processing based on historical images. The controller can combine the historical depth information and the original depth information of the feature points to obtain the depth information of the feature points. In some embodiments, the controller can perform feature point tracking based on historical images and the current image to associate the same feature point in consecutive images, thereby obtaining the depth information of each feature point based on historical depth information and original depth information. For example, the controller can use a Kalman filter method to treat the original depth information as a noisy observation, use the current depth value predicted based on historical depth information as the predicted value, and then perform optimal fusion based on the confidence levels of both to obtain the depth information of the feature points. In some embodiments, the controller can perform smooth interpolation based on historical depth information using a filtering algorithm (such as first-order low-pass filtering or moving average) to ensure the stability and reliability of the output depth information. For example, the controller can use a moving average filtering method to average the original depth information and historical depth information, and use the average value as the depth information of the feature points.

[0070] In this embodiment, the controller can determine the depth information by integrating the historical depth information of the feature points, which can reduce the noise of the depth information, maintain the continuity and reliability of the depth information output, and thus ensure the accuracy of the depth information.

[0071] In an exemplary embodiment, obtaining the depth information of a feature point based on its planar position and disparity includes: obtaining the original depth information of the feature point based on its planar position and disparity; and determining the depth information of the feature point based on the feature point attribute relationship of the target object and the original depth information.

[0072] The original depth information is the depth information of feature points determined based on the current frame image. It represents the straight-line distance between the feature points and the optical center of the camera from the current frame's viewpoint. The feature point attribute relationships are the spatial and logical constraints between feature points, determined by the inherent characteristics of the target object's physiological structure and its motion laws. For example, feature point attribute relationships may include the constraint that two different feature points on the same rigid torso should have similar depth information.

[0073] Optionally, the controller can determine the planar position and disparity of feature points based on the current images (first image and second image), and obtain the original depth information of the feature points based on the planar position and disparity. The controller can combine the feature point attribute relationships of the target object with the original depth information to determine the depth information of the feature points. For example, the controller can remove feature points with abnormal original depth information based on the feature point attribute relationships of the target object; or, the controller can smooth and correct the original depth information of the feature points based on the feature point attribute relationships of the target object to obtain the depth information of each feature point.

[0074] In some embodiments, the controller can also synthesize the feature point attribute relationships of the target object, the historical depth information of the feature points, and the original depth information to obtain the feature point depth information. For example, the controller can synthesize the historical depth information and the original depth information of the feature points to obtain the intermediate depth information of the feature points, and obtain the feature point depth information based on the feature point attribute relationships of the target object and the intermediate depth information.

[0075] In this embodiment, the controller can determine depth information by comprehensively considering the relationship between the feature points of the target object, which can reduce the noise of the depth information and thus improve the accuracy and reliability of the depth information.

[0076] In an exemplary embodiment, detecting the planar positions of multiple preset feature points of a target object from an image includes: identifying the image region where the target object is located in the image through a pre-trained target detection network; and detecting the planar positions of multiple preset feature points of the target object from the image region through a pre-trained regression network.

[0077] The object detection network and regression network can include pre-trained deep learning models. The object detection network can be used for coarse localization to determine the image region where the target object is located; the regression network can be used for fine localization to determine the planar positions of the feature points of the target object. In some embodiments, the object detection network and regression network can be integrated, such as combining them to obtain a feature point detection network for feature point detection. In some embodiments, the object detection network and regression network can also be set independently, that is, coarse localization and fine localization can be performed sequentially by the object detection network and regression network, thereby achieving feature point detection.

[0078] Optionally, the controller can acquire a pre-trained object detection network and use it to detect target objects in the image, thereby identifying the image region where the target object in the cockpit is located. For example, the object detection network can output one or more bounding boxes, each corresponding to a target object in the image, and this bounding box can serve as the image region where the corresponding target object is located. The controller can use a pre-trained regression network to perform feature point detection on the image region, thereby detecting the planar positions of multiple preset feature points of the target object in the image region where the target object is located. For example, the regression network can output the planar positions of multiple preset feature points of the target object, where the planar positions can include two-dimensional coordinates (x, y).

[0079] In this embodiment, the controller can determine the image region where the target object is located through the target detection network, and detect the planar position of feature points from the image region through the regression network. This top-down detection method can improve the accuracy of feature point detection.

[0080] In an exemplary embodiment, detecting the planar positions of multiple preset feature points of a target object from an image includes: detecting the original planar positions of multiple preset feature points of the target object from the image; determining the historical planar position of each feature point, wherein the historical planar position of each feature point is obtained based on historical images acquired for the target object; and for each feature point, obtaining the planar position of the feature point based on the historical planar position and the original planar position of the feature point.

[0081] The original planar position is the planar position of the feature point determined based on the current frame image, representing the planar position of the feature point from the perspective of the current frame image. The historical planar position is the planar position of the same feature point determined based on historical images acquired before the current frame, reflecting the planar position of the feature point from the perspective of historical images.

[0082] For example, the controller can perform detection based on the currently acquired image to determine the original planar position of feature points of a target object in the cockpit. The controller can acquire the historical planar positions of the feature points, which can be obtained by feature point detection processing based on historical images acquired for the target object. The controller can combine the historical planar positions and the original planar positions of the feature points to obtain the planar position of the feature points. In some embodiments, the controller can employ a first-order low-pass filter (such as exponential smoothing) or a Kalman filter to obtain the planar position of the feature points based on the historical planar positions and the original planar positions. For example, the controller can determine the planar position of the feature points based on the final planar position = α × original planar position + (1-α) × historical planar position; where α is a weighting coefficient between 0 and 1. When the confidence of the original planar position is high, α is larger; when the original planar position is occluded or blurred, α becomes smaller, relying more on the prediction of historical trajectories.

[0083] In this embodiment, the controller can determine the planar position by comprehensively considering the historical planar positions of feature points, which can reduce the noise of the planar position, maintain the continuity and reliability of the planar position output, and thus ensure the accuracy of the planar position.

[0084] In an exemplary embodiment, determining the size information of the target object based on the spatial location of each feature point includes: determining the upper body features of the target object based on the spatial location of each feature point, wherein the upper body features include at least one of head and neck height, shoulder width, torso height, or arm length; and obtaining the size information of the target object by predicting based on the upper body features using a size prediction model corresponding to the target object.

[0085] The upper body features can be any features of the target object's upper body that can be used to determine body dimensions, such as, but not limited to, at least one of the following: head and neck height (the three-dimensional straight-line distance between the top of the head and the neck), shoulder width (the three-dimensional straight-line distance between the left and right shoulder features), torso height (the vertical distance from the neck feature point to the theoretical contact surface between the occupant's torso and the seat back), or arm length (the distance from the shoulder point to the elbow point). Upper body features can be determined based on feature points of the target object's upper body, such as by using geometric calculations based on the spatial positions of various feature points. Different upper body features can be calculated using the spatial positions corresponding to different feature points.

[0086] Head and neck height refers to the straight-line distance in three-dimensional space between the top feature point of the head and the neck feature point of the target object (driver or passenger). Head and neck height characterizes the vertical dimension of the human head and upper neck. Shoulder width is the straight-line distance in three-dimensional space between the left and right shoulder feature points of the target object. Shoulder width represents the lateral dimension of the uppermost part of the human torso. Torso height is the vertical distance from the neck feature point of the target object to the estimated contact point between the target object and the seat back. Torso height reflects the length of the torso segment from the neck to the waist and hips when the target object is seated. Arm length may include the arm segment length information of the target object calculated based on shoulder and elbow feature points. The size prediction model is used to predict the size of the target object based on the upper body features of the target object to obtain the size information of the target object. In some embodiments, the size prediction model may include a linear regression model and may also include a pre-trained neural network model.

[0087] Optionally, the controller can determine the upper body features of the target object based on the spatial location of each feature point. These upper body features may include at least one of head and neck height, shoulder width, torso height, or arm length. For example, for each upper body feature, a corresponding calculation rule can be pre-set, defining the feature points required for each upper body feature calculation. The controller can then calculate the corresponding upper body feature based on the spatial location of each required feature point, using distance calculations, such as Euclidean distance.

[0088] The controller can acquire a size prediction model for a target object. This model can correspond to a specific target object; different target objects may correspond to different size prediction models. For example, target objects of different ages or genders may correspond to different size prediction models. In some embodiments, the controller can perform attribute recognition on the target object based on an image to determine its age, gender, and other attribute information. Based on this attribute information, the controller can acquire a size prediction model for the target object. In some embodiments, the correspondence between attribute information and size prediction models can be preset, and the controller can determine the size prediction model for the current target object based on this correspondence. The controller can use the size prediction model to predict the target object's size based on its upper body features, such as its height. For example, the controller can input at least one of head and neck height, shoulder width, torso height, or arm length into the size prediction model to obtain the target object's size information.

[0089] In some embodiments, the controller may also acquire a preset size feature mapping relationship, which may include the correspondence between different upper body features and corresponding size information. The controller can directly determine the size information of the target object based on the upper body features and the size feature mapping relationship. The size feature mapping relationship can be pre-constructed by analyzing the size information and upper body features of each object. In some embodiments, different objects may correspond to different size feature mapping relationships, such as objects of different ages or genders, for which corresponding size feature mapping relationships can be preset. The controller can determine the size feature mapping relationship applicable to the target object based on at least one of the target object's age or gender, and determine the size information in conjunction with the target object's upper body features.

[0090] In this embodiment, the controller can accurately determine the size information of the target object by using a size prediction model based on at least one upper body feature, such as head and neck height, shoulder width, torso height, or arm length. This allows the controller to automatically adjust the equipment in the cockpit according to the size information, achieving rapid and personalized adaptation. This reduces the time and tediousness of manual operation for users, thereby improving the control efficiency of the cockpit.

[0091] In one exemplary embodiment, controlling devices in a cockpit based on size information includes: determining a first control parameter for a seat in the cockpit based on the size information, and adjusting the seat according to the first control parameter.

[0092] The first control parameter can be used to control the seats in the cockpit, and may include, but is not limited to, at least one of the following: seat height parameter, seat rail fore-and-aft parameter, seat back angle parameter, and steering wheel telescopic / tilt angle parameter. Optionally, the controller can determine the first control parameter for the seats in the cockpit based on the size information of the target object. In some embodiments, the mapping relationship between different size information and different seat control parameters can be pre-configured, and the controller can determine the corresponding first control parameter based on the size information of the target object and the mapping relationship. The controller can adjust the seats according to the first control parameter so that the seats fit the target object, ensuring that the target object can obtain a sitting posture that is reasonably positioned and conforms to its body size without any manual operation.

[0093] In this embodiment, the controller can determine the first control parameters of the seat based on the size information of the target object, and adjust the seat according to the first control parameters, thereby achieving rapid personalized adaptation, reducing the time and tediousness of manual operation by the user, and thus improving the control efficiency of the cockpit.

[0094] In one exemplary embodiment, controlling devices in the cockpit based on size information includes: determining a second control parameter for a display device in the cockpit based on the size information, and adjusting the display device according to the second control parameter.

[0095] The second control parameter can be used to control the display device in the cockpit, and may include, but is not limited to, at least one of the following parameters: vertical offset parameter, brightness level parameter, etc. For example, the controller can determine the second control parameter for the display device in the cockpit based on the size information of the target object. In some embodiments, the mapping relationship between different size information and different display device control parameters can be pre-configured, and the controller can determine the corresponding second control parameter based on the size information of the target object and this mapping relationship. The controller can adjust the display device according to the second control parameter so that the display device is adapted to the target object, ensuring that the target object can obtain correctly positioned and comfortably viewed display information without any manual operation.

[0096] In this embodiment, the controller can determine the second control parameters of the display device based on the size information of the target object, and adjust the display device according to the second control parameters, thereby achieving rapid personalized adaptation, reducing the time and tediousness of manual operation by the user, and thus improving the control efficiency of the cockpit.

[0097] This application also provides an application scenario in which the above-mentioned vehicle cockpit control method is applied. Specifically, the application of the vehicle cockpit control method in this scenario is as follows:

[0098] In related technologies, car seats and HUDs (Head-Up Displays) typically require manual adjustment or rely on preset user profiles. While some models are equipped with memory functions, users still need to pre-set and select the corresponding mode. In recent years, with the widespread adoption of DMS (Driver Monitoring System) and OMS (Occupant Monitoring System), personalized settings can also be achieved through facial recognition (Face ID). However, this method has the following inherent drawbacks:

[0099] 1. It can only identify registered users and cannot handle new user scenarios;

[0100] 2. The weak correlation between facial features and height leads to insufficient adjustment precision;

[0101] 3. A large user database needs to be established, resulting in high implementation costs.

[0102] In addition, height can be measured directly using pressure sensors or laser rangefinders, but these methods have significant shortcomings:

[0103] 1. A dedicated sensor needs to be installed, increasing hardware costs and installation complexity;

[0104] 2. Due to cabin space limitations, the measuring device is susceptible to interference from factors such as seat deformation and occupant posture;

[0105] 3. Non-contact real-time monitoring is not possible.

[0106] While human pose estimation techniques in the field of computer vision can detect key points in two-dimensional images, they present unique challenges when directly applied to cockpit environments:

[0107] 1. The cramped cockpit space leads to severe perspective distortion;

[0108] 2. Occupants may be obstructed by objects such as the steering wheel or seat back;

[0109] 3. Differences in seat geometry parameters between different vehicle models affect the measurement benchmark.

[0110] Based on this, this application provides a vehicle cockpit control method that achieves high-precision height estimation without the need for additional hardware by fusing multi-dimensional information from existing in-vehicle vision systems. Specifically, it estimates human height through visual fusion for personalized smart cockpit settings. This method is applicable to real-time estimation of driver or passenger height using human key point data collected by in-vehicle DMS (Driver Monitoring System) / OMS (Occupant Monitoring System) cameras combined with depth information, thereby enabling adaptive adjustment of cockpit equipment such as seat position and HUD (Head-Up Display). The vehicle cockpit control method provided in this application is optimized for the specific scenario of the automotive cockpit, effectively solving practical problems such as spatial constraints and occlusion interference, and providing reliable data support for personalized smart cockpit adjustments. Furthermore, the vehicle cockpit control method provided in this application achieves non-contact height measurement by reusing image data acquired by DMS / OMS cameras and combining human key point detection and depth information fusion algorithms, which solves the problem of increased costs caused by existing cockpit occupant height estimation methods relying on additional sensors. Furthermore, by introducing depth estimation modules such as monocular / binocular / TOF (Time of Flight) and establishing multidimensional regression equations between human key points (head height, shoulder width, etc.) and height, the robustness of measurement in complex cockpit environments can be improved, overcoming the deficiency of insufficient estimation accuracy caused by statistical models ignoring three-dimensional spatial information.

[0111] In some embodiments, such as Figure 5As shown, in the vehicle cockpit control method provided in this application, feature point information (i.e., planar position) can be obtained based on RGB (Red-Green-Blue) images through a feature point extraction model, and depth information can be determined through a depth estimation model. The depth information can be obtained through TOF or a binocular camera. The controller can align and fuse the depth information and feature point information to determine the size information of the target object, such as the height of the target object. Based on the height of the target object, the controller can automatically control the seat motor through a height-sitting posture recommendation model to adjust the seat in the cockpit. The controller can also automatically control the HUD motor through a height-sitting posture-HUD model to adjust the HUD in the cockpit.

[0112] The vehicle cockpit control method provided in this application includes multiple preset feature points for detecting a target object, which may include upper body feature points. This allows for high-precision height estimation within a limited cockpit space, utilizing only the upper body feature points of the target object, and enables intelligent adaptation of the cockpit system. Specifically, this may include the following steps:

[0113] S1: Image acquisition and upper body feature point detection.

[0114] A dual-camera system deployed on the A-pillar or roof of the vehicle simultaneously acquires RGB images and depth information of the occupants. Using a pre-trained human pose estimation model, multiple human feature points of the occupants' upper bodies are accurately detected from the RGB images. In some embodiments, such as... Figure 6 As shown, the human body can include multiple feature points. In this application, the detected feature points can be limited to the upper body, such as the top of the head, neck, left shoulder, and right shoulder. Optionally, the left elbow, right elbow, etc., can also be included to increase the feature dimension.

[0115] S2: Recovery of the three-dimensional coordinates of upper body feature points.

[0116] Using the parallax principle of binocular cameras, the depth value (Z coordinate) of each upper body feature point detected in step S1 is calculated in three-dimensional space. By combining the 2D (Two-Dimensional) coordinates and depth information of the feature points in the RGB image, all upper body feature points are transformed from the image coordinate system to the three-dimensional camera coordinate system with the camera as the origin, and their three-dimensional coordinates (X,Y,Z) are obtained.

[0117] S3: Height estimation based on upper body geometric features.

[0118] The specific calculation process may include the following:

[0119] S3.1: Extraction of upper body feature dimensions.

[0120] In a 3D camera coordinate system, calculate one or more of the following upper body geometric feature dimensions:

[0121] 1. Head and neck height (H_head): Calculates the three-dimensional Euclidean distance between the feature points on the top of the head and the feature points on the neck.

[0122] 2. Shoulder width (W_shoulders): Calculates the three-dimensional Euclidean distance between the feature points of the left and right shoulders.

[0123] 3. Torso Height (H_torso): Estimates the vertical height from the neck to the seat back contact surface. This height can be indirectly calculated by the relationship between the three-dimensional coordinates of the neck feature points and a predefined seat back planar model (determined based on the seat adjustment angle).

[0124] 4. Arm length (L_arm): If an elbow feature point is detected, the magnitude of the arm length-related vector formed by "shoulder-elbow" or "shoulder-neck-elbow" can be calculated as an auxiliary feature.

[0125] S3.2: Construct a height estimation model.

[0126] A multiple regression model is established to map the extracted upper body geometric features to the estimated height (H_estimated) of the occupant. This model is represented as:

[0127] H_estimated=f(H_head,W_shoulders,H_torso,...)+C

[0128] Where: H_head, W_shoulders, H_torso, etc., are the upper body feature dimensions of the input; f(...) is a mapping function, which can be a linear regression function (such as a*H_head+b*W_shoulders+c*H_torso) or a nonlinear model (such as one trained through a neural network or support vector machine); C is a constant offset. The parameters of this model (such as a, b, c, C) are obtained in advance through machine learning training on a large-scale human body size dataset, which contains the correspondence between upper body dimensions and total height for individuals of different heights.

[0129] S4: Intelligent Adaptation of Cockpit Systems.

[0130] The estimated height H_estimated is used as input to drive the automatic adaptation of the cockpit system:

[0131] S4a: Height-Seat Fitting Model.

[0132] Establish the mapping relationship between height and ideal seat position parameters using the function F_seat(H):

[0133] [Seat_Height,Seat_Slide,Seatback_Angle,Steering_Column_Extension]=F_seat(H_estimated)

[0134] Wherein, F_seat(H) can be the seat and steering wheel adjustment function; Seat_Height is the seat height; Seat_Slide is the seat forward / backward sliding parameter; Seatback_Angle is the seat back angle; and Steering_Column_Extension is the steering column extension / retraction amount. Based on the model output, the controller automatically adjusts the seat and steering wheel to the optimal safe and comfortable position that matches the current occupant's height.

[0135] S4b: Height-HUD settings adapt to the model.

[0136] Establish the mapping relationship between height and HUD projection parameters using the function F_HUD(H):

[0137] [HUD_Vertical_Offset,HUD_Brightness_Level]=F_HUD(H_estimated)

[0138] Wherein, F_HUD(H) is the head-up display adjustment function; HUD_Vertical_Offset is the HUD vertical offset parameter; and HUD_Brightness_Level is the HUD brightness level. Based on the model output, the controller automatically adjusts the vertical offset angle of the HUD projector to ensure that the virtual image is always within the natural eye level of the occupant of that height.

[0139] In some embodiments, the feature point detection in step S1 aims to accurately and in real-time locate the upper body feature points necessary for height estimation from the images captured by the cockpit camera. Specifically, this includes:

[0140] S1.1: Network Model Architecture

[0141] A top-down feature point detection process based on a deep convolutional neural network is adopted, which includes two core sub-networks:

[0142] 1. Human body detection network:

[0143] Function: First, locate one or more bounding boxes containing the occupants in the input RGB image. This is especially important in OMS (Occupant Monitoring System) scenarios (such as when there are passengers in the front row) to distinguish between the driver and passengers.

[0144] Implementation: Lightweight object detection networks, such as YOLOv5s (You Only Look Once version 5small) or SSD-MobileNet (Single Shot Multi Box Detector-Mobile Net), are employed to ensure real-time performance on in-vehicle computing platforms. The network is trained to recognize the "human" category.

[0145] 2. Feature Point Regression Network:

[0146] Function: For each obtained human body bounding box, accurately regress the pixel coordinates of each feature point of the upper body.

[0147] Implementation: HRNet (High-Resolution Net) or its lightweight variant is used as the backbone network. The advantage of HRNet is that it consistently maintains high-resolution feature representations, rather than recovering details through downsampling and then upsampling, which is crucial for accurately locating small-scale feature points such as the top of the head and neck. Figure 7 As shown, the input layer of HRNet receives RGB images for detecting upper body feature points. The high-resolution backbone of the HRNet terminal includes basic convolutional units (Conv.Unit) and current-stage channel feature maps. In the multi-resolution stage, four parallel rectangular regions represent four multi-resolution processing stages. For each stage, there is a high-resolution branch (TopBranch), middle-resolution branches, a fusion module, and output heads. The network's final output layer is a series of heatmaps, each corresponding to a specific feature point (e.g., "left shoulder"). The pixel with the highest value in the heatmap is the predicted location of that feature point. This method is more spatially robust than directly regressing coordinates.

[0148] S1.2: Feature point definition and upper body topology

[0149] This application defines a set of feature points specifically optimized for estimating the upper body height in the cockpit, containing multiple core feature points whose connections constitute the upper body topological skeleton:

[0150] Head (1 point):

[0151] Top of the head: The highest point of the head, which is the key endpoint for estimating the height of the head and neck.

[0152] Torso (3 points):

[0153] Neck: Located directly below the chin where it connects to the torso, it is usually defined as an offset point above the midpoint of the left and right shoulders, and is the core hub connecting the head and torso.

[0154] Left shoulder: The point where the left arm connects to the torso.

[0155] Right shoulder: The point where the right arm connects to the torso.

[0156] (The upper part of the torso is formed by the left shoulder, right shoulder and neck, and is used to calculate shoulder width and infer torso posture).

[0157] Arm (2 points):

[0158] Left elbow and right elbow: As auxiliary feature points, they can be used to improve the stability of pose estimation and introduce arm length-related features into the model to increase the dimension of height estimation.

[0159] S1.3: Model Optimization in Cockpit Scenarios (Data Augmentation and Training)

[0160] To ensure that the general feature point detection model can perfectly adapt to the complex in-vehicle environment, the following targeted optimizations were performed:

[0161] Data augmentation: During the model training phase, extensive use of augmentation strategies simulating cockpit environments is employed to improve the model's robustness. This includes:

[0162] Lighting variations: Simulates nighttime driving, tunnel entry and exit, and strong light reflection from car windows.

[0163] Occlusion simulation: Randomly simulates occlusion of the upper body area by the steering wheel, seat headrests, interior items, or even another occupant.

[0164] Posture diversity: Ensure that training data includes various driving and riding postures, such as leaning forward, leaning back, and turning sideways.

[0165] Domain-specific training: The network is pre-trained and fine-tuned using a large dataset collected and precisely labeled in real-world cockpit environments. This enables the model to learn prior knowledge of feature point locations in scenarios such as seatbelt entanglement and specific seat bolstering.

[0166] S1.4: Feature Point Tracking and Filtering (Post-processing)

[0167] To ensure the temporal stability and spatial smoothness of the output feature points and avoid jitter in single-frame detection, a post-processing algorithm is introduced:

[0168] Feature point tracking: Lightweight tracking algorithms (such as optical flow or Kalman filtering) are used to associate the same feature points in consecutive video frames. This not only eliminates jitter but also allows for quick and accurate re-locking of feature points after brief occlusion.

[0169] Confidence filtering: For feature points whose confidence in the network output is lower than a preset threshold, the detection results of the current frame are not used directly. Instead, the feature points are smoothed by interpolation through filtering algorithms (such as first-order low-pass filtering or moving average) based on their motion history to ensure that the output feature point coordinate stream is stable and reliable.

[0170] For step S2, feature point depth estimation can be achieved based on binocular vision. The purpose of this step is to convert the 2D pixel feature points detected in S1 into precise 3D coordinates (X, Y, Z) in a 3D camera coordinate system with the camera as the origin, where the Z coordinate is the depth value, representing the vertical distance from the feature point to the camera plane. Specifically, this includes:

[0171] S2.1: Binocular System Calibration and Image Correction

[0172] This forms the basis for all subsequent calculations, designed to eliminate lens distortion and align the images from the two cameras to an ideal state of coplanar alignment, including:

[0173] 1. Camera calibration:

[0174] By photographing a calibration board with a known pattern (such as a checkerboard), the intrinsic parameter matrices and distortion coefficients of the left and right cameras are obtained respectively.

[0175] The intrinsic parameter matrix contains core information such as focal length (fx, fy) and principal point coordinates (cx, cy), which are used to associate pixel coordinates with camera coordinates.

[0176] The distortion coefficient is used to correct image distortion caused by the optical characteristics of the lens.

[0177] 2. Stereo calibration and correction:

[0178] Through stereo calibration, extrinsic parameter matrices describing the relative positions of the left and right cameras are obtained, including the rotation matrix R and the translation vector T. The x-component of T is the horizontal distance between the optical centers of the two cameras—the baseline length, which is the core parameter for calculating depth.

[0179] Based on the calibration results, stereo correction is performed. By calculating a set of mapping relationships, the image planes of the left and right cameras are reprojected onto a completely coplanar and row-aligned plane. After correction, for any point in the left image, its corresponding point in the right image must lie in the same row (i.e., the same v-coordinate). This simplifies the two-dimensional search problem into a one-dimensional horizontal search, greatly improving the efficiency of subsequent matching.

[0180] S2.2: Stereo Matching of Feature Points

[0181] This step aims to find the corresponding point in the right image for each upper body feature point in the left image.

[0182] 1. Matching basis: Using the left and right views after S2.1 correction, there is only horizontal displacement (parallax) between them.

[0183] 2. Matching Strategy: For sparse features such as feature points, an efficient and robust strategy is adopted:

[0184] Template-based local matching: For a detected feature point p_left(u,v) in the left image, a small image patch (e.g., 16x16 pixels) is cropped with p_left(u,v) as the center and used as the template.

[0185] One-dimensional search: On the same row (same v coordinate) of the right image, slide the template horizontally within a preset disparity range [d_min, d_max].

[0186] Similarity metric: Calculate the similarity between the template and each candidate location in the right image. Preferably, we use zero-mean normalized cross-correlation as the metric function, which has good robustness to changes in illumination.

[0187] Subpixel interpolation: After finding the position with the highest similarity, its disparity d is an integer pixel value. To obtain subpixel-level accuracy, quadratic curve fitting (or a similar interpolation algorithm) is performed on the best matching point and its neighborhood points to estimate a more accurate subpixel disparity d_subpixel.

[0188] S2.3: Depth Calculation and 3D Coordinate Reconstruction

[0189] This is a direct calculation process for converting parallax into depth and 3D (Three-Dimensional) coordinates, including:

[0190] 1. Depth Calculation: Based on the triangulation principle of binocular vision, the depth value Z of feature point P is calculated using the following formula:

[0191] Z=(f*B) / d_subpixel

[0192] in:

[0193] Z: Depth value of the feature point relative to the camera (unit: millimeters); f: Focal length of the camera (in pixels), usually taken from fx in the intrinsic parameter matrix; B: Baseline length, i.e., the horizontal distance between the optical centers of the left and right cameras (unit: millimeters); d_subpixel: Calculated subpixel disparity (unit: pixels). Disparity d is inversely proportional to depth Z. The larger the disparity (the farther the distance between points in the left and right images), the closer the object is to the camera, and the smaller the depth Z; conversely, the smaller the disparity d.

[0194] 2. 3D Coordinate Reconstruction: After obtaining the depth Z, the 2D pixel coordinates can be back-projected into 3D space using the camera model to obtain the 3D coordinates (X, Y, Z) of the feature point in the left camera coordinate system:

[0195] X=(u-cx)*Z / fx

[0196] Y=(v-cy)*Z / fy

[0197] Where (u,v) are the pixel coordinates of the feature point in the left image, and (cx,cy) are the principal point coordinates (the coordinates of the intersection of the camera's optical axis and the image plane in the pixel coordinate system, which can be understood as the "center reference point" of the image. Physically, the camera's optical axis is the central axis of the lens, which is perpendicular to the image plane, and the intersection of the optical axis and the image plane is the principal point).

[0198] S2.4: Depth optimization for feature points

[0199] Since feature point detection itself may have an error of a few pixels, and the human body surface is not rich in texture, direct matching may introduce noise. The following optimization strategies can be introduced:

[0200] Left-right consistency check: After completing the matching from the left image to the right image, perform a matching from the right image to the left image. Only when the two matching results are consistent (loop check) is the matching point considered reliable; otherwise, it is considered a mismatch and discarded.

[0201] Spatiotemporal filtering:

[0202] Spatial consistency constraints: Utilizing the rigidity of the human skeleton and known geometric relationships, the depth of all feature points within the same frame is checked. For example, the depth values ​​of the left and right shoulders should be very close, and the depth value of the top of the head should be similar to that of the neck. Abnormal depth values ​​that significantly deviate from the overall distribution will be corrected or removed.

[0203] Temporal filtering: Kalman filtering or moving window averaging is used to smooth the depth values ​​of the same feature point in consecutive frames. This effectively suppresses jitter in single-frame depth estimation, outputting a stable and smooth depth sequence, providing reliable input for subsequent height calculation.

[0204] The vehicle cockpit control method provided in this application:

[0205] (1) By reusing existing DMS (Driver Monitoring System) and OMS (Occupant Monitoring System) cameras, there is no need to install additional dedicated sensor equipment, which reduces hardware costs and avoids the installation complexity and space occupation problems caused by adding new sensors. This solution makes full use of existing cockpit perception hardware resources and improves system integration.

[0206] (2) Using computer vision algorithms to extract key points of the human body (such as shoulder width, head height, etc.) and combining them with depth information to estimate height. Compared with traditional height measurement methods (such as laser ranging or ultrasonic measurement), this method has the advantage of non-contact measurement, which will not interfere with the normal activities of the driver or passengers and improves the user experience.

[0207] (3) The regression equation between human body key points and height established by statistical learning methods can achieve fast and accurate height estimation. Compared with simple linear regression based on a single feature (such as head height), this method comprehensively considers the correlation of multiple key points, thus improving the estimation accuracy and reliability.

[0208] (4) By linking the height estimation results with the seat adjustment system and the HUD (Head-Up Display) system, adaptive adjustment of seat position, angle, and HUD display height can be achieved. Compared with manual adjustment, this automatic adjustment not only improves the convenience of use, but also optimizes the ergonomic settings according to the user's height, thereby improving driving comfort and safety.

[0209] (5) Combined with the Face ID function, the system can remember the height information and personalized settings of different users, realize quick switching and personalized adaptation in multi-user scenarios, and improve the intelligence level of the cockpit system and user experience.

[0210] (6) It supports multiple depth information acquisition methods (monocular estimation, binocular estimation, TOF, etc.), has good compatibility and adaptability, and can flexibly select the implementation scheme according to the hardware configuration of different vehicle models, thus improving the scalability and application scope of the technology.

[0211] It should be understood that although the steps in the flowcharts of the above embodiments are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the above embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages in other steps. It is understood that the steps in different embodiments can be freely combined as needed, and all non-contradictory solutions formed by such combinations are within the scope of protection of this application.

[0212] Based on the same inventive concept, this application also provides a vehicle cockpit control device for implementing the above-described vehicle cockpit control method. The solution provided by this device is similar to the implementation described in the above method; therefore, the specific limitations of one or more vehicle cockpit control device embodiments provided below can be found in the limitations of the vehicle cockpit control method described above, and will not be repeated here.

[0213] In one exemplary embodiment, such as Figure 8 As shown, a vehicle cockpit control device 800 is provided, including: a planar position determination module 801, a spatial position determination module 802, a size information determination module 803, and an equipment control module 804, wherein:

[0214] The planar position determination module 801 is used to acquire images of target objects in the cockpit and detect the planar positions of multiple preset feature points of the target objects from the images.

[0215] The spatial location determination module 802 is used to determine the depth information of each feature point, and for each feature point, obtain the spatial location of the feature point based on the planar location and depth information of the feature point;

[0216] The size information determination module 803 is used to determine the size information of the target object based on the spatial position of each feature point;

[0217] Equipment control module 804 is used to control equipment in the cockpit based on dimensional information.

[0218] In some embodiments, the images include a first image and a second image acquired by a binocular camera installed in the cockpit; the spatial position determination module 802 is further configured to obtain a positional mapping relationship between the first image and the second image, the positional mapping relationship being pre-calibrated for the binocular camera; for each feature point, based on the positional mapping relationship, determine the disparity of the feature point between the first image and the second image; and obtain the depth information of the feature point based on the planar position and disparity of the feature point.

[0219] In some embodiments, the spatial position determination module 802 is further configured to obtain a first planar position of the feature point in the first image and a second planar position of the feature point in the second image; based on the position mapping relationship, map the first planar position to the second image to obtain a first disparity of the feature point; based on the position mapping relationship, map the second planar position to the first image to obtain a second disparity of the feature point; and based on the first disparity and the second disparity, obtain the disparity of the feature point between the first image and the second image.

[0220] In some embodiments, the spatial location determination module 802 is further configured to obtain the original depth information of the feature point based on the planar position and disparity of the feature point; obtain the historical depth information of the feature point, wherein the historical depth information of the feature point is obtained based on historical images acquired for the target object; and obtain the depth information of the feature point based on the historical depth information and the original depth information of the feature point.

[0221] In some embodiments, the spatial location determination module 802 is further configured to obtain the original depth information of the feature points based on the planar position and disparity of the feature points; and determine the depth information of the feature points based on the feature point attribute relationship of the target object and the original depth information.

[0222] In some embodiments, the planar position determination module 801 is further configured to identify the image region where the target object is located in the image through a pre-trained target detection network; and to detect the planar position of multiple preset feature points of the target object from the image region through a pre-trained regression network.

[0223] In some embodiments, the planar position determination module 801 is further configured to detect the original planar positions of multiple preset feature points of the target object from the image; determine the historical planar position of each feature point, wherein the historical planar position of each feature point is obtained based on the detection of historical images acquired for the target object; and for each feature point, obtain the planar position of the feature point based on the historical planar position and the original planar position of the feature point.

[0224] In some embodiments, the size information determination module 803 is further configured to determine the upper body features of the target object based on the spatial location of each feature point, wherein the upper body features include at least one of head and neck height, shoulder width, torso height, or arm length; and to obtain the size information of the target object by predicting based on the upper body features using the size prediction model corresponding to the target object.

[0225] In some embodiments, the device control module 804 is further configured to: determine a first control parameter for a seat in the cockpit based on size information, and adjust the seat according to the first control parameter; determine a second control parameter for a display device in the cockpit based on size information, and adjust the display device according to the second control parameter.

[0226] The various modules in the cockpit control device of the aforementioned vehicle can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in the processor of the controller in hardware form or independent of it, or stored in the memory of the controller in software form, so that the processor can call and execute the corresponding operations of each module.

[0227] In one exemplary embodiment, a controller is provided, the internal structure of which can be shown in the following diagram. Figure 9 As shown, the controller includes a processor, memory, input / output (I / O) interfaces, and a communication interface. The processor, memory, and I / O interfaces are connected via a system bus, and the communication interface is also connected to the system bus via the I / O interfaces. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides the environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The I / O interfaces are used for exchanging information between the processor and external devices. The communication interface is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, mobile cellular networks, Near Field Communication (NFC), or other technologies. When the computer program is executed by the processor, it implements a vehicle cockpit control method.

[0228] Those skilled in the art will understand that Figure 9 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the controller to which the present application is applied. A specific controller may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0229] In one embodiment, a controller is also provided, including a memory and a processor, the memory storing a computer program, the processor executing the computer program to implement the steps in the above method embodiments.

[0230] In one embodiment, a vehicle is also provided, including a cockpit and a controller, the controller including a memory and a processor, the memory storing a computer program, the processor executing the computer program to implement the steps in the above method embodiments.

[0231] In one embodiment, a computer-readable storage medium is provided storing a computer program that, when executed by a processor, implements the steps in the above method embodiments.

[0232] In one embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps in the above method embodiments.

[0233] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data must comply with relevant regulations.

[0234] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, artificial intelligence (AI) processors, etc., and are not limited to these.

[0235] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.

[0236] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.

Claims

1. A vehicle cockpit control method, characterized in that, The method includes: Acquire images of a target object in the cockpit, and detect the planar positions of multiple preset feature points of the target object from the images; Determine the depth information of each feature point, and for each feature point, obtain the spatial position of the feature point based on its planar position and depth information; The size information of the target object is determined based on the spatial location of each feature point; The equipment in the cockpit is controlled based on the size information.

2. The method according to claim 1, characterized in that, The images include a first image and a second image captured by the binocular cameras installed in the cockpit; determining the depth information of each feature point includes: Obtain the positional mapping relationship between the first image and the second image, wherein the positional mapping relationship is pre-calibrated for the binocular camera; For each feature point, based on the position mapping relationship, the disparity of the feature point between the first image and the second image is determined; Furthermore, the depth information of the feature point is obtained based on its planar position and disparity.

3. The method according to claim 2, characterized in that, The step of determining the disparity of the feature points between the first image and the second image based on the position mapping relationship includes: Obtain the first planar position of the feature point in the first image, and the second planar position of the feature point in the second image; Based on the position mapping relationship, the first planar position is mapped to the second image to obtain the first disparity of the feature point; Based on the position mapping relationship, the second planar position is mapped to the first image to obtain the second disparity of the feature point; Based on the first disparity and the second disparity, the disparity of the feature point between the first image and the second image is obtained.

4. The method according to claim 2, characterized in that, The step of obtaining the depth information of the feature point based on its planar position and disparity includes: Based on the planar position and disparity of the feature point, the original depth information of the feature point is obtained; The historical depth information of the feature points is obtained, and the historical depth information of the feature points is obtained based on historical images collected for the target object. The depth information of the feature point is obtained based on the historical depth information and the original depth information of the feature point.

5. The method according to claim 2, characterized in that, The step of obtaining the depth information of the feature point based on its planar position and disparity includes: Based on the planar position and disparity of the feature point, the original depth information of the feature point is obtained; The depth information of the feature points is determined based on the feature point attribute relationship of the target object and the original depth information.

6. The method according to any one of claims 1 to 5, characterized in that, Determining the size information of the target object based on the spatial position of each feature point includes: Based on the spatial location of each feature point, the upper body features of the target object are determined, and the upper body features include at least one of head and neck height, shoulder width, torso height, or arm length. The size information of the target object is obtained by using the size prediction model corresponding to the target object and predicting based on the upper body features.

7. The method according to any one of claims 1 to 5, characterized in that, The device that controls the cockpit based on the size information includes at least one of the following: Based on the size information, a first control parameter is determined for the seat in the cabin, and the seat is adjusted according to the first control parameter; Based on the size information, a second control parameter is determined for the display device in the cockpit, and the display device is adjusted according to the second control parameter.

8. A vehicle cockpit control device, characterized in that, The device includes: The planar position determination module is used to acquire images of a target object in the cockpit and detect the planar positions of multiple preset feature points of the target object from the images. The spatial location determination module is used to determine the depth information of each feature point, and for each feature point, obtain the spatial location of the feature point based on the planar location and depth information of the feature point; The size information determination module is used to determine the size information of the target object based on the spatial position of each feature point; The equipment control module is used to control the equipment in the cockpit based on the size information.

9. A controller comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 7.

10. A vehicle, characterized in that, Includes the cockpit and the controller as described in claim 9.