Method for identifying physical features of driver and passenger and automatically adjusting seat and safety air bag
Through VitPose-based attitude recognition algorithm and deep learning technology, accurate prediction of driver's body shape is achieved, and through the automatic adjustment of seats and airbags, the problem that the existing system cannot dynamically adapt to driver's body shape changes is solved, significantly improving driving experience and safety.
Patent Information
- Application Number
- CN202510114453.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-24
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-01-24
AI Technical Summary
The existing intelligent seat adjustment system cannot dynamically adapt to the driver's body shape changes and attitude adjustment, resulting in insufficient seat adjustment accuracy and single function, and failing to consider the impact of the overall driving environment on driver comfort and safety.
VitPose-based attitude recognition algorithm is adopted, combined with deep learning algorithms to accurately predict the driver's height and body shape, and automatic adjustment of seats and airbags is achieved through steps such as image capture and data set construction, key bone point annotation, two-dimensional bone point recognition system construction, and three-dimensional reconstruction based on two-dimensional bone point.
It realizes personalized and accurate comfort and safety experience for the driver, avoids the cumbersomeness of traditional manual operations, significantly improves the user's driving experience and safety, and optimizes the settings of seats and airbags through dynamic adjustment strategies and real-time monitoring mechanisms, reducing fatigue and secondary injury risks.
Smart Images

Figure CN120071309A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of intelligent driving, covering multiple disciplines such as computer vision, machine learning, human pose recognition, and automotive human-computer interaction. Specifically, it relates to a method for recognizing the physical characteristics of vehicle occupants and automatically adjusting seats and airbags. Background Art
[0002] Today, with the rapid development of intelligent driving technology, the personalized adaptation needs of drivers have become an important research direction for improving the driving experience and driving safety. Traditional vehicle facility adjustment methods (such as seat position, airbag configuration, etc.) mostly rely on manual adjustment by the driver. This method is not only cumbersome and time-consuming but also prone to discomfort or safety hazards due to improper adjustment.
[0003] To simplify this process, some existing technologies have proposed intelligent seat adjustment solutions. For example, the patent of CN110316027 stores the information of the seat and rearview mirror positions set manually by the user for the first time in the vehicle system for automatic restoration when used next time. However, this method is still limited to fixed memory adjustment after the initial setting and fails to make adaptive adjustments according to the changes in the driver's body shape or posture. In addition, CN107232822 proposes an intelligent seat adjustment system based on pressure sensors and posture sensors, which adjusts the seat position by detecting the pressure and posture changes borne by the seat. Although this technology provides a posture recognition function, the sensors it relies on have limited accuracy and relatively single functions, and cannot accurately predict the driver's height and body shape.
[0004] Based on the above analysis, the deficiencies of the existing technologies can be summarized as follows: (1) Lack of dynamic body shape adaptability: Most existing systems are based on fixed initial settings or single sensor information and cannot make real-time adjustments according to the changes in the driver's body shape and posture; (2) Limited adjustment accuracy: Although some systems use sensors to detect seat pressure or posture, they cannot accurately predict the driver's body shape, resulting in insufficient accuracy of seat adjustment; (3) Single function: Existing technologies mainly focus on the basic position adjustment of the seat and do not consider the impact of other key factors (such as bone point postures) in the overall driving environment on the driver's comfort and safety. In this context, there is an urgent need for a more intelligent in-vehicle seat adjustment system that can accurately predict the driver's height and body shape through advanced posture recognition algorithms and machine learning models and automatically adjust the seat position to meet the personalized needs of different drivers and improve the driving experience and driving safety; (4) Lack of convenience and security: It is not deployed locally, so it cannot perform instant and non-intrusive recognition of the driver, resulting in insufficient convenience, and the information is not stored locally, so it cannot guarantee the driver's personal privacy information. Summary of the Invention
[0005] In the context of the continuous development of intelligent driving and in-vehicle systems, the intelligent adjustment technology based on physical characteristics has become an important innovative direction for enhancing the driving experience and driving safety of drivers. Traditional adjustment methods mostly rely on drivers to manually adjust facilities such as seats and rearview mirrors to adapt to personal height, body shape, and comfort requirements. This method is not only cumbersome and time-consuming but also easily leads to a poor driving environment due to improper adjustment, affecting the field of vision and the accuracy of operations, thereby reducing driving safety. At the same time, some existing intelligent adjustment technologies mainly focus on single functions, such as the memory or pressure-sensing adjustment of seat positions, and fail to achieve the comprehensive adaptation of the overall driving environment and lack the ability of dynamic real-time adjustment.
[0006] Through the pose recognition algorithm based on VitPose, this invention captures the key bone points in multiple frames of images and combines deep learning algorithms to accurately predict the height and body shape of the driver. It not only solves the problem of lacking consideration for the body shape differences of drivers in the existing technology but also enables the automatic adjustment of seats and airbags, allowing each driver to enjoy personalized and precise comfort and safety experiences. This intelligent adjustment method avoids the cumbersome traditional manual operations and significantly improves the driving experience and safety of users.
[0007] To achieve the above invention purpose, this invention proposes a method for recognizing the physical characteristics of passengers and automatically adjusting seats and airbags, including the following steps:
[0008] S1 Image capture and dataset construction: Continuously capture image sequences of the driver getting into the vehicle through an in-vehicle camera, analyze the continuous image sequences using an action recognition algorithm based on 2D CNN + LSTM to obtain the action characteristics of the driver and form a key frame dataset, screen the key frame dataset, and divide the screened key frame dataset into a training set and a test set according to a certain proportion;
[0009] S2 Key bone point annotation: Use an automatic annotation tool to annotate the bone points of each key frame image I i to generate a set of annotated bone point coordinates P label ={(x label , y label )}, where x label represents the horizontal coordinate of a certain bone point on the image plane, and y label represents the vertical coordinate of a certain key bone point on the image plane. This annotated data is used for the training of the subsequent two-dimensional bone point recognition system;
[0010] S3 Construction of a two-dimensional bone point recognition system: Extract features from the key frame image I i to generate a preliminary set of bone point coordinates P i and the corresponding bone point connection diagram S i; Combine with the in-vehicle physical layout, and generate an optimization result set through multi-modal fusion and self-attention mechanism and Analyze the dynamic changes of the skeletal points in multiple frames of images, and use spatio-temporal convolution to generate dynamic optimization results and Output a set of two-dimensional skeletal point coordinates and the corresponding skeletal point connection diagram S final ;
[0011] S4 Obtain in-vehicle space data based on the three-dimensional reconstruction of two-dimensional skeletal points, and use the modeling tool Blender to model the in-vehicle three-dimensional space; utilize the set of two-dimensional skeletal point coordinates Combine with the JOTR algorithm to calculate the three-dimensional skeletal point coordinates Use the self-supervised learning framework of the generative adversarial network for Calibrate to generate the calibrated three-dimensional coordinates Calculate the Euclidean distance d between the three-dimensional skeletal points based on the calibration results ij , and map it to the driver's height and body shape parameters ShapeParams through a multi-layer perceptron MLP, and output the driver's accurate height and body shape parameters ShapeParams;
[0012] S5 Seat automatic adjustment According to the driver's height Body shape parameters ShapeParams, analyze the positions and sitting postures of the driver's head, shoulders and waist, and dynamically adjust the height, backrest angle and seat cushion angle of the seat to match the driver's body shape requirements; adopt a dynamic adjustment strategy, real-time monitor the driver's posture changes during long-term driving, and respond intelligently to optimize the seat settings and reduce fatigue.
[0013] S6 Airbag automatic adjustment By analyzing the driver's three-dimensional skeletal point coordinates Body shape parameters ShapeParams and real-time sitting posture information, dynamically optimize the deployment position, deployment range and triggering strategy of the airbag; adopt a real-time monitoring mechanism to ensure that the airbag can accurately protect the driver's key parts, such as the head, chest and shoulders, when deployed, and reduce the risk of secondary injury during a collision.
[0014] Furthermore, the image capture and dataset construction in the above step S1 specifically include:
[0015] S1.1 The 2D convolutional neural network extracts spatial features from each frame of the continuous image sequence, such as the human body contour, key limb parts and their movement trajectories;
[0016] S1.2 Process the temporal information in the continuous image sequence using the LSTM model, capture the continuous action features of the driver during the boarding process, analyze the transitional features of actions in a set of adjacent frames to identify the turning points of actions.
[0017] S1.3 Analyze the time curve of the change in action amplitude, mark the key frames at the moment of sharp amplitude change, and capture the starting frame and ending frame of each action.
[0018] S1.4 Develop an adaptive screening strategy to select the key frame images I i , and form the key frame dataset {I i}, and divide the screened key frame dataset into a training set and a test set according to a certain proportion to ensure the generalization ability and accuracy of the model.
[0019] Furthermore, the above adaptive screening strategy includes the following:
[0020] S1.4.1) Quantification of action fluency: Quantify the action fluency by calculating the optical flow between two adjacent frames of images.
[0021] Let the image frames I t and I t+1 be two consecutive frames of images respectively. The calculation method of the optical flow F t is as follows:
[0022]
[0023] where T is the time interval, and F t is the displacement vector between two frames, indicating the movement of the object between frames; if the displacement exceeds the threshold, it is considered that a large change has occurred in the action, which does not meet the fluency requirement.
[0024] Calculate the mean square error MSE of the optical flow between adjacent frames. Let F t and F t+1 be the optical flows of adjacent frames, and calculate their mean square error:
[0025]
[0026] Select the index. If MSE(F t , F t+1 ) is higher than the threshold 0.15, it is considered that the action is not fluent, and this frame is excluded.
[0027] S1.4.2) Quantification of the visibility of limb parts: Use OpenPose to perform fast skeleton detection on the image to obtain the preliminary visibility score of key parts; the visibility V(p i)Quantify according to the detection confidence score and image quality, with the scoring range being [0, 1]; use the spatial information of the fixed structures inside the vehicle, such as the seats and the steering wheel, to determine whether a certain part is occluded; for the occluded area O of the vehicle interior structure and the total visible area V(p i ), quantify it using the following occlusion index:
[0028]
[0029] Among them, |V(p i )) ∩ O| is the area of the occluded part, and |V(p i ))| is the total visible area of the part. If the occlusion ratio is greater than the set threshold of 40%, then this part is considered invisible;
[0030] S1.4.3) Occlusion situation assessment: Set the percentage of the occluded area as the evaluation index to determine whether the fixed structures inside the vehicle (such as seats and steering wheels) cause significant occlusion to the key parts of the driver (such as upper limbs and legs); process the image through the image segmentation technology Mask R-CNN to extract the vehicle interior structure area R car and the driver area R driver , and calculate the intersection and overlapping area of the two areas; let |R car ∩R driver | be the intersection area of the vehicle interior structure and the driver area, and |R driver | be the total area of the driver area. Then the occlusion degree is:
[0031]
[0032] If the occlusion ratio exceeds 40%, then this frame is considered to have a significant occlusion problem.
[0033] Furthermore, the specific method for data annotation and bone point selection in the above step S2 is as follows:
[0034] S2.1 Use the CVAT automatic annotation tool to perform bone point annotation on each selected key frame image I i . According to the special scenes inside the vehicle and the continuous actions of the driver from opening the door to holding the steering wheel, design a dynamic annotation strategy. The annotation scheme is as follows:
[0035] Opening the door stage: Focus on annotating bone points such as shoulders, elbows, wrists, knees, and ankles to capture the movement trajectory of the driver reaching out to pull the door and ensure that the core posture features of this action are recorded. This stage is the only stage that shows the complete knee-to-ankle stage of the driver;
[0036] Sitting down stage: Focus on annotating key bone points such as the spine and knees to obtain the sitting posture angle and body rotation state of the driver.
[0037] Seat belt stage: Focus on capturing the subtle changes in the positions of upper body bone points such as the shoulders, elbows, and waist, accurately record the driver's side body movements, and provide data support for subsequent analysis;
[0038] Steering wheel holding stage: Mark key bone points such as the shoulders, elbows, wrists, and neck to ensure accurate recording of the final driving sitting posture;
[0039] The key bone points marked in S2.2 include the following parts: (1) Nose; (2) Neck; (3) Left shoulder; (4) Right shoulder; (5) Chest; (6) Left elbow; (7) Right elbow; (8) Left hip; (9) Right hip; (10) Left knee; (11) Right knee; (12) Left ankle; (13) Right ankle.
[0040] Furthermore, the above step S3 includes:
[0041] In the feature extraction stage of S3.1, the VitPose network is used to extract features from each key frame image I i to generate a preliminary feature representation related to bone points;
[0042] Input the key frame image I i , assuming it is an image of size H×W×C, where H and W are the height and width of the image respectively, and C is the number of channels; divide the image into image patches of size P×P, and the number of pixels in each image patch is P 2 ·C. Assuming there are a total of N image patches, the calculation formula for N is:
[0043]
[0044] Each image patch is embedded into a d-dimensional feature vector space through a linear transformation to obtain the embedding vector e ij ∈R d , and the specific formula is as follows:
[0045] e ij =W patch P ij +b patch
[0046] where, is the weight of the linear transformation, b patch is the bias term. Stack all the embedding vectors e ij into a matrix E∈R N×d , and input it into the Transformer encoder. The Transformer encoder captures the global context information of the image through the self-attention mechanism, and its core calculation formula is:
[0047]
[0048] Among them, Q, K, and V are matrices of query, key, and value respectively, and d k is the scaling factor of the feature dimension. After being processed by multiple Transformer layers, the output feature matrix is:
[0049] z i = Transformer(E), z i ∈R N×d
[0050] The output feature matrix z i contains the global context information of the image;
[0051] S3.2 Pose Estimation Layer (PoseEstimationLayer, SE_Layer) further predicts the preliminary skeleton point coordinates and connection diagram of the driver based on the feature matrix z i ;
[0052] The feature matrix z i ∈R N×d generated by the Transformer encoder and the in-vehicle physical layout S car are input into SE_Layer. Through multimodal fusion, S car is fused with z i and input into the regression network g pose to predict the preliminary skeleton point coordinates:
[0053] P i = g pose (concat(z i , S car ))), P i = {(x i , y i )}
[0054] Among them, P i = {(x i , y i )} represents the set of preliminary skeleton point coordinates, including the two-dimensional positions of each skeleton point. g pose is the regression function used to extract the skeleton point coordinates from the feature matrix. concat means splicing the two kinds of information together. According to the set of skeleton point coordinates P i , the skeleton point connection diagram S i is constructed, which is generated through the topological relationship between the skeleton points and describes the skeleton point features of the driver:
[0055] S i = (V, E), V = P i , E = {(v k , v j )∣vk , v j ∈ V and satisfies the connection relationship}
[0056] The output result includes the initial set of skeletal point coordinates P i and the corresponding skeletal point connection diagram S i ;
[0057] S3.3 Scene-aware Feature Integration Layer (SF_Layer) further optimizes the features of skeletal points through the in-vehicle environment perception mechanism;
[0058] Input the initial set of skeletal point coordinates P i , the initial skeletal point connection diagram S i and the in-vehicle physical layout S car , perform linear transformations on the skeletal point connection diagram S i and the in-vehicle physical layout S car respectively, and map them to the same high-dimensional feature space:
[0059]
[0060] Among them, and are the weight matrices of the linear transformation, and are the bias terms, and the self-attention mechanism is used to calculate the correlation between the skeletal point features and the in-vehicle physical layout:
[0061]
[0062] Among them, d represents the feature dimension, S′ i (S′ car ) T reflects the degree of association between each skeletal point and the in-vehicle structure, softmax is the normalized attention weight, Attn i is the weighted influence result of the in-vehicle physical layout on the skeletal point features, and the corresponding skeletal point connection diagram is constructed according to the optimized set of skeletal point coordinates
[0063]
[0064] Output the optimized set of skeletal point coordinates and the corresponding skeletal point connection diagram
[0065] S3.4 Spatio-Temporal Layer (ST_Layer) optimizes the temporal consistency and spatial coherence of skeletal points by analyzing the changes in skeletal points between consecutive multi-frame images;
[0066] Input the optimized skeletal point coordinate sets of consecutive multi-frames and the corresponding skeletal point connection diagrams Integrate them into Stack the multi-frame feature data into spatio-temporal feature blocks with dimensions of T×N×d, where T represents the number of frames in the time dimension, and 3D convolution is used to capture temporal and spatial features to optimize the dynamic coherence of skeletal points:
[0067]
[0068] where, W ST ∈R d′×d×k×k×T is a 3D convolution kernel, d′ is the output feature dimension, k is the spatial dimension of the convolution kernel, T is the range of the time dimension, and z ST ∈R N×d′ is a high-dimensional feature matrix containing spatio-temporal context information, which not only considers the position changes of skeletal points between multi-frame images but also captures the temporal coherence of the driver's actions; after spatio-temporal convolution processing, the skeletal point coordinates are dynamically optimized using dynamic weights, and the formula is as follows:
[0069]
[0070] where, is the skeletal point coordinate of the t-th frame, and α t is the dynamic weight, which is calculated from the spatio-temporal feature correlation of each frame of skeletal points:
[0071]
[0072] where, MLP is a multi-layer perceptron, which extracts the correlation score f ST between the skeletal point features of each frame and the overall dynamic optimization objective from the spatio-temporal feature z t for subsequent weight calculation; based on the dynamically optimized skeletal points construct the corresponding skeletal point connection diagrams
[0073]
[0074] The topological relationship here is the same as but is recalculated based on the positions of the dynamically optimized skeletal points.
[0075] Output the dynamically optimized skeletal point coordinate sets and the corresponding skeletal point connection diagram It retains the consistency in time and space;
[0076] S3.5 Proportion Adjustment Layer (PA_Layer) performs human body proportion correction on the dynamically optimized skeletal point set and the corresponding skeletal point connection diagram to ensure that it meets the human anatomy standards.
[0077] Input the dynamically optimized skeletal point coordinate set and the corresponding skeletal point connection diagram Predict the physical size of the skeletal points from the features through a regression network:
[0078] h i = MLP(z ST )
[0079] According to the standard human body proportion h max , calculate the scale factor λ of each skeletal point i :
[0080]
[0081] λ i is the scaling factor representing the physical size of each skeletal point relative to the standard human body proportion, used to ensure that the skeletal point prediction results can adapt to drivers of different body types. Therefore, the scale factor λ of each skeletal point i is calculated independently. Adjust the skeletal point coordinates using the scale factor:
[0082]
[0083] According to the corrected skeletal point coordinates Construct the corresponding skeletal point connection diagram S final .
[0084] Output the set of skeletal point coordinates after proportion correction and the corresponding skeletal point connection diagram S final , as the final two-dimensional skeletal point recognition result.
[0085] Furthermore, the above step S4 for the 3D reconstruction based on 2D skeletal points specifically includes:
[0086] According to the 2D skeletal point coordinate set Combined with the internal structure information of the vehicle and the JOTR algorithm to complete the 3D skeletal point reconstruction, and finally generate a 3D skeletal point connection diagram and extract the body type parameters of the driver.
[0087] S4.1 Input the interior space data (such as seat height, width, depth, and interior length and width) for modeling, and use the modeling tool Blender to build the constraint structure of the three-dimensional space in the vehicle;
[0088] S4.2 Using JOTR algorithm combined with 2D skeleton point coordinate set Calculate the 3D coordinates of each bone point Provide input for subsequent 3D reconstruction:
[0089]
[0090] Initial 3D bone points There is a certain error, so the self-supervised learning framework of the Generative Adversarial Network (GAN) is used to correct it and convert the three-dimensional coordinates Projected onto a two-dimensional plane, we get the projection
[0091]
[0092] The projection function is calculated based on camera parameters and perspective relationship.
[0093] S4.3 Discriminator D projects the points onto the two-dimensional plane and the original 2D point The error feedback generator G between them minimizes the loss function L self , continuously optimize the generator so that the three-dimensional coordinates of its output are more consistent with the two-dimensional projection of the real image points:
[0094]
[0095] After the correction is completed, the generator G outputs the corrected 3D bone points:
[0096]
[0097] S4.4 Based on the corrected 3D bone point set Generate 3D bone point connection diagram Its topological relationship is similar to the two-dimensional connection graph S final Same as, but recalculated in 3D space, computing the Euclidean distance d between 3D bone points ij , used to characterize the driver's limb length characteristics:
[0098]
[0099] Among them, d ij represents the three-dimensional distance between the i-th bone point and the j-th bone point, and the distance feature set {d ij} Input to the multi-layer perceptron (MLP) and mapped to the driver's height and body shape parameters ShapeParams:
[0100]
[0101] The final output is a more accurate driver height and their body shape parameters ShapeParams.
[0102] Furthermore, the dynamic adjustment strategy for step S5 above is as follows:
[0103] S5.1 Seat front-back adjustment: Ensure the appropriate distance between the driver and the pedals and steering wheel to achieve the best control and comfort;
[0104] S5.2 Seat height adjustment: Ensure the driver has the best field of vision and reduce interference from the roof or other components;
[0105] S5.3 Seat backrest angle adjustment: Ensure sufficient support for the back and reduce driver fatigue;
[0106] S5.4 Seat and steering wheel linkage adjustment: Ensure that the driver can comfortably operate the steering wheel and avoid unnatural arm extension;
[0107] S5.5 Seat belt adjustment: Ensure that the seat belt can accurately restrain the driver's body and provide the best safety protection.
[0108] Furthermore, the above step S6 includes:
[0109] S6.1 Using three-dimensional skeletal point coordinates Extract the set P of key part coordinates of the driver key :
[0110]
[0111] Among them, respectively represent the calibrated three-dimensional coordinates of the head, chest and shoulders, represents the calibrated three-dimensional coordinate of the key part k;
[0112] S6.2 According to the body shape parameters ShapeParams, calculate the range of the driver's body shape characteristics:
[0113]
[0114] Among them, BodyWidth represents the horizontal distance between the left and right endpoints of the driver's shoulders, and are respectively the x-axis values of the calibrated three-dimensional coordinates of the right shoulder and the left shoulder; BodyHeight represents the vertical distance from the driver's head to the hip, and respectively represent the z-axis values of the calibrated three-dimensional coordinates of the head and the hip;
[0115] S6.3 Based on the coordinates P of the driver's key parts key , calculate the optimal center position P deploy and the coverage radius R deploy :
[0116]
[0117] wherein, P deploy is the center point of airbag deployment, calculated based on the average coordinates of the driver's key parts, and R deploy is the coverage radius of the airbag, determined by the maximum three-dimensional distance of the key parts;
[0118] S6.4 Combine the data of in-vehicle collision sensors to dynamically adjust the triggering conditions of the airbag; the calculation formula of the triggering threshold T trigger is:
[0119] T trigger = w 1 F collision + w 2 θ impact + w 3 Δ pose
[0120] wherein, F collision is the collision force, θ impact is the collision angle, Δ pose is the driver's sitting posture deviation, and w 1 , w 2 is a weight parameter used to adjust the influencing factors of the triggering conditions;
[0121] The deployment force P of the airbag inflate is calculated according to the driver's body type parameters and the severity of the collision:
[0122] P inflate = f(ShapeParams, Collision Severity)
[0123] S6.5 Dynamically adjust the airbag deployment direction and deployment strategy by real-time monitoring of the driver's sitting posture changes;
[0124] The adjusted airbag deployment angle α adjusted = α default + kΔ pose is expressed as:
[0125] α adjusted = α default + k·Δ pose
[0126] Among them, α default is the default deployment angle, k is the sitting posture correction coefficient, and ·Δ pose is the detected sitting posture deviation amount; the deployment center point P deploy of the airbag, the coverage range R deploy the trigger threshold T trigger and the deployment force P inflate and the deployment angle α adjusted after the driver's dynamic adjustment are obtained to provide the adjustment parameters of the airbag, ensuring that the key parts of the driver can be accurately protected when the airbag deploys and reducing the risk of secondary injury during a collision.
[0127] Compared with the prior art, the present invention has the following advantages:
[0128] 1. Efficient action capture and data annotation process: Multiple frames of images are captured through dual-view cameras and an adaptive screening strategy, and the CVAT semi-automatic annotation tool is used to improve the data annotation efficiency, ensuring that the captured body shape feature data is more comprehensive and accurate, thereby enhancing the prediction performance of the model.
[0129] 2. Improved accuracy of pose recognition in the vehicle interior occlusion environment: The improved VitPose pose recognition algorithm combines the occlusion perception mechanism of special scenes in the vehicle interior, and can accurately identify the bone points occluded by the seat, steering wheel, etc., realizing more accurate two-dimensional bone point recognition.
[0130] 3. Enhanced accuracy of 3D reconstruction: The JOTR algorithm is used to achieve 3D reconstruction based on two-dimensional bone points, combined with the vehicle interior space data, improving the accuracy of driver body shape and height prediction, and overcoming the influence of vehicle interior space limitations and occlusion on body shape measurement.
[0131] 4. Automated driving position adjustment: The system automatically adjusts the seat and seat belt positions according to the driver's body shape data, enabling the driver to obtain the best safe driving posture, improving driving comfort and safety, and solving the inconvenience and insufficient adaptability of traditional manual adjustment.
[0132] 5. Precise deployment of the airbag: Through the analysis of three-dimensional bone points and body shape parameters, the deployment position, coverage range and trigger strategy of the airbag are dynamically adjusted to ensure accurate protection of the driver's head, chest, shoulders and other key parts during a collision, reducing the risk of secondary injury and significantly improving the passive safety performance of the vehicle.
[0133] 6. Privacy protection and local data processing: Support in-vehicle deployment and local data storage, without relying on cloud processing, ensuring the user's autonomous control of personal data and effectively protecting privacy and security. Description of the Drawings
[0134] Figure 1 Schematic diagram of the process for constructing the two-dimensional skeleton point recognition system ICDPose of the present invention;
[0135] Figure 2 Schematic diagram of the final effect structure of the present invention. Specific embodiments
[0136] The technical solution of the present invention will be described in detail below, but the protection scope of the present invention is not limited to the described embodiments.
[0137] As Figure 1 、 Figure 2 shown, a method for recognizing the physical characteristics of vehicle occupants and automatically adjusting the seat and airbag of the present invention includes the following steps:
[0138] S1 Image capture and dataset construction: Continuously capture image sequences of the driver getting into the vehicle through an in-vehicle camera, analyze the continuous image sequences using an action recognition algorithm based on 2D CNN + LSTM, obtain the action characteristics of the driver to form a key frame dataset, screen the key frame dataset, and divide the screened key frame dataset into a training set and a test set according to a ratio;
[0139] S2 Key skeleton point annotation: Use an automatic annotation tool to annotate the skeleton points of each key frame image I i to generate a set of annotated skeleton point coordinates P label ={(x label , y label )}, where x label represents the horizontal coordinate of a certain skeleton point on the image plane, and y label represents the vertical coordinate of a certain key skeleton point on the image plane. This annotation data is used for the subsequent training of the two-dimensional skeleton point recognition system;
[0140] S3 Construction of the two-dimensional skeleton point recognition system: Extract features from the key frame image I i to generate a preliminary set of skeleton point coordinates P i and the corresponding skeleton point connection diagram S i ; Combine the in-vehicle physical layout, such as the spatial positions of fixed structures such as seats, steering wheels, and rearview mirrors, and generate an optimized result set through multi-modal fusion and self-attention mechanism and Analyze the dynamic changes of the skeleton points in multiple frames of images, use spatio-temporal convolution to generate a dynamically optimized result and Through human body proportion correction, output the two-dimensional skeleton point coordinate set and the corresponding skeleton point connection diagram S final ;
[0141] S4 Obtain in-vehicle space data based on 3D reconstruction of 2D skeletal points, such as data on vehicle seat height, width, depth, length and width of the interior space, etc., and use the modeling tool Blender to model the 3D in-vehicle space; utilize the set of 2D skeletal point coordinates Combine with the JOTR algorithm to deduce 3D skeletal point coordinates Use the self-supervised learning framework of the generative adversarial network for Perform correction to generate corrected 3D coordinates Calculate the Euclidean distance d between 3D skeletal points based on the correction result ij and map it to the driver's height through the multi-layer perceptron MLP and body shape parameters ShapeParams, and output the driver's accurate height and body shape parameters ShapeParams;
[0142] S5 Seat automatic adjustment According to the driver's height and body shape parameters ShapeParams, analyze the positions and sitting postures of the driver's head, shoulders and waist, and dynamically adjust the height, backrest angle and seat cushion angle of the seat to match the driver's body shape requirements; adopt a dynamic adjustment strategy, real-time monitor the driver's posture changes during long-term driving, and respond intelligently to optimize the seat settings and reduce fatigue.
[0143] S6 Airbag automatic adjustment By analyzing the driver's 3D skeletal point coordinates body shape parameters ShapeParams and real-time sitting posture information, dynamically optimize the deployment position, deployment range and triggering strategy of the airbag; adopt a real-time monitoring mechanism to ensure that the airbag can accurately protect the driver's key parts, such as the head, chest and shoulders, when deployed, and reduce the risk of secondary injury during a collision.
[0144] As a preferred embodiment of the present invention, the image capture and dataset construction in S1 specifically include:
[0145] S1.1 The 2D convolutional neural network extracts spatial features from each frame of the continuous image sequence, such as the human body contour, key limb parts and their movement trajectories;
[0146] S1.2 Use the LSTM model to process the temporal information in the continuous image sequence, capture the continuous action features of the driver during the boarding process, analyze the action transitional features in a group of adjacent frames to identify the turning points of the actions;
[0147] S1.3 Analyze the time curve of the action amplitude change, mark the key frames at the moment of sharp amplitude change, and capture the starting frame and aborting frame of each action;
[0148] S1.4 Develop an adaptive screening strategy to select key-frame images I i , and form a key-frame dataset {I i}. Divide the screened key-frame dataset into a training set and a test set according to a certain proportion to ensure the generalization ability and accuracy of the model.
[0149] As a preferred embodiment of the present invention, the adaptive screening strategy includes the following:
[0150] S1.4.1) Quantification of action fluency: Calculate the optical flow between two adjacent frames of images to quantify the action fluency;
[0151] Let image frames I t and I t+1 be two consecutive frames of images respectively. The calculation method of the optical flow F t is as follows:
[0152]
[0153] where T is the time interval, and F t is the displacement vector between two frames, indicating the movement of the object between frames; if the displacement exceeds the threshold, it is considered that a large change has occurred in the action, which does not meet the fluency requirement;
[0154] Calculate the mean square error MSE of the optical flow between adjacent frames. Let F t and F t+1 be the optical flows of adjacent frames, and calculate their mean square error:
[0155]
[0156] Select the index. If MSE(F t , F t+1 ) is higher than the threshold of 0.15, it is considered that the action is not fluent, and this frame is excluded;
[0157] S1.4.2) Quantification of the visibility of limb parts: Use OpenPose to perform rapid skeleton detection on the image to obtain the preliminary visibility score of key parts; the visibility V(p i ) of each part is quantified according to the detection confidence score and the image quality, and the scoring range is [0,1]; use the spatial information of the fixed structures in the vehicle, such as the seat and the steering wheel, to judge whether a certain part is occluded; the occluded area O of the vehicle structure and the total visible area V(p i ) of this part are quantified using the following occlusion index:
[0158]
[0159] where |V(p i) ∩ O| is the occluded area of this part, |V(p i )| is the total visible area of this part. If the occlusion ratio is greater than the set threshold of 40%, then this part is considered invisible;
[0160] S1.4.3) Occlusion situation assessment: Set the percentage of the occlusion area as the evaluation index to determine whether the fixed structures in the vehicle (such as seats, steering wheels) cause significant occlusion to the driver's key parts (such as upper limbs, legs); Process the image through the image segmentation technology Mask R-CNN to extract the vehicle interior structure area R car and the driver area R driver , calculate the intersection and overlapping area of the two areas; Let |R car ∩R driver | be the intersection area of the vehicle interior structure and the driver area, |R driver | be the total area of the driver area, then the occlusion degree is:
[0161]
[0162] If the occlusion ratio exceeds 40%, then this frame is considered to have a significant occlusion problem.
[0163] As a preferred embodiment of the present invention, the specific method for step S2 data annotation and bone point selection is as follows:
[0164] Use the CVAT automatic annotation tool to perform bone point annotation on each selected key frame image I i , and design a dynamic annotation strategy according to the special scene in the vehicle and the entire continuous action of the driver from opening the door to holding the steering wheel. The specific annotation scheme is as follows:
[0165] Opening the door stage: Focus on annotating bone points such as shoulders, elbows, wrists, knees, and ankles to capture the movement trajectory of the driver reaching out to pull the door, and ensure that the core posture features of this action are recorded. This stage is the only stage that shows the driver's complete knee to ankle stage;
[0166] Sitting down stage: Focus on annotating key bone points such as the spine and knees to obtain the sitting posture angle and body rotation state of the driver.
[0167] Fastening the seat belt stage: Focus on capturing the subtle changes in the positions of upper body bone points such as shoulders, elbows, and waist, and accurately record the driver's side body movement to provide data support for subsequent analysis;
[0168] Holding the steering wheel stage: Focus on annotating bone points such as shoulders, elbows, wrists, and neck to ensure accurate recording of the final driving sitting posture;
[0169] The marked key skeletal points include the following parts: (1) nose; (2) neck; (3) left shoulder; (4) right shoulder; (5) chest; (6) left elbow; (7) right elbow; (8) left hip; (9) right hip; (10) left knee; (11) right knee; (12) left ankle; (13) right ankle.
[0170] The connection information between the marked key skeletal points will be used for subsequent skeletal ratio conversion and length calculation. These connections not only help capture the detailed characteristics of the driver's body structure but also provide the necessary parsing information for the system to support more accurate pose estimation and body type analysis.
[0171] Given the spatial limitations of the in-vehicle environment, some skeletal points may be occluded by structures such as doors and seats. The system dynamically classifies and strategically processes this according to the technology supported by CVAT:
[0172] Fully visible points: refer to the skeletal points that are completely exposed during the driver's actions, such as the head and shoulders. Such skeletal points are directly marked.
[0173] Occluded points: For example, during actions such as opening the door or turning around, some parts of the arms or legs may be occluded by in-vehicle structures. When there is a part that is partially occluded but the general position of the skeletal point can still be inferred, click "Switch occluded property" to present it in a dashed state indicating that this key point is partially occluded.
[0174] Invisible points: Such as the situation where the ankles are occluded by the seat after sitting down, click "Switch outside property" to present it in a hidden state indicating that this key point is invisible in this frame.
[0175] As a preferred embodiment of the present invention, step S3 includes:
[0176] In the feature extraction stage, the VitPose network is used to extract features from each key frame image I i to generate a preliminary feature representation related to the skeletal points;
[0177] Input the key frame image I i , assumed to be an image of size H×W×C, where H and W are the height and width of the image respectively, and C is the number of channels; divide the image into image patches of size P×P, and the number of pixels in each image patch is P 2 ·C. Assuming there are a total of N image patches, the calculation formula for N is:
[0178]
[0179] Each image patch is embedded into a d-dimensional feature vector space through a linear transformation to obtain the embedding vector e of each image patch ij∈R d , the specific formula is as follows:
[0180] e ij = W patch P ij + b patch
[0181] Among them, is the weight of the linear transformation, b patch is the bias term. Stack all the embedding vectors e ij into a matrix E ∈ R N×d , and input it into the Transformer encoder. The Transformer encoder captures the global context information of the image through the self-attention mechanism. Its core calculation formula is:
[0182]
[0183] Among them, Q, K, and V are the matrices of query, key, and value respectively, and d k is the scaling factor of the feature dimension. After being processed by multiple Transformer layers, the output feature matrix is:
[0184] z i = Transformer(E), z i ∈ R N×d
[0185] The output feature matrix z i contains the global context information of the image;
[0186] S3.2 Pose Estimation Layer (PoseEstimationLayer, SE_Layer) further predicts the preliminary skeletal point coordinates and connection diagram of the driver based on the feature matrix z i ;
[0187] Input the feature matrix z i ∈ R N×d generated by the Transformer encoder and the in-vehicle physical layout S car into SE_Layer. Through multimodal fusion, fuse S car with z i , and input it into the regression network g pose to predict the preliminary skeletal point coordinates:
[0188] P i = g pose (concat(z i , S car )), P i = {(x i , y i)}
[0189] Among them, P i ={(x i , y i )} represents the set of preliminary skeletal point coordinates, including the two-dimensional positions of each skeletal point. g pose is a regression function used to extract skeletal point coordinates from the feature matrix. Concat means splicing two kinds of information together. According to the set of skeletal point coordinates P i , construct the skeletal point connection graph S i , generated through the topological relationship between skeletal points, describing the skeletal point features of the driver:
[0190] S i =(V, E), V = P i , E ={(v k , v j )∣v k , v j ∈V and satisfy the connection relationship}
[0191] The output results include the set of preliminary skeletal point coordinates P i and the corresponding skeletal point connection graph S i ;
[0192] S3.3 Scene-aware Feature Integration Layer (SF_Layer) further optimizes the features of skeletal points through the in-vehicle environment perception mechanism;
[0193] Input the set of preliminary skeletal point coordinates P i , the preliminary skeletal point connection graph S i and the in-vehicle physical layout S car , perform linear transformations on the skeletal point connection graph S i and the in-vehicle physical layout S car respectively, and map them to the same high-dimensional feature space:
[0194]
[0195] Among them, and are the weight matrices of the linear transformation, and are the bias terms, and use the self-attention mechanism to calculate the correlation between the skeletal point features and the in-vehicle physical layout:
[0196]
[0197] Among them, d represents the feature dimension, S′ i (S′ car )T Reflects the degree of association between each skeletal point and the in-vehicle structure. Softmax is the normalized attention weight, Attn i is the weighted influence result of the in-vehicle physical layout on the skeletal point features, based on the optimized set of skeletal point coordinates Construct the corresponding skeletal point connection diagram
[0198]
[0199] Output the optimized set of skeletal point coordinates and the corresponding skeletal point connection diagram
[0200] S3.4 Spatio-Temporal Layer (ST_Layer) optimizes the temporal consistency and spatial coherence of skeletal points by analyzing the changes in skeletal points between consecutive frames of images;
[0201] Input the optimized set of skeletal point coordinates for consecutive frames and the corresponding skeletal point connection diagram Integrate them into Stack the multi-frame feature data into spatio-temporal feature blocks with dimension T×N×d, where T represents the number of frames in the time dimension. Use 3D convolution to capture temporal and spatial features and optimize the dynamic coherence of skeletal points:
[0202]
[0203] where, W ST ∈R d′×d×k×k×T is a 3D convolution kernel, d ′ is the output feature dimension, k is the spatial dimension of the convolution kernel, T is the range of the time dimension, z ST ∈R N×d′ is a high-dimensional feature matrix containing spatio-temporal context information, which considers both the position changes of skeletal points between multi-frame images and captures the temporal coherence of the driver's actions; after spatio-temporal convolution processing, the skeletal point coordinates are dynamically optimized using dynamic weights, and the formula is as follows:
[0204]
[0205] where, is the skeletal point coordinate of the t-th frame, α t is the dynamic weight, calculated from the spatio-temporal feature correlation of skeletal points in each frame:
[0206]
[0207] Among them, MLP is a multi-layer perceptron, which extracts the correlation score f between the skeletal point features of each frame and the overall dynamic optimization objective from the spatio-temporal feature z ST for subsequent weight calculation; construct the corresponding skeletal point connection diagram according to the dynamically optimized skeletal points t
[0208]
[0209] The topological relationship here is the same as but is recalculated based on the positions of the dynamically optimized skeletal points.
[0210] Output the set of coordinates of the dynamically optimized skeletal points and the corresponding skeletal point connection diagram which retains the consistency in time and space;
[0211] The S3.5 Proportion Adjustment Layer (PA_Layer) performs human body proportion correction on the set of dynamically optimized skeletal points and the corresponding skeletal point connection diagram to ensure that it meets the human anatomy standards.
[0212] Input the set of coordinates of the dynamically optimized skeletal points and the corresponding skeletal point connection diagram Predict the physical size of the skeletal points from the features through a regression network:
[0213] h i = MLP(z ST )
[0214] According to the standard human body proportion h max , calculate the scale factor λ of each skeletal point i :
[0215]
[0216] λ i is the scaling factor representing the physical size of each skeletal point relative to the standard human body proportion, which is used to ensure that the skeletal point prediction results can adapt to drivers of different body types. Therefore, the scale factor λ of each skeletal point i is calculated independently. Adjust the skeletal point coordinates using the scale factor:
[0217]
[0218] Construct the corresponding skeletal point connection diagram S according to the corrected skeletal point coordinates final .
[0219] The output includes a set of scale-corrected bone point coordinates. And the corresponding bone point connection diagram S final , as the final two-dimensional skeleton point recognition result.
[0220] As a preferred embodiment of the present invention, step S4 specifically includes the following steps of three-dimensional reconstruction based on two-dimensional skeleton points:
[0221] According to the two-dimensional bone point coordinate set Combine the internal structure information of the vehicle with the JOTR algorithm to complete the 3D skeleton point reconstruction and finally generate a 3D skeleton point connection diagram And extract the driver's body parameters.
[0222] S4.1 Input the interior space data (such as seat height, width, depth, and interior length and width) for modeling, and use the modeling tool Blender to build the constraint structure of the three-dimensional space in the vehicle;
[0223] S4.2 Using JOTR algorithm combined with 2D skeleton point coordinate set Calculate the 3D coordinates of each bone point Provide input for subsequent 3D reconstruction:
[0224]
[0225] Initial 3D bone points There is a certain error, so the self-supervised learning framework of the Generative Adversarial Network (GAN) is used to correct it and convert the three-dimensional coordinates Projected onto a two-dimensional plane, we get the projection
[0226]
[0227] The projection function is calculated based on camera parameters and perspective relationship.
[0228] S4.3 Discriminator D projects the points onto the two-dimensional plane and the original 2D point The error feedback generator G between them minimizes the loss function L self , continuously optimize the generator so that the three-dimensional coordinates of its output are more consistent with the two-dimensional projection of the real image points:
[0229]
[0230] After the correction is completed, the generator G outputs the corrected 3D bone points:
[0231]
[0232] S4.4 Generate a 3D skeletal point connection diagram based on the corrected set of 3D skeletal points Generate a 3D skeletal point connection diagram Its topological relationship is the same as that of the 2D connection diagram S final but is recalculated in 3D space, calculating the Euclidean distance d between 3D skeletal points ij to characterize the limb length characteristics of the driver:
[0233]
[0234] where d ij represents the 3D distance between the i-th skeletal point and the j-th skeletal point. The set of distance features {d ij} is input into a multi-layer perceptron (MLP) and mapped to the height of the driver and body shape parameters ShapeParams:
[0235]
[0236] Finally, a more accurate driver height and its body shape parameters ShapeParams are output.
[0237] Based on the predicted results of the driver's height and body shape, the system can automatically optimize and adjust multiple position parameters of the car seat, the position of the seat belt, and the deployment strategy of the airbag, etc. The system formulates a dynamic adjustment plan based on the ISO 7250-1:2017 "Basic Anthropometric Technical Design" standard and the SAE J826-2021 "Equipment for Defining and Measuring Vehicle Seat Adjustment" data to achieve precise matching of the driver's body shape requirements. At the same time, all data is processed and stored locally on the vehicle-mounted side, effectively protecting the driver's personal privacy and security.
[0238] As a preferred embodiment of the present invention, the dynamic adjustment strategy in step S5 is as follows:
[0239] S5.1 Seat front-back adjustment: Ensure the appropriate distance between the driver and the pedal and the steering wheel to achieve the best control and comfort.
[0240] Objective: Ensure the appropriate distance between the driver and the pedal and the steering wheel to achieve the best control and comfort.
[0241] Input parameters: Driver's leg length (vertical distance from the hip to the knee), distance from the knee to the seat (straight-line distance from the knee to the front edge of the seat), pedal distance (distance from the driver's sole to the accelerator and brake pedals).
[0242] Adjustment method:
[0243] Leg length data acquisition: Calculate the thigh length and knee position of the driver to deduce the optimal angle of the leg (usually between 90° and 120° for the knee angle).
[0244] Forward and backward adjustment value:
[0245] For drivers with long legs (thigh length > 50 cm): The seat is moved forward by about 5 - 30 cm;
[0246] For drivers with short legs (thigh length < 45 cm): The seat is moved backward by about 5 - 7 cm, and the distance between the knee and the front edge of the seat should be greater than 30 cm.
[0247] Dynamic adjustment: According to the movement of the driver's foot, if it is detected that the driver has knee discomfort during emergency braking or acceleration, the seat can automatically fine-tune the front and back position by 1 - 2 cm.
[0248] S5.2 Seat height adjustment: Ensure that the driver obtains the best field of vision and reduce the interference of the roof or other components;
[0249] Goal: Ensure that the driver obtains the best field of vision and reduce the interference of the roof or other components.
[0250] Input parameters: Height, safety distance between the roof and the head (vertical distance from the driver's head to the roof).
[0251] Adjustment method:
[0252] Height data acquisition: By detecting the shoulder and head positions of the driver, calculate the eye height and ensure the best angle of the line of sight with the windshield.
[0253] Height adjustment value:
[0254] For drivers with a height greater than 180 cm: The seat is raised by 5 - 8 cm to ensure that the eye position is parallel to the center line of the window and the head is more than 5 cm away from the roof;
[0255] For drivers with a height below 160 cm: The seat is lowered by 5 - 7 cm to ensure that the eye position is not lower than the lower edge of the window and the head is more than 7 cm away from the roof.
[0256] Dynamic feedback: Real-time monitor the change of the driver's line of sight through an in-vehicle camera or sensor and dynamically adjust the seat height (for example, adjust 3 - 5 cm) to ensure that the field of vision is not affected by changes in the vehicle interior environment.
[0257] S5.3 Seat backrest angle adjustment: Ensure sufficient support for the back and reduce driver fatigue;
[0258] Goal: Ensure sufficient support for the back and reduce driver fatigue.
[0259] Input parameters: Driver's back curvature (natural bending angle between the waist and chest), sitting posture angle (angle between the back and the seat).
[0260] Adjustment method:
[0261] Back data acquisition: Detect the bone point positions of the driver's waist and chest, calculate the angle between the two, obtain the natural back curvature, and adjust the angle of the seat backrest.
[0262] Backrest angle adjustment value:
[0263] For a relatively upright posture (back angle < 100°): The seat backrest reclines 5 - 7°, so that the back reaches the natural bending angle;
[0264] For a relatively relaxed posture (back angle > 110°): The seat backrest tilts forward 5 - 7°, ensuring a comfortable angle (100° - 110°) is formed between the back and the seat.
[0265] Dynamic feedback: If the driver stays in a relaxed posture for a long time, the seat will automatically detect and finely adjust the backrest angle (usually 1 - 3°) through sensors to reduce the burden on the back.
[0266] S5.4 Seat and steering wheel linkage adjustment: Ensure that the driver can comfortably operate the steering wheel and avoid unnatural arm extension;
[0267] Goal: Ensure that the driver can comfortably operate the steering wheel and avoid unnatural arm extension.
[0268] Input parameters: Arm length (distance from the shoulder to the wrist), distance from the shoulder to the steering wheel (based on bone point recognition data).
[0269] Adjustment method:
[0270] Arm length data acquisition: Calculate the relative positions of the driver's shoulder and elbow through bone points to determine the appropriate arm angle.
[0271] Seat front - back adjustment value:
[0272] For drivers with long arms (arm length > 60 cm): The seat moves forward 3 - 5 cm to ensure that the elbow bending angle remains around 120°, facilitating flexible operation of the steering wheel;
[0273] For drivers with short arms (arm length < 55 cm): The seat moves backward 3 - 5 cm to ensure that the driver can naturally hold the steering wheel and both hands can operate comfortably.
[0274] Steering wheel angle adjustment: If the seat position is adjusted, the height and angle of the steering wheel are automatically adjusted by 1 - 3° to keep it within the ideal arm operation range.
[0275] S5.5 Seat Belt Adjustment: Ensure that the seat belt can accurately restrain the driver's body and provide the best safety protection.
[0276] Objective: Ensure that the seat belt can accurately restrain the driver's body and provide the best safety protection.
[0277] Input Parameters: Shoulder height (vertical distance from the shoulder to the ground), chest and waist circumference (distance between the chest and the waist).
[0278] Adjustment Method:
[0279] Shoulder Belt Adjustment: Automatically adjust the height of the shoulder belt of the seat belt to ensure that the shoulder belt is aligned with the center of the driver's shoulder. Usually, the shoulder belt should be located 3 - 5 cm from the center line of the shoulder.
[0280] Waist Belt Adjustment: Adjust the position of the waist belt according to the waist circumference data to ensure that the waist belt fits the pelvic area. If the waist circumference is large, the seat will automatically fine-tune the position to avoid the seat belt being too tight.
[0281] Dynamic Adjustment: During driving, if the driver's posture changes significantly (such as bending, turning, etc.), the system will automatically adjust the tension and position of the seat belt to ensure maximum protection for the driver in the event of a collision.
[0282] As a preferred embodiment of the present invention, step S6 includes:
[0283] S6.1 Use three-dimensional bone point coordinates Extract the set of key part coordinates P of the driver key :
[0284]
[0285] Among them, respectively represent the calibrated three-dimensional coordinates of the head, chest, and shoulders, represents the calibrated three-dimensional coordinates of the key part k;
[0286] S6.2 Calculate the driver's body shape characteristic range according to the body shape parameters ShapeParams:
[0287]
[0288] Among them, BodyWidth represents the horizontal distance between the left and right endpoints of the driver's shoulders, and are respectively the x-axis values of the calibrated three-dimensional coordinates of the right shoulder and the left shoulder; BodyHeight represents the vertical distance from the driver's head to the buttocks, and respectively represent the z-axis values of the calibrated three-dimensional coordinates of the head and the buttocks;
[0289] S6.3 Calculate the optimal center position P key of the airbag deployment based on the key part coordinates P of the driver deploy and the coverage radius R deploy :
[0290]
[0291] where P deploy is the center point of the airbag deployment, calculated based on the average coordinates of the driver's key parts, and R deploy is the coverage radius of the airbag, determined by the maximum three-dimensional distance of the key parts;
[0292] S6.4 Combine the in-vehicle collision sensor data to dynamically adjust the triggering conditions of the airbag; the calculation formula for the triggering threshold T trigger is:
[0293] T trigger = w 1 F collision + w 2 θ impact + w 3 Δ pose
[0294] where F collision is the collision force, θ impact is the collision angle, Δ pose is the driver's sitting posture deviation, and w 1 , w 2 are weight parameters used to adjust the influencing factors of the triggering conditions;
[0295] The deployment force P of the airbag inflate is calculated according to the driver's body shape parameters and the severity of the collision:
[0296] P inflate = f(ShapeParams, Collision Severity)
[0297] S6.5 Dynamically adjust the airbag deployment direction and deployment strategy by real-time monitoring of the driver's sitting posture changes;
[0298] The adjusted airbag deployment angle α adjusted = α default + kΔ pose is expressed as:
[0299] α adjusted = α default + k·Δ pose
[0300] where α default is the default deployment angle, k is the sitting posture correction coefficient, ·Δpose is the detected sitting posture deviation amount; obtain the deployment center point P of the airbag deploy , coverage range R deploy , trigger threshold T trigger and deployment force P inflate , deployment angle α after the driver's dynamic adjustment adjusted , provide the adjustment parameters of the airbag to ensure that the key parts of the driver can be accurately protected when the airbag deploys, and reduce the risk of secondary injury during a collision.
Claims
1. A method for identifying the physical features of a driver and passenger and automatically adjusting the seat and airbag, characterized in that: The following steps are involved: S1 Image capture and dataset construction: The in-car camera is used to capture a continuous sequence of images of the entire process of the driver getting on the car. The action recognition algorithm based on 2D CNN+LSTM is used to analyze the continuous image sequence, obtain the driver's action features to form a key frame dataset, and the key frame dataset is screened. The screened key frame dataset is divided into a training set and a test set according to the proportion; S2 key skeleton point annotation uses automatic annotation tools to annotate each key frame image I i Mark the skeleton points and generate the marked skeleton point coordinate set P label ={(x label ,y label )},x label Represents the horizontal coordinate of a bone point on the image plane, y label Indicates the vertical coordinate of a key bone point on the image plane. This annotation data is used for subsequent training of the two-dimensional bone point recognition system. S3 2D Skeleton Point Recognition System Construction for Key Frame Image I i Perform feature extraction to generate a preliminary skeleton point coordinate set P i And the corresponding bone point connection diagram S i ; Combined with the physical layout of the car, the optimization result set is generated through multimodal fusion and self-attention mechanism and Analyze the dynamic changes of skeleton points in multiple frames and use spatiotemporal convolution to generate dynamic optimization results and Output a set of two-dimensional bone point coordinates through human body proportion correction And the corresponding bone point connection diagram S final ; S4 obtains the interior space data based on the three-dimensional reconstruction of the two-dimensional skeleton points, and uses the modeling tool Blender to model the three-dimensional space inside the car; Combined with JOTR algorithm to calculate the coordinates of three-dimensional bone points A self-supervised learning framework using generative adversarial networks Perform correction to generate corrected three-dimensional coordinates Calculate the Euclidean distance d between the 3D bone points based on the correction results ij , and is mapped to the driver's height and body shape through a multi-layer perceptron MLP, outputting the driver's exact height and shape parameters ShapeParams; S5 seats automatically adjust according to driver's height ShapeParams: analyzes the position and sitting angle of the driver's head, shoulders, and waist, and dynamically adjusts the seat height, backrest angle, and cushion angle to match the driver's body shape. Adopting dynamic adjustment strategy, it monitors the driver's posture changes during long driving in real time and responds intelligently to optimize the seat settings to reduce fatigue; S6 airbag automatic adjustment by analyzing the driver's 3D bone point coordinates Body shape parameters ShapeParams and real-time sitting posture information are used to dynamically optimize the deployment position, deployment range and triggering strategy of the airbag. A real-time monitoring mechanism is used to ensure that the airbag can accurately protect the driver's key parts, such as the head, chest and shoulders, when deployed, reducing the risk of secondary injuries during a collision.
2. The method for identifying the physical features of a driver and passenger and automatically adjusting the seat and airbag according to claim 1, characterized in that: The step S1 of image capture and data set construction specifically includes: S1.1 2D convolutional neural network extracts spatial features from each frame of a continuous image sequence, such as the outline of the human body, the position of key limbs and their movement trajectory; S1.2 uses the LSTM model to process the temporal information in the continuous image sequence, captures the continuous action characteristics of the driver when getting on the car, and analyzes the action transition characteristics in a set of adjacent frames to identify the turning points of the action; S1.3 analyzes the time curve of the change of the action amplitude, marks the key frame at the moment when the amplitude changes sharply, and captures the start frame and stop frame of each action; S1.4 Develop an adaptive screening strategy to select key frame images I i , forming a key frame dataset {I i }, the filtered key frame dataset is divided into training set and test set in proportion to ensure the generalization ability and accuracy of the model.
3. The method for identifying the physical features of a driver and passenger and automatically adjusting the seat and airbag according to claim 2 is characterized in that: The adaptive screening strategy includes the following: S1.4.1) Action smoothness quantification: quantify the action smoothness by calculating the optical flow between two adjacent frames; Assume that image frame I t and I t+1 They are two consecutive frames of images, and the optical flow F t The calculation method is: Where T is the time interval, F t It is the displacement vector between two frames, indicating the movement of the object between frames; if the displacement exceeds the threshold, it is considered that the action has changed significantly and does not meet the fluency requirement; Calculate the mean square error (MSE) of the optical flow between adjacent frames, and set F t and F t+1 For the optical flow of adjacent frames, calculate its mean square error: Select the indicator, if MSE(F t , F t+1 ) is higher than the threshold of 0.15, the motion is considered to be unsmooth and the frame is discarded; S1.4.2) Quantification of body part visibility: Use OpenPose to perform fast skeleton detection on the image to obtain preliminary visibility scores of key parts; the visibility V(p i ) is quantified based on the detection confidence score and image quality, with a score range of [0,1]; the spatial information of the fixed structure inside the car is used to determine whether a certain part is blocked; the blocked area O of the structure inside the car and the total visible area V of the part (p i ), quantized using the following occlusion index: Among them, |V(p i )∩O| is the area of the part blocked, |V(p i )| is the total visible area of the part. If the occlusion ratio is greater than the set threshold of 40%, the part is considered invisible; S1.4.3) Occlusion assessment: Set the percentage of occlusion area as the evaluation index to determine whether the fixed structure inside the car has a significant occlusion on the driver's key parts; use the image segmentation technology Mask R-CNN to process the image and extract the structure area R inside the car. car and driver area R driver , calculate the intersection and overlapping area of the two regions; let |R car ∩R driver| is the intersection area of the interior structure and the driver area, |R driver | is the total area of the driver area, then the degree of occlusion is: If the occlusion ratio exceeds 40%, the frame is considered to have a large occlusion problem.
4. The method for identifying the physical features of a driver and passenger and automatically adjusting the seat and airbag according to claim 1, characterized in that: The specific method of data annotation and skeleton point selection in step S2 is: S2.1 uses the CVAT automatic annotation tool to automatically annotate each key frame image I i Skeleton point annotation is performed, and a dynamic annotation strategy is designed based on the special scenes in the car and the driver's continuous actions from opening the door to holding the steering wheel. The annotation scheme is as follows: Door opening stage: focus on marking the shoulder, elbow, wrist, knee, ankle and other bone points to capture the driver's motion trajectory of reaching out to open the door, ensuring that the core posture features of the action are recorded. This stage is the only stage that shows the driver's complete knee to ankle stage; Sitting stage: focus on marking key bone points such as the spine and knees to obtain the driver's sitting angle and body rotation status; Seat belt fastening stage: focus on capturing subtle changes in the position of upper body bone points such as shoulders, elbows and waist, accurately record the driver's sideways movements, and provide data support for subsequent analysis; Steering wheel holding stage: focus on marking the skeletal points such as shoulders, elbows, wrists and neck to ensure accurate recording of the final driving posture; The key bone points annotated in S2.2 include the following parts: (1) nose; (2) neck; (3) left shoulder; (4) right shoulder; (5) chest; (6) left elbow; (7) right elbow; (8) left hip; (9) right hip; (10) left knee; (11) right knee; (12) left ankle; (13) right ankle.
5. The method for identifying the physical features of a driver and passenger and automatically adjusting the seat and airbag according to claim 1, characterized in that: The step S3 comprises: S3.1 Feature extraction stage: Using the VitPose network to extract each key frame image I i Perform feature extraction to generate preliminary feature representation related to skeleton points; Input key frame image I i , assuming an image of size H×W×C, where H and W are the height and width of the image respectively, and C is the number of channels; divide the image into image blocks of size P×P, and the number of pixels in each image block is P 2 C, assuming there are N image blocks in total, the calculation formula for N is: Each image block is embedded into the d-dimensional feature vector space through linear transformation to obtain the embedding vector e of each image block ij ∈R d , the specific formula is as follows: e ij =W patch P ij +b patch in, is the weight of the linear transformation, b patch is the bias term, which transforms all embedding vectors e ij Stacked into a matrix E∈R N×d , is input to the Transformer encoder, which captures the global context information of the image through the self-attention mechanism. Its core calculation formula is: where Q, K, and V are the matrices of query, key, and value, respectively. k is the scaling factor of the feature dimension. After being processed by multiple Transformer layers, the feature matrix is output: With i =Transformer(E),z i ∈R N×d The output feature matrix z i Contains the global context information of the image; S3.2 Pose estimation layer: based on feature matrix z i Further predict the driver's preliminary skeleton point coordinates and connection diagram; The feature matrix z generated by the Transformer encoder i ∈R N×d and the physical layout of the vehicle car Input to SE-Layer, through multi-modal fusion, S car With z i Fusion and input into the regression network g pose , predict the initial bone point coordinates: P i =g pose (concat(z i ,S car )),P i ={(x i ,y i )} Among them, P i ={(x i ,y i )} represents the initial set of bone point coordinates, including the two-dimensional position of each bone point, g pose It is a regression function used to extract the coordinates of the bone points from the feature matrix. concat means to concatenate the two types of information together according to the bone point coordinate set P i , construct the skeleton point connection graph S i , generated through the topological relationship between skeleton points, describing the skeleton point features of the driver: S i =(V, E), V = P i , E={(v k , v j )|v k , v j ∈V, and satisfy the connection relationship} The output results include a preliminary set of bone point coordinates P i And the corresponding bone point connection diagram S i ; S3.3 Scene perception and feature integration layer: further optimizes the features of skeleton points through the in-vehicle environment perception mechanism; Input the initial bone point coordinate set P i , Preliminary skeleton point connection diagram S i and the physical layout of the vehicle car , for the skeleton point connection diagram S i and the physical layout of the vehicle car Perform linear transformations respectively and map them to the same high-dimensional feature space: in, and is the weight matrix of the linear transformation, and is a bias term, which uses the self-attention mechanism to calculate the correlation between the skeleton point features and the physical layout of the car: Among them, d represents the feature dimension, S′ i (S′ car ) T reflects the degree of association between each skeleton point and the structure inside the car, softmax is the normalized attention weight, Attn i It is the weighted influence of the physical layout of the car on the skeleton point features. According to the optimized skeleton point coordinate set Construct the corresponding bone point connection diagram Output optimized bone point coordinate set And the corresponding bone point connection diagram S3.4 spatiotemporal integration layer: optimizes the temporal consistency and spatial coherence of the skeleton points by analyzing the changes of the skeleton points between consecutive multi-frame images; Input a continuous multi-frame optimized set of bone point coordinates And the corresponding bone point connection diagram Integrate it into Multi-frame feature data The stack is a spatiotemporal feature block with a dimension of T×N×d, where T represents the number of frames in the time dimension. 3D convolution is used to capture temporal and spatial features and optimize the dynamic coherence of skeleton points: Among them, W ST ∈R d′×d×k×k×T is a 3D convolution kernel, d′ is the output feature dimension, k is the spatial dimension of the convolution kernel, T is the range of the time dimension, z ST ∈R N×d′ It is a high-dimensional feature matrix containing spatiotemporal context information. After spatiotemporal convolution, the dynamic weights are used to dynamically optimize the coordinates of the bone points. The formula is as follows: in, is the coordinate of the bone point in the tth frame, α t is the dynamic weight, which is calculated by the correlation of the spatiotemporal features of the skeleton points in each frame: Among them, MLP is a multi-layer perceptron, which is based on the spatiotemporal features z ST Extract the correlation score f between the skeleton point features of each frame and the overall dynamic optimization target t , used for subsequent weight calculation; according to the dynamically optimized bone points Construct the corresponding bone point connection diagram Output dynamically optimized bone point coordinate set And the corresponding bone point connection diagram It preserves consistency in time and space; S3.5 scale adjustment layer: dynamically optimized bone point set And the corresponding bone point connection diagram Perform body proportion correction to ensure it conforms to human anatomy standards; Input the dynamically optimized bone point coordinate set And the corresponding bone point connection diagram Predict the physical size of the skeleton points from the features through a regression network: h i =MLP(from ST ) According to the standard human body proportions h max , calculate the scale factor λ for each bone point i : λ i It is a scaling factor that represents the physical size of each bone point relative to the standard human body proportion. It is used to ensure that the bone point prediction results can adapt to drivers of different body shapes. The scale factor is used to adjust the bone point coordinates: According to the corrected bone point coordinates Construct the corresponding skeleton point connection graph S final ; The output includes a set of scale-corrected bone point coordinates. And the corresponding bone point connection diagram S final , as the final two-dimensional skeleton point recognition result.
6. The method for identifying the physical features of a driver and passenger and automatically adjusting the seat and airbag according to claim 1, characterized in that: The three-dimensional reconstruction based on the two-dimensional skeleton points in step S4 specifically includes: According to the two-dimensional bone point coordinate set Combine the internal structure information of the vehicle with the JOTR algorithm to complete the 3D skeleton point reconstruction and finally generate a 3D skeleton point connection diagram And extract the driver's body parameters; S4.1 Input the interior space data of the vehicle for modeling, and use the modeling tool Blender to construct the constraint structure of the three-dimensional space inside the vehicle; S4.2 Using JOTR algorithm combined with 2D skeleton point coordinate set Calculate the 3D coordinates of each bone point Provide input for subsequent 3D reconstruction: The three-dimensional coordinates are corrected by the self-supervised learning framework of the generative adversarial network. Projected onto a two-dimensional plane, we get the projection S4.3 Discriminator D projects the points onto the two-dimensional plane and the original 2D point The error feedback generator G between them minimizes the loss function L self , continuously optimize the generator so that the three-dimensional coordinates of its output are more consistent with the two-dimensional projection of the real image points: After the correction is completed, the generator G outputs the corrected 3D bone points: S4.4 Based on the corrected 3D bone point set Generate 3D bone point connection diagram Its topological relationship is similar to the two-dimensional connection graph S final Same as, but recalculated in 3D space, computing the Euclidean distance d between 3D bone points ij , used to characterize the driver's limb length characteristics: Among them, d ij represents the three-dimensional distance between the i-th bone point and the j-th bone point, and the distance feature set {d ij } Input to the multi-layer perceptron (MLP) and mapped to the driver's height And shape parameters ShapeParams: Final output driver height And its shape parameters ShapeParams.
7. The method for identifying the physical features of a driver and passenger and automatically adjusting the related system according to claim 1, characterized in that: The dynamic adjustment strategy of step S5 is as follows: S5.1 Seat front and rear adjustment: ensures the driver is at a suitable distance from the pedals and steering wheel to achieve optimal control and comfort; S5.2 Seat height adjustment: ensures the driver has the best field of vision and reduces interference from the roof or other components; S5.3 Seat back angle adjustment: ensures adequate back support and reduces driver fatigue; S5.4 seat and steering wheel linkage adjustment: ensures the driver can operate the steering wheel comfortably and avoids unnatural arm extension; S5.5 Seat belt adjustment: Ensures that the seat belt can accurately restrain the driver's body and provide the best safety protection.
8. The method for identifying the physical features of a driver and passenger and automatically adjusting the seat and airbag according to claim 1, characterized in that: The step S6 comprises: S6.1 Using 3D bone point coordinates Extract the coordinate set P of the driver's key parts key : in, Represent the corrected 3D coordinates of the head, chest and shoulders respectively, represents the corrected three-dimensional coordinates of the key part k; S6.2 Calculate the driver's body shape feature range based on the body shape parameters ShapeParams: Among them, BodyWidth represents the horizontal distance between the left and right endpoints of the driver's shoulders. and are the x-axis values of the corrected three-dimensional coordinates of the right shoulder and the left shoulder respectively; BodyHeight represents the vertical distance from the top of the driver's head to the hips, and Represent the z-axis values of the corrected three-dimensional coordinates of the head and hips respectively; S6.3 According to the coordinates of the driver's key parts P key , calculate the optimal center position P of airbag deployment deploy and coverage radius R deploy : Among them, P deploy is the center point of airbag deployment, calculated based on the average coordinates of the driver’s key parts, R deploy is the coverage radius of the airbag, which is determined by the maximum three-dimensional distance of the key parts; S6.4 dynamically adjusts the triggering conditions of the airbags based on the collision sensor data in the vehicle; the triggering threshold T trigger The calculation formula is: T trigger =w1F collision +w2θ impact +w3Δ pose Among them, F collision is the collision force, θ impact is the collision angle, Δ pose is the deviation of the driver's sitting posture, w1, w2, w3 are weight parameters used to adjust the influencing factors of the triggering conditions; Airbag deployment force P inflate Calculated based on the driver's body parameters and collision severity: P inflate =f(ShapeParams,Collision Severity) S6.5 dynamically adjusts the airbag deployment direction and deployment strategy by monitoring the driver's sitting posture changes in real time; Adjusted airbag deployment angle α adjusted =α default +kΔ pose It is expressed as: a adjusted =a default +k·D pose Among them, α default is the default unfolding angle, k is the sitting posture correction coefficient, ·Δ pose is the monitored sitting posture deviation; the deployment center point P of the airbag is obtained deploy , Coverage R deploy , trigger threshold T trigger and the expansion force P inflate 、Deployment angle α after dynamic adjustment by the driver adjusted , providing the adjustment parameters of the airbag to ensure that the airbag can accurately protect the driver's key parts when deployed, reducing the risk of secondary injuries during a collision.
Citation Information
Patent Citations
Airbag self-adaptive control method based on driver posture real-time 3D modeling
CN111703393A
Vehicle seat adjusting method, device and equipment based on human skeleton recognition and medium
CN118418851A
Automobile driving seat adjusting system and method based on driver and passenger characteristics
CN118894023A
Adjusting method and device for vehicle seat and vehicle
CN119218065A
Cited By
Seat comfort optimization calculation method based on sensor and AI modeling
CN120716536A
High-speed rail seat comfort automatic adaptation control method based on passenger behavior habit machine learning
CN120986478A
Vehicle personalized adaptation system and method
CN121516012A