Pedestrian crossing intention prediction method fusing skeleton information in natural driving environment
By extracting pedestrian bone point data and combining machine learning models, mobile and static posture features are developed, which solves the transparency and interpretability problems of autonomous vehicles in the prediction of pedestrian crossing intentions, and improves prediction accuracy and decision-making reliability.
Patent Information
- Application Number
- CN202510163655.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-14
- Publication Date
- 2025-07-01
AI Technical Summary
The lack of transparent and explainable methods in the prediction of pedestrian crossing intentions by existing autonomous vehicles has led to insufficient operational reliability in complex environments, making it difficult to understand pedestrian behavioral motivations, and affecting traffic safety.
The OpenPose algorithm is used to extract pedestrian bone point data, combine it with the XGBoost model for prediction, and provide global and local interpretation through SHAP and LIME interpretation technology to develop dynamic and static pose features to improve prediction accuracy and interpretability.
It improves the accuracy and interpretability of pedestrian crossing intention prediction, enhances the reliability of decision-making of autonomous vehicles in complex environments, and enhances the public's trust in autonomous driving technology.
Smart Images

Figure CN120236263A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method for predicting pedestrian crossing intention by integrating skeletal information in a natural driving environment, and belongs to the field of pedestrian behavior prediction. Background Art
[0002] With the increasing development momentum of the autonomous driving vehicle industry in recent years, especially the continuous efforts of technology companies represented by Tesla, Google, Baidu, etc. to quickly launch highly automated vehicles, the future traffic environment will become more complex. As an important participant in the traffic environment, pedestrians have flexible and variable movements and are greatly affected by the surrounding environment, which is extremely likely to cause potential traffic safety hazards. Therefore, correctly understanding the behavior and motivation of pedestrians is crucial for ensuring traffic safety. However, compared with human drivers, current autonomous driving vehicles have significant gaps in risk perception and decision-making processes, which may lead to their inability to fully understand the behavior of other road users. Therefore, verifying the reliability of autonomous driving vehicles operating in various complex environments is an essential link before large-scale promotion. An example of the definition given for the term "scenario" is the response of a specified test vehicle to a given action of a pedestrian crossing the road. Thus, understanding pedestrian crossing intention is crucial for the safe driving of autonomous driving vehicles and further improving public acceptance.
[0003] The purpose of predicting pedestrian crossing intention is to predict whether a target pedestrian will cross the road at a certain future moment. The booming development of computer vision technology has brought powerful tools for pedestrian target detection and trajectory prediction. However, predicting pedestrian intention requires a deeper semantic understanding of human physical and mental activities. Even in existing related research, although there are many advanced algorithms that can provide high prediction accuracy, the decision-making mechanism behind the algorithms still lacks necessary explanations. In addition, although existing pedestrian intention prediction research has begun to use skeletal information to improve prediction effects, it often ignores the clear physical definition of many pose-related features, thus hindering the popularization of existing methods. Therefore, it is necessary to seek a more general, transparent, and interpretable method to provide a basis for the autonomous driving system to make judgments and decisions during the interaction between vehicles and pedestrians in order to gain public trust. Summary of the Invention
[0004] The purpose of the present invention is to overcome the defects of the above-mentioned existing technologies and provide a method for predicting pedestrian crossing intention by integrating skeletal information in a natural driving environment, which can endow a series of dynamic and static pose features developed based on skeletal information with clear physical meanings, provide global and local explanations for the specific effects of various influencing factors on the prediction effect, and at the same time improve the prediction accuracy and interpretability, and provide effective suggestions for improving future autonomous driving vehicle scenario tests.
[0005] The present invention adopts the following technical solutions:
[0006] The present invention provides a method for predicting pedestrian crossing intention by integrating skeletal information in a natural driving environment, including the following steps:
[0007] S1: Prepare a natural driving basic data set: The data set used contains video segments captured from the perspective of a driving recorder in a natural driving environment, and is divided into segments of 5 - 15 seconds to better focus on the interaction with a certain pedestrian or certain groups of pedestrians. Based on the video segments, the position coordinates of the pedestrian identification box and the feature text annotations are extracted to form data set a, which is stored in XML format;
[0008] S2: Extract pedestrian skeletal point data: Use the OpenPose algorithm to extract the coordinates and recognition confidence of 25 skeletal points of all recognizable pedestrians in the video, and store them in JSON format corresponding to the video serial number and frame serial number respectively, and record them as data set b;
[0009] S3: Target matching between multi-source data: Since the annotation storage format in the basic data set is different from the skeletal point data format, and the IDs of pedestrians in the two data sources also lack a corresponding relationship, specific rules need to be formulated to achieve the matching of the same target between different data sources;
[0010] S4: Extract dynamic and static pose features: Based on the skeletal point data extracted in S2, develop new dynamic and static pose features that are helpful for judging pedestrian crossing intention, including the front, rear, left, and right body orientations of the pedestrian relative to the vehicle, the smaller value of the knee bending angles of the left and right legs, and the relative ratios of the step length and shoulder width in the driver's perspective to the overall body width, a total of 7 static variables. After sampling by extracting frames at fixed intervals, extract 3 dynamic variables including whether there is an action of alternating the front and rear of the two legs, whether there is a switching of the front and rear orientations, and whether there is a switching of the left and right orientations within 1 second for the pedestrian;
[0011] S5: Train a machine learning model for prediction and explain the specific influence of each variable: Take whether the pedestrian crosses the street marked in the basic data set as the dependent variable, use XGBoost to construct a pedestrian crossing intention prediction model, debug the model using the grid search method on the training set to obtain the parameter combination with the best effect, then verify the model effect on the test set, use the machine learning explainable technology SHAP to sort the importance of each variable, and give a global explanation of the specific influence of each variable. Use LIME to give a local explanation of the influence effect of each variable in a specific instance, and combine the two explanations to analyze the influence mechanism behind the pedestrian crossing decision.
[0012] Further, the vertex coordinates of the target pedestrian identification box in S1 are denoted as [x a1 ,y a1 ,x a2 ,ya2 , x a1 is the horizontal coordinate pixel value of the lower left vertex of the identification box, y a1 is the vertical coordinate pixel value of the lower left vertex of the identification box, x a2 is the horizontal coordinate pixel value of the upper right vertex of the identification box, y a2 is the vertical coordinate pixel value of the upper right vertex of the identification box.
[0013] Furthermore, the bone point coordinates and confidence levels extracted in S2 are recorded as [x bi , y bi , α bi , where i = 1, 2, … 24, x bi is the horizontal coordinate pixel value of the i-th bone point, y bi is the vertical coordinate pixel value of the i-th bone point, α bi is the recognition confidence level of the i-th bone point, ranging from 0 to 1. The 25 bone points and their corresponding specific joints are described as:
[0014] Bone point label Corresponding joint point definition Bone point label Corresponding joint point definition 0 Nose 13 Left knee 1 Base of neck 14 Left ankle 2 Right shoulder 15 Right eye 3 Right elbow 16 Left eye 4 Right wrist 17 Right ear 5 Left shoulder 18 Left ear 6 Left elbow 19 Left big toe 7 Left wrist 20 Left little toe 8 Midpoint of hips 21 Left heel 9 Right endpoint of hips 22 Right big toe 10 Right knee 23 Right little toe 11 Right ankle 24 Right heel 12 Left endpoint of hips
[0015] Furthermore, the feature text annotations in S1 include three types of features corresponding to the video frame numbers:
[0016] 1), Pedestrian features:
[0017]
[0018] 2), Environmental information:
[0019]
[0020] 3), Vehicle state: stopped, slow driving, fast driving, decelerating, accelerating.
[0021] Furthermore, in S3, the specific rule for matching the same target in different data sources is that the relative error between the coordinate differences of the pedestrian position reference points in different data sources is less than Δ%;
[0022] Among them, according to the pedestrian identification box coordinates [x a1 , y a1 , x a2 , y a2 given in the basic dataset, the pedestrian position reference point in the calibration dataset a is the midpoint of the identification box, denoted as M: (x a , y a ):
[0023] Meanwhile, the length of the calibration identification box is L = |y a1 - y a2 |, and the width is W = |xa1 -x a2 |;
[0024] Based on the recognized bone point data, the reference point of the pedestrian position in dataset b is calibrated as the centroid of the recognizable bone point framework, denoted as G: (x b , y b ),
[0025] where, δ(x bi , y bi ) is an indicator function for evaluating the recognizability of the i-th bone point. When the coordinate values of (x bi , y bi ) are specific pixel values rather than 0, the indicator function counts as 1;
[0026] Finally, the relative errors δ x and δ y between the coordinate differences of the pedestrian position reference points in different data sources are obtained:
[0027] Furthermore, the static pose variables in S4 are calibrated by the following method:
[0028] 1), The orientation of the pedestrian's body: Front (the eyes and nose of the pedestrian's front face appear in the picture with a relatively high recognition confidence, that is, α b0 , α b15 , α b16 are all greater than 0.5), Back (any one of α b0 , α b15 , α b16 is less than or equal to 0.5), Left (the number of recognizable body joint points on the left side of the pedestrian is greater than the number of recognizable body joint points on the right side, that is, α b5 , α b6 , α b7 , α b12 , α b13 , α b14 , α b16 , α b18 , α b19 , α b20 , α b21 with more than 0.5), Right (the number of recognizable body joint points on the right side of the pedestrian is greater than the left side, that is, α b2 , α b3 , α b4 , α b9 , α b10 , α b11 , α b15 , α b17 , α b22, α b23 , α b24 There are more with a value greater than 0.5);
[0029] 2) Knee flexion angle: After obtaining the knee joint flexion angles of the left and right legs of the pedestrian respectively, the smaller value of the two is taken as the value of this variable;
[0030]
[0031]
[0032] 3) The step length ratio takes the value of
[0033] 4) The shoulder width ratio takes the value of
[0034] Furthermore, the frame rate of the video material in the basic dataset is v frames per second. When expanding the sample in S4, the fixed interval selected is That is, starting from the first frame of the video when the pedestrian appears, one sample data is taken every 0.5 seconds for the processing of dynamic pose features. The front-back relationship and body orientation of the left and right legs of the pedestrian in the j-th frame are respectively compared with those in the th frame and the th frame. If there is a change, then the corresponding actions of whether there is an alternating movement of the two legs back and forth, whether there is a switching of the front-back orientation, and whether there is a switching of the left-right orientation are marked as "yes".
[0035] Furthermore, the ratio of the training set to the test set in S5 is 7:3. During the process of grid search using the GridSearchCV function in Python, 5-fold cross-validation is selected to determine the optimal parameter combination. The specific parameters involved in optimizing the XGBoost model are: the number of trees n_estimators, the minimum loss reduction value gamma, the maximum depth of the tree max_depth, the minimum number of samples in a leaf node min_child_weight, the sampling ratio of the tree subsample, the column sampling ratio of the samples colsample_bytree, the regularization parameter lambda (L2 regularization) and alpha (L1 regularization), and the learning rate learning_rate. The parameters adjusted on the training set are applied to the test set for model metric verification to examine the prediction accuracy.
[0036] Further, in S5, the SHAP package in Python is used to sort the variables that play a key role in the model according to their importance, and the positive or negative impact of each variable on the global result is given. In addition, by applying the LIME package in Python, the importance ranking and qualitative explanation of the variables that have a significant impact locally are obtained for specific instances in the sample. Combining the two analyzes the specific influence mechanism of pedestrian crossing decisions, providing a reference basis for the scenario testing of autonomous vehicles.
[0037] Compared with the prior art, the present invention has the following technical advantages:
[0038] The present invention provides a method for predicting pedestrian crossing intention by integrating skeletal information in a natural driving environment, and proposes a complete workflow for feature development, model construction and interpretive analysis. First, in the feature development stage, the dynamic and static pose features generated based on the 25 skeletal point data not only have clear physical meanings, can provide accurate descriptions of individual behaviors, but also introduce multi-dimensional feature vectors for the subsequent prediction model. In the model construction stage, the accuracy of the prediction result is ensured by optimizing the machine learning model.
[0039] The present invention provides a method for predicting pedestrian crossing intention by integrating skeletal information in a natural driving environment. In the interpretive analysis stage, combined with methods such as SHAP and LIME, a qualitative explanation is provided for the specific effects of each influencing factor on the prediction result, thus realizing the organic coordination between feature extraction, model optimization and result interpretation. This collaborative optimization process can significantly improve the standardization and reliability levels in the scenario testing of autonomous vehicles and address the public's doubts about the decision-making mechanism behind new autonomous driving technologies. Description of the Drawings
[0040] Figure 1 It is a schematic diagram of the skeletal information applied in the present invention.
[0041] Figure 2 It is a schematic diagram of the method flow of the present invention. Detailed Embodiments
[0042] The present invention will be described in detail below with reference to the drawings and specific embodiments. Features such as component models, material names, connection structures, control methods, algorithms, etc. that are not clearly stated in this technical solution are regarded as common technical features disclosed in the prior art.
[0043] Embodiment 1
[0044] This embodiment provides a method for predicting pedestrian crossing intention by integrating skeletal information in a natural driving environment, which can evaluate the rationality of the cycling route of the navigation platform and optimize the information presentation form and content thereof.
[0045] As Figure 1 and 2 shown, a method for predicting pedestrian crossing intention by integrating skeletal information in a natural driving environment includes the following steps:
[0046] S1: Prepare the natural driving basic dataset: The dataset used contains video segments taken from the perspective of a driving recorder in a natural driving environment, and is divided into segments of 5 - 15 seconds to better focus on the interaction with a certain pedestrian or certain groups of pedestrians. In addition, dataset a also contains the vertex coordinates [x a1 , y a1 , x a2 , y a2 of the pedestrian identification box corresponding to the video frame number, as well as 21 types of feature text annotations including pedestrian features, environmental information, and vehicle status, stored in XML format.
[0047] S2: Extract pedestrian skeletal point data: As Figure 1 shown, use the OpenPose algorithm to extract the coordinates and recognition confidence [x bi , y bi , α bi of 25 skeletal points of all recognizable pedestrians in the video, and store them in JSON format corresponding to the video number and frame number respectively, and record it as dataset b.
[0048] S3: Target matching between multi-source data: Since the annotation storage format in the basic dataset is different from the skeletal point data format, and there is also a lack of correspondence between the IDs of pedestrians in the two data sources, specific rules need to be formulated to achieve the matching of the same target between different data sources. The specific rule is that the midpoint M: (x a , y a ) of the pedestrian identification box in dataset a and the centroid G: (x b , y b ) of the recognizable skeletal point frame in dataset b have a relative error less than Δ%. When matching, select a specific target in dataset a, traverse and search all recognizable pedestrian data in dataset b, and find the corresponding target that meets the conditions to successfully match. This rule can be applied to the target matching between any dataset with pedestrian identification box coordinates and datasets containing skeletal information, and has strong universality.
[0049] S4: Extraction of dynamic and static posture features: Based on the skeleton point data extracted in S2, develop new dynamic and static posture features that help judge the intention of pedestrians to cross the road, including the front, rear, left, and right body orientations of pedestrians relative to the vehicle, the smaller value of the knee bending angles of the left and right legs, and the relative ratios of the step length and shoulder width to the entire body width from the driver's perspective, a total of 7 static variables. After sampling at a fixed interval, extract 3 dynamic variables within 1 second, including whether there is an action of alternating the front and rear of the two legs, whether there is a switch in the front and rear orientations, and whether there is a switch in the left and right orientations of the pedestrian.
[0050] S5: Train a machine learning model for prediction and explain the specific influence of each variable: Use whether the pedestrian crosses the road annotated in the basic dataset as the dependent variable, and use XGBoost to construct a prediction model for the intention of pedestrians to cross the road. Divide the sample set into a training set and a test set according to 7:3. Apply the grid search function GridSearchCV in Python to debug the model on the training set to obtain the optimal parameter combination, and then verify the model effect on the test set. Apply the machine learning interpretability technique SHAP to rank the importance of each variable and give a global explanation of the specific influence of each variable. Apply LIME to give a local explanation of the influence effect of each variable in a specific instance. Combine the two explanations to analyze the influence mechanism behind the pedestrian's decision to cross the road, providing a reference basis for optimizing the scenario testing of autonomous driving vehicles in the future.
[0051] The following uses specific embodiments to illustrate the present invention:
[0052] S1: Prepare a natural driving basic dataset: The dataset used in this example contains 346 video clips taken from the perspective of a driving recorder in a natural driving environment, and is divided into clips of 5 - 15 seconds to better focus on the interaction with a certain pedestrian or group of pedestrians. In addition, dataset a also contains the vertex coordinates [x a1 , y a1 , x a2 , y a2 of the pedestrian identification box corresponding to the video frame number, as well as a total of 21 feature text annotations including pedestrian features, environmental information, and vehicle status, stored in XML format.
[0053] The specific names and variable calibrations of the feature text annotations in dataset a are shown in Table 1 - 1:
[0054] Table 1 - 1 Basic variable and calibration description table
[0055]
[0056] S2: Extract pedestrian skeleton point data: As Figure 1As shown, use the OpenPose algorithm to extract the coordinates and recognition confidence of 25 skeleton points of all recognizable pedestrian bodies appearing in the video [x bi , y bi , α bi , and store them in JSON format corresponding to the video sequence number and frame sequence number respectively, denoted as dataset b.
[0057] S3: Target matching between multi-source data: Since the annotation storage format in the basic dataset is different from the skeleton point data format, and there is also a lack of correspondence between the pedestrian IDs in the two data sources, specific rules need to be formulated to achieve the matching of the same target between different data sources. The specific rule is that the relative error between the coordinate differences of the midpoint M: (x a , y a ) of the pedestrian identification box in dataset a and the centroid G: (x b , y b ) of the recognizable skeleton point frame in dataset b is less than Δ% (specifically 20% in this example).
[0058] Among them, according to the pedestrian identification box coordinates [x a1 , y a1 , x a2 , y a2 given in the basic dataset, the pedestrian position reference point in dataset a is calibrated as the midpoint of the identification box, denoted as M: (x a , y a ):
[0059] At the same time, the length of the identification box is calibrated as L = |y a1 - y a2 |, and the width is calibrated as W = |x a1 - x a2 |;
[0060] According to the recognized skeleton point data, the pedestrian position reference point in dataset b is calibrated as the centroid of the recognizable skeleton point frame, denoted as G: (x b , y b ),
[0061] Among them, δ(x bi , y bi ) is an indicator function for evaluating the recognizability of the i-th skeleton point. When the coordinate values of (x bi , y bi ) are specific pixel values rather than 0, the indicator function counts as 1;
[0062] Finally, obtain the relative errors δ x and δy : According to the rule, when both δ x and δ y are less than 20%, the matching result is successful.
[0063] When matching, select a specific target in dataset a, traverse and search all recognizable pedestrian data in dataset b, and find the corresponding target that meets the conditions to successfully match. This rule can be applied to the target matching between any datasets with pedestrian identification box coordinates and containing skeleton information, and has strong universality.
[0064] S4: Extraction of dynamic and static pose features: Based on the skeleton point data extracted in S2, develop new dynamic and static pose features that are helpful for judging the pedestrian's intention to cross the road, including the four body orientations of the pedestrian relative to the vehicle, namely front, back, left, and right, the smaller value of the knee bending angles of the left and right legs, and the relative ratios of the step length and shoulder width relative to the entire body width from the driver's perspective, a total of 7 static variables. The calculation methods and value ranges of the static pose variables are shown in Table 1-2:
[0065] Table 1-2 Description of the calculation methods and value ranges of static pose variables
[0066]
[0067]
[0068] The frame rate of the video material in the basic dataset is v frames per second (specifically 30 frames per second in this example). The fixed interval selected during the upsampling in S4 is (specifically 15 frames in this example), that is, starting from the frame of the video when the pedestrian first appears, take a sample data every 0.5 seconds for the processing of dynamic pose features, and compare the front-back relationship and body orientation of the pedestrian's left and right legs in the j-th frame with those in the th frame (specifically the (j - 15)-th frame in this example) and the th frame (specifically the (j + 15)-th frame in this example). If there is a change, then mark whether there is an action of alternating the front and back of the legs, whether there is a switch in the front-back orientation, and whether there is a switch in the left-right orientation as "yes". The calculation methods and value ranges of the dynamic pose variables are shown in Table 1-3:
[0069] Table 1-3 Description of the calculation methods and value ranges of dynamic pose variables
[0070] Variable Calculation method Value range Legs alternate The front - back relationship of the pedestrian's left and right legs changes within 1 second 0 or 1 Front - back orientation change The front - back orientation of the pedestrian changes within 1 second 0 or 1 Left - right orientation change The left - right orientation of the pedestrian changes within 1 second 0 or 1
[0071] S5: Train a machine learning model for prediction and explain the specific impact of each variable: Using whether the pedestrian crosses the street marked in the basic dataset as the dependent variable, build a pedestrian crossing intention prediction model using XGBoost. Divide the sample set into a training set and a test set at a ratio of 7:3. Apply the GridSearchCV function in Python on the training set, select 5-fold cross-validation to debug the model, and obtain the parameter combination with the best effect. The specific parameters involved in optimizing the XGBoost model are: the number of trees n_estimators, the minimum loss reduction value gamma, the maximum depth of the tree max_depth, the minimum number of samples per child min_child_weight, the sampling ratio of the tree subsample, the column sampling ratio of the sample colsample_bytree, the regularization parameter lambda (L2 regularization) and alpha (L1 regularization), and the learning rate learning_rate. Apply the parameters adjusted on the training set to the test set to verify the model metrics and examine the prediction accuracy.
[0072] Use the SHAP package in Python to sort the variables that play a key role in the model according to their importance, and give the positive or negative impact of each variable on the global result. In addition, by applying the LIME package in Python, obtain the importance ranking and qualitative explanation of the variables that have a significant impact locally for specific instances in the sample. Combine the two to analyze the specific impact mechanism of influencing pedestrian crossing decisions, providing a reference basis for the scenario test of autonomous vehicles.
[0073] The above description of the embodiments is to enable those of ordinary skill in the art to understand and use the invention. It is obvious that those skilled in the art can easily make various modifications to these embodiments and apply the general principles described herein to other embodiments without creative efforts. Therefore, the present invention is not limited to the above embodiments, and the improvements and modifications made by those skilled in the art without departing from the scope of the present invention should be within the protection scope of the present invention.
Claims
1. A method for predicting pedestrian crossing intention by integrating skeleton information in a natural driving environment, characterized by: The prediction steps are as follows: S1. Prepare a natural driving basic dataset: The dataset used contains video clips shot from the perspective of a dashcam in a natural driving environment, and is divided into 5-15 second segments to better focus on the interaction with a certain pedestrian or certain groups of pedestrians. The position coordinates of the pedestrian identification box and the feature text annotations are extracted based on the video clips to form dataset a; S2. Extract pedestrian skeleton point data: Use the OpenPose algorithm to extract the coordinates and recognition confidence of 25 skeleton points of all identifiable pedestrians appearing in the video to form data set b; S3, target matching between multi-source data: Based on the data set a collected in step S1 and the data set b collected in step S2, the same target between different data sources of data set a and data set b is matched; S4, dynamic and static posture feature extraction: based on the skeleton point data extracted in step S2, dynamic and static posture features for judging the pedestrian's intention to cross the street are established; The static posture features include: the front, back, left, and right body orientation features of the person relative to the vehicle. the smaller value of the left leg knee flexion angle and the right leg knee flexion angle; The relative ratios of step length and shoulder width to the entire body width from the driver's perspective; The dynamic posture feature includes: extracting the action variable of whether the pedestrian's legs alternate forward and backward within 1 second after sampling frames at fixed intervals; Is there a switch variable for front and back orientation? Whether there is a switch variable for left and right orientation; S5: Train the machine learning model to make predictions and explain the specific impact of each variable: Take the prediction result of whether the pedestrians marked in data set a cross the street as the dependent variable, use XGBoost to build a pedestrian crossing intention prediction model, apply the grid search method on the training set to debug the model, and get the best parameter combination. The model effect was then verified on the test set. The machine learning interpretable technology SHAP was used to rank the importance of each variable and give a global explanation of the specific impact of each variable. LIME was used to give a local explanation of the impact of each variable in a specific instance. The two explanations were combined to analyze the influencing mechanism behind pedestrian crossing decisions.
2. The method for predicting pedestrian crossing intention by integrating skeleton information in a natural driving environment according to claim 1, characterized in that: Dataset a also contains the vertex coordinates of the pedestrian identification box corresponding to the video frame number and feature text annotations, which are stored in XML format; The vertex coordinates of the target pedestrian identification box are recorded as,x a1 ,y a1 ,x a2 ,y a2 -; x a1 is the horizontal coordinate pixel value of the vertex of the lower left corner of the identification box, y a1 is the vertical coordinate pixel value of the lower left corner of the identification box, x a2 is the horizontal coordinate pixel value of the top right corner of the identification box, y a2 The vertical coordinate pixel value of the upper right corner of the identification box.
3. The method for predicting pedestrian crossing intention by integrating skeleton information in a natural driving environment according to claim 1, characterized in that: The confidence in step S2 is recorded as,x bi ,y bi ,α bi -; where i = 1, 2, ... 24, x bi is the horizontal coordinate pixel value of the i-th bone point, y bi is the vertical coordinate pixel value of the i-th bone point, α bi is the recognition confidence of the i-th bone point; The 25 skeleton points are numbered 0-24, and the corresponding joint points are defined as: 0 nose, 1 base of neck, 2 right shoulder, 3 right elbow, 4 right wrist, 5 left shoulder, 6 left elbow, 7 left wrist, 8 midpoint of hip, 9 right end point of hip, 10 right knee, 11 right ankle, 12 left end point of hip, 13 left knee, 14 left ankle, 15 right eye, 16 left eye, 17 right ear, 18 left ear, 19 left big toe, 20 left little toe, 21 left heel, 22 right big toe, 23 right little toe, 24 right heel.
4. The method for predicting pedestrian crossing intention by integrating skeleton information in a natural driving environment according to claim 1, 2 or 3, characterized in that: The feature text annotations in S1 include pedestrian features, environmental information, vehicle status and variable content corresponding to the video frame sequence number, as follows: 1) Pedestrian characteristics are: Crossing the street, variable content: yes, no; Pedestrian response, the variable content is: moving forward unimpeded, walking fast, walking slowly; Age, the variable content is: children, youth, middle-aged, and elderly; Gender, the variable content is: male, female; Moving direction, the variable contents are: perpendicular to the curb, parallel to the road; Backpack, the variable content is: yes, no; Using a mobile phone, the variable content is: yes, no; Carrying other obvious items, the variable content is: yes, no; Pushing a stroller, the variable contents are: yes, no; Riding a non-motorized vehicle, the variable content is: yes, no; The specific number of people around: 1, 2, 3, 4, 5, 6; 2) Environmental information: The location of the scene, the variable content is: street, parking lot, garage; For pedestrian crossing, the variable contents are: yes, no; Pay attention to the pedestrian sign, the variable content is: yes, no; The stop sign has variable contents: yes, no; Number of lanes: 1, 2, 3, 4, 5, 6, Located at the intersection, the variable content is: yes, no; Channelization facilities, the variable content is: yes, no; Signal control facility, variable content: yes, no; The direction of the traffic flow in which the vehicle is located. The variable content is: one-way street, two-way street; 3) Vehicle status: stopped, slow driving, fast driving, deceleration, acceleration.
5. The method for predicting pedestrian crossing intention by integrating skeleton information in a natural driving environment according to claim 1, 2 or 3, characterized in that: In step S3, the same target between different data sources of dataset a and dataset b is matched; the relative error between the coordinate difference of the pedestrian position reference point in different data sources is less than Δ%. The specific steps are as follows: 1) According to the pedestrian identification frame coordinates given in the natural driving basic dataset, x a1 ,y a1 ,x a2 ,y a2 -; The reference point of the pedestrian position in the calibration data set a is the midpoint of the identification frame, denoted by M:(x a ,y a ) The length of the calibration mark frame is L = |y a1 -y a2 |, width is W = |x a1 -x a2 |; 2) According to the identified skeleton point data, the pedestrian position reference point in the calibration data set b is the center of gravity of the identifiable skeleton point frame, denoted as G:(x b ,y b ), in, δ(x bi ,y bi ) is the indicator function for evaluating the identifiability of the i-th skeleton point. When (x bi ,y bi ) is a specific pixel value instead of 0, the indicator function count is 1; 3) Finally, the relative error δ between the coordinate differences of the pedestrian position reference points in different data sources is obtained x and δ y :
6. The method for predicting pedestrian crossing intention by integrating skeleton information in a natural driving environment according to claim 1, 2 or 3, characterized in that: The static posture variables in S4 are calibrated by the following method: 1) According to the joint points after the skeleton points are marked, the pedestrian's face is in the picture when the pedestrian's body is facing forward and the recognition confidence is greater than 0.5; that is, the recognition confidence of the nose, right eye, and left eye is greater than 0.5; Back: Any one of the recognition confidences of the nose, right eye, and left eye is less than or equal to 0.5; Left side: the recognition confidence of left shoulder, left elbow, left wrist, left end of hip, left knee, left ankle, left eye, left ear, left big toe, left little toe and left foot are all greater than 0.5; Right side: the recognition confidence of the right shoulder, right elbow, right wrist, right end of the hip, right knee, right ankle, right eye, right ear, right big toe, right little toe, and right heel are all greater than 0.5; 2) Knee bending angle: Obtain the knee bending angles of the pedestrian’s left and right legs respectively and take the smaller value as the value of this variable; Right knee angle: Left knee angle: 3) The step ratio is: 4) Shoulder width ratio is: 。 7. The method for predicting pedestrian crossing intention by integrating skeleton information in a natural driving environment according to claim 1, characterized in that: In step S4, The frame rate of the video material in dataset a is v frames / second, and the fixed interval selected when expanding the sample in step S4 is set to The front-to-back relationship of the left and right legs and the body orientation of the pedestrian in the jth frame are respectively compared with the Frame and Compare and judge the situations in the frames; If there is a change, whether there is an alternating forward and backward movement of the legs, whether there is a switching of the front and back directions, and whether there is a switching of the left and right directions are marked as: yes.
8. The method for predicting pedestrian crossing intention by integrating skeleton information in a natural driving environment according to claim 1, The ratio of training set to test set in S5 is 7:
3. In the process of grid search using the GridSearchCV function in Python, 5-fold cross validation is selected to determine the optimal parameter combination to calibrate the XGBoost model results and obtain the final pedestrian crossing intention prediction model. The parameters that need to be adjusted are: The number of trees n_estimators, the minimum loss reduction value gamma, the maximum depth of the tree max_depth, the minimum number of subsamples min_child_weight, the sampling ratio of the tree subsample, the column sampling ratio of the sample colsample_bytree, the regularization parameters lambda and alpha, and the learning rate learning_rate. The parameters adjusted on the training set are applied to the test set to verify the model indicators and examine the prediction accuracy.