Automatic driving system based on pedestrian motivation and control method
By using an autonomous driving system based on pedestrian motivation, combined with video processing and pedestrian intent reasoning, the vehicle speed is dynamically adjusted, solving the problem of insufficient accuracy in pedestrian detection and improving traffic safety and efficiency.
Patent Information
- Application Number
- CN202511257095.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-03
- Publication Date
- 2025-11-14
Smart Images

Figure CN120942366A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of autonomous driving technology, and in particular to an autonomous driving system and control method based on pedestrian motivation. Background Technology
[0002] With increasing attention being paid to autonomous driving technology, more and more scholars are conducting extensive research on self-driving cars, and the technology is becoming increasingly mature. Currently, autonomous vehicles have begun to be tested on some open roads, testing shared roads with pedestrians and other road users. Because pedestrians are an important part of road users, and they often do not follow traffic rules when crossing roads, their behavior is highly unpredictable.
[0003] Due to factors such as differences in camera angle, changes in lighting, cluttered backgrounds, changes in pedestrian posture, and occlusion interference, the accuracy of pedestrian detection needs to be improved. Most existing human-vehicle interaction algorithms only consider the pedestrian detection results, that is, they take braking or steering actions to avoid collisions when a pedestrian is detected in front, which affects traffic efficiency.
[0004] To address the above issues, this application presents an autonomous driving system and control method based on pedestrian motivation. Summary of the Invention
[0005] The technical problem this invention aims to solve is to address the shortcomings of existing technologies by providing an autonomous driving system and control method based on pedestrian motivation. The system includes a data acquisition module, a pedestrian detection module, a pedestrian intention reasoning module, and a vehicle control module. The data acquisition module acquires video sequences and environmental information of the traffic intersection when the autonomous vehicle approaches an unsignalized intersection. The pedestrian detection module detects pedestrians through video processing and pedestrian recognition strategies and outputs inter-frame images of pedestrians. The pedestrian intention reasoning module classifies pedestrians' crossing motivations based on their trajectory and posture information. The vehicle control module dynamically adjusts the vehicle speed according to the pedestrian motivation category and automatically brakes when necessary to ensure pedestrians safely cross the intersection. The system achieves adaptive yielding to pedestrians at unsignalized intersections, effectively improving traffic safety and the adaptability of autonomous driving.
[0006] To achieve the above objectives, the present invention provides the following technical solution:
[0007] An autonomous driving system based on pedestrian motivation, the system comprising:
[0008] The system comprises a data acquisition module, a pedestrian detection module, a pedestrian intent reasoning module, and a vehicle control module.
[0009] The data acquisition module is used to determine whether the intersection is a traffic light-free intersection by using a camera device when the intelligent driving vehicle approaches the intersection. If it is a traffic light-free intersection, it acquires a video sequence of the traffic intersection.
[0010] The pedestrian detection module is configured with a pedestrian recognition strategy, which is used to detect pedestrians based on the video sequence of the traffic intersection, determine whether there are pedestrians crossing the intersection, and output pedestrian inter-frame images if there are pedestrians crossing the intersection.
[0011] The pedestrian intention reasoning module is configured with a motivation prediction strategy, which is used to predict the pedestrian's intention to cross the street based on the pedestrian inter-frame images and output the pedestrian motivation category.
[0012] The vehicle control module is used to adaptively control the intelligent driving vehicle based on the pedestrian's motivation category.
[0013] The data acquisition module includes:
[0014] The intersection recognition unit is used to determine whether an intelligent driving vehicle has entered an intersection without traffic lights, based on a geographic information system and the vehicle's current location.
[0015] The environmental sensing unit is used to acquire information about the surrounding environment at traffic intersections.
[0016] The pedestrian detection module includes:
[0017] The video sequence processing unit is used to process the traffic intersection video sequence and output the traffic intersection inter-frame image.
[0018] The pedestrian recognition unit is used to determine whether there are pedestrians crossing the intersection based on the inter-frame images of the traffic intersection. If there are pedestrians crossing, it outputs the inter-frame images of the pedestrians.
[0019] The pedestrian recognition strategy includes video processing logic and pedestrian detection logic. The video processing logic is configured within the video sequence processing unit, and the pedestrian detection logic is configured within the pedestrian recognition unit.
[0020] The video processing logic includes:
[0021] Based on the video sequence of the traffic intersection captured by the camera equipment, the inter-frame images of the traffic intersection are acquired frame by frame, and differential denoising processing is performed on the inter-frame images of the traffic intersection.
[0022] The differentially denoised traffic intersection frame image is adjusted based on the surrounding environment information to obtain an optimized intersection image. The surrounding environment information includes ambient light and ambient weather.
[0023] Background subtraction is performed on the optimized intersection image, and the contour features of the optimized intersection image are calculated through morphological opening operations.
[0024] The pedestrian detection logic includes:
[0025] A pedestrian recognition network is constructed, and the contour features of the optimized intersection image are used as the input parameters of the pedestrian recognition network. The pedestrian recognition network is trained to determine whether there are pedestrians crossing the intersection. If there are pedestrians, the pedestrian inter-frame image is output.
[0026] The pedestrian recognition network includes:
[0027] The sampling layer is used to expand the planar dimension of the input parameters according to bilinear interpolation, transforming contour features of different levels and scales into a uniform scale;
[0028] Convolutional layers are used to expand the channels of the sampling layer's output and fuse the outputs of the sampling layer using a 1x1 convolutional fusion channel;
[0029] The detection layer is used to activate the output of the convolutional layer. The output of the activated convolutional layer is backpropagated according to the loss function. After the convergence condition is met, the inter-frame image of the pedestrian is output. The activation process includes planar dimension activation mapping and stereo dimension activation mapping.
[0030] The pedestrian intent reasoning module includes:
[0031] The pedestrian trajectory analysis unit is used to extract the trajectory information of pedestrians based on their positions in the inter-frame images of pedestrians.
[0032] A pedestrian pose recognition unit is used to analyze the pedestrian's body pose based on the inter-frame images of the pedestrian;
[0033] The motivation classification unit is used to classify the pedestrian's crossing motivation based on the trajectory information and body posture, and output the pedestrian's motivation category, wherein the motivation category includes preparing to cross the street, waiting to cross the street, and having no intention to cross the street;
[0034] The motivation prediction strategy includes information analysis logic and information integration logic. The information analysis logic is used to process the pedestrian inter-frame images and extract pedestrian features from the pedestrian inter-frame images. The information integration logic is used to fuse the pedestrian features and process the fused pedestrian features through a neural network algorithm to output the pedestrian's motivation category.
[0035] The information analysis logic is configured within the pedestrian trajectory analysis unit and the pedestrian posture recognition unit, and the information integration logic is configured within the motivation classification unit.
[0036] The information analysis logic includes:
[0037] The pixel displacement vector corresponding to the pedestrian pixel in the inter-frame image is calculated by optical flow method, and the trajectory information of the pedestrian is determined based on the pixel displacement vector.
[0038] The inter-frame images of pedestrians are skeletonized by morphological operations to obtain pedestrian skeleton maps. The pedestrian skeleton maps are then learned and optimized by adaptive graph convolution to obtain pedestrian orientation features and pedestrian dynamic features.
[0039] The pedestrian orientation features and pedestrian dynamic features are weighted and concatenated, and the pedestrian's body posture is calculated using the sigmoid activation function.
[0040] The information integration logic includes:
[0041] A pedestrian crossing intention prediction network is constructed, and the trajectory information and body posture are used as input parameters of the pedestrian crossing intention prediction network. The pedestrian crossing intention prediction network processes the input parameters and outputs the pedestrian's motivation category.
[0042] The pedestrian crossing intention prediction network includes:
[0043] The action classification layer calculates the similarity between the input parameters and the standard dataset of pedestrian crossing behavior by using a normalized embedded Gaussian function, and enhances and classifies the similarity by using an adjacency matrix and a mask to calculate the pedestrian motion vector.
[0044] The intent reasoning layer constructs a pedestrian walking intent reasoning module, takes the pedestrian motion vector as the input parameter of the pedestrian walking intent reasoning module, transforms the input parameter into real-time pedestrian description features at the intersection through semantic conversion, and infers from the real-time pedestrian description features to obtain pedestrian behavior intent information;
[0045] The prediction layer constructs a pedestrian future trajectory prediction module, using the pedestrian behavioral intention information as the input parameter of the pedestrian future trajectory prediction module. The input parameter is trained through a weight matrix and a bias matrix, and the category of pedestrian motivation judgment is output.
[0046] The vehicle control module includes:
[0047] The vehicle status acquisition unit is used to acquire the current status information of the intelligent driving vehicle in real time.
[0048] The pedestrian motivation analysis unit is used to receive the pedestrian motivation category output by the pedestrian intention reasoning module, and further analyze and judge the pedestrian's intention to cross the street, and determine whether the pedestrian has the intention to cross immediately or just stay on the side of the road to wait.
[0049] The vehicle speed adjustment unit dynamically adjusts the vehicle's speed based on the output of the pedestrian motivation analysis unit.
[0050] An automatic braking unit is used to perform automatic braking in emergency situations.
[0051] An autonomous driving control method based on pedestrian motivation, the method comprising:
[0052] When an intelligent driving vehicle approaches a traffic intersection, it uses camera equipment to monitor the intersection in real time and obtain video sequences of the intersection.
[0053] Based on the pedestrian recognition strategy at the intersection, pedestrian detection is performed on the video sequence of the traffic intersection to determine whether there are pedestrians crossing the intersection. If there are pedestrians, the inter-frame images of pedestrians at the intersection are acquired and the vehicle is controlled to slow down.
[0054] A pedestrian motivation prediction network is constructed. The inter-frame images of pedestrians at the intersection are used as input parameters of the pedestrian motivation prediction network. The pedestrian motivation prediction network infers the input parameters to predict the pedestrians' intention to cross the street and outputs the category of the pedestrian motivation judgment.
[0055] The intelligent driving vehicle is controlled based on the category determined by the pedestrian's motivation.
[0056] Compared with the prior art, the beneficial effects of the present invention are:
[0057] 1. This invention addresses the feature extraction stage of pedestrian detection methods by using the contour features of inter-frame images at traffic intersections as input parameters for a pedestrian recognition network. It provides geometric feature information of the contours through morphological and background subtraction methods, thereby enhancing the descriptive ability of the contour features and improving the accuracy of pedestrian detection.
[0058] 2. This invention identifies and analyzes the dynamic behavior patterns of pedestrians, predicts their future movement trajectories, and makes corresponding decisions and plans based on these predictions. By considering the motives and intentions of pedestrians, collisions with pedestrians can be avoided, and the safety and efficiency of the traffic system can be improved. Attached Figure Description
[0059] Other features, objects, and advantages of the invention will become more apparent from the following detailed description of non-limiting embodiments with reference to the accompanying drawings:
[0060] Figure 1 This is a flowchart illustrating an autonomous driving intersection parking control method based on pedestrian motivation, according to Embodiment 1 of the present invention.
[0061] Figure 2 This is a flowchart of differential denoising processing for inter-frame images of traffic intersections in Embodiment 1 of the present invention;
[0062] Figure 3 This is a schematic diagram of the environmental conditioning process in Embodiment 1 of the present invention;
[0063] Figure 4 This is a diagram of the pedestrian recognition network structure in Embodiment 1 of the present invention;
[0064] Figure 5 This is a block diagram of an autonomous driving intersection parking control system based on pedestrian motivation, according to Embodiment 2 of the present invention. Detailed Implementation
[0065] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments.
[0066] The term "embodiment" as used herein means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a mutually exclusive, independent, or alternative embodiment. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.
[0067] Example 1:
[0068] Please see Figure 1 The present invention provides an embodiment of an automated driving control method based on pedestrian motivation, the specific steps of which are as follows:
[0069] S1: When the intelligent driving vehicle approaches an intersection, the vehicle's camera equipment is activated to identify whether the current intersection is an intersection without traffic lights;
[0070] S2: If it is determined to be an intersection without traffic lights, collect the traffic intersection video sequence;
[0071] S3: Process the traffic intersection video sequence according to the pedestrian recognition strategy, detect whether there are pedestrians in the intersection area, and if pedestrians are detected passing through, extract the pedestrian inter-frame image;
[0072] S4: Process the pedestrian inter-frame images according to the motivation prediction strategy and output the pedestrian motivation category;
[0073] S5: Adaptive control of the intelligent driving vehicle based on the pedestrian's motivation category;
[0074] When an autonomous vehicle approaches an intersection, it first activates its onboard camera system, including high-resolution cameras mounted at the front or sides of the vehicle, to collect real-time environmental information about the intersection. The triggering conditions for the camera system are typically determined by the vehicle's driving status and pre-set geographic information system (GIS) data. When the vehicle reaches the intersection area, the camera system automatically activates and begins capturing image data of the intersection ahead. In determining whether it is an intersection without traffic lights, it calls upon pre-loaded road network map data and combines it with the visual information from the camera system to determine if the current intersection lacks traffic lights. This determination process also includes image recognition algorithms that extract and match features from the intersection images captured by the cameras, detecting the presence of traffic light characteristics such as their shape, location, and color. If no traffic lights are detected, the current intersection is determined to be an intersection without traffic lights.
[0075] The specific steps of the pedestrian recognition strategy are as follows:
[0076] S3.1: Based on the video sequence of the traffic intersection captured by the camera equipment, obtain the inter-frame images of the traffic intersection frame by frame, and perform differential noise reduction processing on the inter-frame images of the traffic intersection.
[0077] In this step, the inter-frame images of the traffic intersection video sequence are first acquired frame by frame by the camera device. In order to eliminate the interference caused by noise and environmental changes, the difference operation is performed on every two frames to extract dynamic target information and eliminate static background information. The principle of differential denoising is to eliminate static noise in the dynamic environment by calculating the pixel difference between the current frame and the previous frame. Through differential denoising, redundant information in the background can be effectively reduced, the motion characteristics of pedestrians can be enhanced, and the accuracy of pedestrian detection can be improved.
[0078] Please see Figure 2 The flowchart of differential denoising processing for inter-frame images of traffic intersections according to an embodiment of the present invention is as follows: The specific steps of differential denoising processing for inter-frame images of traffic intersections in S3.1 are as follows:
[0079] S3.1.1: Based on Gaussian difference filtering, feature information is extracted from the traffic intersection inter-frame image of the current frame and the previous frame to obtain a four-layer two-dimensional Gaussian difference filtered traffic intersection inter-frame image.
[0080] S3.1.2: The four-layer two-dimensional Gaussian difference filtered traffic intersection inter-frame image is divided into four directions. High-frequency noise is filtered out of the segmented four-layer two-dimensional Gaussian difference filtered traffic intersection inter-frame image. The four-way noise-reduced four-layer two-dimensional Gaussian difference filtered traffic intersection inter-frame images are fused to obtain the high-frequency noise-reduced four-layer two-dimensional Gaussian difference filtered traffic intersection inter-frame image.
[0081] S3.1.3: Fuse the two-dimensional Gaussian difference filtered traffic intersection inter-frame images after each layer of denoising to obtain the denoised traffic intersection inter-frame images;
[0082] S3.2: Adjust the inter-frame image of the traffic intersection after differential denoising according to the surrounding environment information to obtain an optimized intersection image, wherein the surrounding environment information includes ambient light and ambient weather;
[0083] In this step, the adjustment is mainly based on two important environmental factors: ambient light and ambient weather, ensuring effective pedestrian detection under different lighting and weather conditions. By adaptively adjusting the image's lighting compensation and weather condition correction, image optimization under complex environmental conditions is achieved. This technology not only maintains image clarity under different weather conditions but also ensures stable output of inter-frame images under extreme lighting conditions through complex environmental information inference and adaptive gain adjustment.
[0084] Please see Figure 3 The schematic diagram of the environment adjustment process for inter-frame images at traffic intersections according to an embodiment of the present invention, wherein the specific steps of S3.2 are as follows:
[0085] S3.2.1: Ambient light has a significant impact on the images captured by the camera equipment, especially in cases of excessively strong or weak light, making it difficult to guarantee recognition accuracy. By obtaining the current ambient light intensity, gamma compensation is performed on the light intensity of the image.
[0086] S3.2.2: After adjusting the image brightness, the image is enhanced by a contrast enhancement method based on a bilateral filter. Its core is to preserve edge information while smoothing non-edge areas.
[0087] S3.2.3: Under different weather conditions, images will be affected to varying degrees. For example, image clarity decreases under hazy weather, and water droplets and snowflakes may cause image blurring on rainy or snowy days. Therefore, the influence of weather can be corrected by atmospheric scattering models.
[0088] S3.3: Perform background subtraction on the optimized intersection image, and calculate the contour features of the optimized intersection image through morphological opening operations. The formula for calculating the contour features of the optimized intersection image is as follows:
[0089]
[0090] in, Let represent the contour features of the optimized intersection image, ... This represents the weight value of the i-th Gaussian distribution. Let X represent the probability density function of a Gaussian distribution, and let X represent the optimized intersection image. This represents the mean of the Gaussian distribution in the optimized intersection image. This represents the Gaussian distribution variance of the optimized intersection image. Indicates corrosion operation. This indicates a dilation operation, where T represents the average pixel density of the optimized intersection image.
[0091] In this step, the core of background subtraction lies in separating the foreground target, i.e., pedestrians, from each frame of the video sequence, while the background is ignored through modeling. The specific steps are as follows:
[0092] S3.3.1: A Gaussian Mixture Model (GMM) is used to model multi-frame intersection images. This model can dynamically adapt to changes in ambient light, making it particularly suitable for video image processing in complex scenes. Each pixel is treated as a data point, and a Gaussian distribution is used for fitting, thereby continuously updating the background model over a long period to ensure the accuracy of the background.
[0093] S3.3.2: By comparing the pixel distribution in the current frame with that in the background model, it is determined which pixels belong to the foreground and which belong to the background. This can eliminate the influence of noise in dynamic environments and obtain more accurate foreground targets. For each frame, the pixel values of the background model are compared with those of the current frame. Pixels with large differences are considered foreground, while those with small differences are identified as background. In this way, fixed or unchanging objects (such as buildings, road markings, etc.) can be removed, while information of moving targets is preserved. Through background subtraction, the efficiency of target detection can be significantly improved because the parts of the background that do not need to be processed have been eliminated, reducing the burden of subsequent calculations.
[0094] S3.3.3: Morphological processing is performed on the foreground image after background subtraction, specifically including the opening operation. The opening operation is a combination of erosion and dilation, which aims to remove noise from the image, smooth object edges, and retain larger foreground targets. The erosion operation scans the image with small structuring elements to gradually erode noise points and small isolated targets in the foreground, effectively filtering out interference information other than pedestrians. The dilation operation restores the edges of foreground targets to maintain the complete outline of pedestrians.
[0095] S3.4: Construct a pedestrian recognition network and use the contour features of the optimized intersection image as the input parameters of the pedestrian recognition network. Train the pedestrian recognition network to determine whether there are pedestrians crossing the intersection. If there are pedestrians, output the pedestrian inter-frame image.
[0096] Please see Figure 4 The pedestrian recognition network structure diagram of this invention includes:
[0097] The sampling layer is used to describe contour features using elliptical Fourier descriptors. It expands the planar dimension of the input parameters using bilinear interpolation, transforming contour features of different levels and scales into a uniform scale.
[0098] Convolutional layers are used to expand the channels of the sampling layer's output and fuse the outputs of the sampling layer using a 1x1 convolutional fusion channel;
[0099] The detection layer is used to activate the output of the convolutional layer. The error is backpropagated on the activated output of the convolutional layer according to the loss function. After the convergence condition is met, the inter-frame image of pedestrians at the intersection is output. The activation process includes planar dimension activation mapping and stereo dimension activation mapping.
[0100] The formula for calculating the loss function is as follows:
[0101]
[0102] Where Loss represents the loss function. This represents the coordinate loss of the convolutional layer output after activation processing. Indicates the coordinate loss weights. This represents the confidence loss of the convolutional layer output after activation processing. Indicates the confidence loss weight. This represents the class loss of the convolutional layer output after activation processing. This represents the category loss weights, where the coordinate loss of the convolutional layer output after activation is calculated using the following formula:
[0103]
[0104] Where m represents a single pedestrian detection box, B represents the total number of pedestrian detection boxes in the pedestrian recognition network, S represents the plane dimension of the convolutional layer output, and n represents a single coordinate in the plane output by the convolutional layer. This represents the pedestrian detection discriminant function. This represents the x-coordinate of the plane output of the convolutional layer. This represents the average horizontal coordinate of the output plane of the convolutional layer. This represents the ordinate of the plane output of the convolutional layer. The average ordinate of the plane output of the convolutional layer, where V represents the 3D dimension of the convolutional layer output, and j represents a single 3D coordinate of the convolutional layer output. Represents the vertical coordinates of the convolutional layer output. This represents the average vertical coordinate of the output plane of the convolutional layer;
[0105] The specific steps of the motivation prediction strategy are as follows:
[0106] S4.1: Calculate the pixel displacement vector corresponding to the pedestrian pixel in the inter-frame image of the pedestrian using the optical flow method, and determine the trajectory information of the pedestrian based on the pixel displacement vector;
[0107] S4.2: The inter-frame images of pedestrians are skeletonized through morphological operations to obtain a pedestrian skeleton map. The pedestrian skeleton map is then learned and optimized using adaptive graph convolution to obtain pedestrian orientation features and pedestrian dynamic features.
[0108] S4.3: The pedestrian orientation features and pedestrian dynamic features are weighted and concatenated, and the pedestrian's body posture is calculated using the sigmoid activation function;
[0109] S4.4: Construct a pedestrian crossing intention prediction network, using the trajectory information and body posture as input parameters of the pedestrian crossing intention prediction network, and process the input parameters through the pedestrian crossing intention prediction network to output the pedestrian's motivation category;
[0110] The pedestrian crossing intention prediction network includes:
[0111] In the action classification layer, a normalized embedded Gaussian function is applied to calculate the similarity between pedestrian posture features and the standard pedestrian posture dataset. The similarity is enhanced and classified through adjacency matrix and mask to obtain pedestrian motion vectors. The pedestrian motion vectors include direction vectors, position vectors, and action vectors. The direction vectors include pedestrian lateral movement, pedestrian longitudinal movement, and pedestrian stationary state. The position vectors include the middle of the intersection, the sides of the intersection, and away from the intersection. The action vectors include walking state and standing state.
[0112] The intent reasoning layer constructs a pedestrian walking intent reasoning module. The pedestrian motion vector is used as the input parameter of the pedestrian walking intent reasoning module. The input parameter is transformed into real-time pedestrian description features at the intersection through semantic conversion. The real-time pedestrian description features are used for reasoning to obtain pedestrian behavior intent information. The pedestrian behavior intent information includes pedestrian turning behavior intent information, pedestrian detour behavior intent information, pedestrian crossing behavior intent information, and pedestrian waiting behavior intent information.
[0113] The prediction layer constructs a pedestrian future trajectory prediction module, which takes the pedestrian behavioral intention information as the input parameter of the pedestrian future trajectory prediction module, and performs inference on the input parameter through the weight matrix and bias matrix to output the category of pedestrian motivation judgment;
[0114] The specific steps of S5 are as follows:
[0115] S5.1: If it is determined that the pedestrian has an explicit intention to cross the street, stop the vehicle based on the distance of the intelligent driving vehicle from the intersection and the vehicle speed; explicit intention to cross the street refers to preparing to cross the street.
[0116] S5.2: If the pedestrian's implicit intention to cross the street is determined, the vehicle will decelerate based on the distance between the intelligent driving vehicle and the intersection and the vehicle speed; implicit intention to cross the street refers to waiting to cross the street.
[0117] S5.3: If it is determined that the pedestrian has no intention of crossing the street, control the intelligent driving vehicle to maintain a constant speed through the intersection.
[0118] Example 2:
[0119] Please see Figure 5 The present invention provides an embodiment: an autonomous driving system based on pedestrian motivation, the system comprising:
[0120] The system comprises a data acquisition module, a pedestrian detection module, a pedestrian intent reasoning module, and a vehicle control module.
[0121] The data acquisition module is used to determine whether the intersection is a traffic light-free intersection by using a camera device when the intelligent driving vehicle approaches the intersection. If it is a traffic light-free intersection, it acquires a video sequence of the traffic intersection.
[0122] The pedestrian detection module is configured with a pedestrian recognition strategy, which is used to detect pedestrians based on the video sequence of the traffic intersection, determine whether there are pedestrians crossing the intersection, and output pedestrian inter-frame images if there are pedestrians crossing the intersection.
[0123] The pedestrian intention reasoning module is configured with a motivation prediction strategy, which is used to predict the pedestrian's intention to cross the street based on the pedestrian inter-frame images and output the pedestrian motivation category.
[0124] The vehicle control module is used to adaptively control the intelligent driving vehicle based on the pedestrian's motivation category.
[0125] The data acquisition module includes:
[0126] The intersection recognition unit is used to determine whether an intelligent driving vehicle has entered an intersection without traffic lights, based on a geographic information system and the vehicle's current location.
[0127] The environmental sensing unit is used to acquire information about the surrounding environment at traffic intersections.
[0128] The pedestrian detection module includes:
[0129] The video sequence processing unit is used to process the traffic intersection video sequence and output the traffic intersection inter-frame image.
[0130] The pedestrian recognition unit is used to determine whether there are pedestrians crossing the intersection based on the inter-frame images of the traffic intersection. If there are pedestrians crossing, it outputs the inter-frame images of the pedestrians.
[0131] The pedestrian recognition strategy includes video processing logic and pedestrian detection logic. The video processing logic is configured within the video sequence processing unit, and the pedestrian detection logic is configured within the pedestrian recognition unit.
[0132] The pedestrian intent reasoning module includes:
[0133] The pedestrian trajectory analysis unit is used to extract the trajectory information of pedestrians based on their positions in the inter-frame images of pedestrians.
[0134] A pedestrian pose recognition unit is used to analyze the pedestrian's body pose based on the inter-frame images of the pedestrian;
[0135] The motivation classification unit is used to classify the pedestrian's crossing motivation based on the trajectory information and body posture, and output the pedestrian's motivation category, wherein the motivation category includes preparing to cross the street, waiting to cross the street, and having no intention to cross the street;
[0136] The motivation prediction strategy includes information analysis logic and information integration logic. The information analysis logic is used to process the pedestrian inter-frame images and extract pedestrian features from the pedestrian inter-frame images. The information integration logic is used to fuse the pedestrian features and process the fused pedestrian features through a neural network algorithm to output the pedestrian's motivation category.
[0137] The information analysis logic is configured within the pedestrian trajectory analysis unit and the pedestrian posture recognition unit, and the information integration logic is configured within the motivation classification unit.
[0138] The vehicle control module includes:
[0139] The vehicle status acquisition unit is used to acquire the current status information of the intelligent driving vehicle in real time.
[0140] The pedestrian motivation analysis unit is used to receive the pedestrian motivation category output by the pedestrian intention reasoning module, and further analyze and judge the pedestrian's intention to cross the street, and determine whether the pedestrian has the intention to cross immediately or just stay on the side of the road to wait.
[0141] The vehicle speed adjustment unit dynamically adjusts the vehicle's speed based on the output of the pedestrian motivation analysis unit.
[0142] An automatic braking unit is used to perform automatic braking in emergency situations.
[0143] Although embodiments of the present invention have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those skilled in the art can make changes, modifications, substitutions and variations to the above embodiments within the scope of the present invention.
Claims
1. An autonomous driving system based on pedestrian motivation, characterized in that, The system includes: The system comprises a data acquisition module, a pedestrian detection module, a pedestrian intent reasoning module, and a vehicle control module. The data acquisition module is used to determine whether the intersection is a traffic light-free intersection by using a camera device when the intelligent driving vehicle approaches the intersection. If it is a traffic light-free intersection, it acquires a video sequence of the traffic intersection. The pedestrian detection module is configured with a pedestrian recognition strategy, which is used to detect pedestrians based on the video sequence of the traffic intersection, determine whether there are pedestrians crossing the intersection, and output pedestrian inter-frame images if there are pedestrians crossing the intersection. The pedestrian intention reasoning module is configured with a motivation prediction strategy, which is used to predict the pedestrian's intention to cross the street based on the pedestrian inter-frame images and output the pedestrian motivation category. The vehicle control module is used to adaptively control the intelligent driving vehicle based on the pedestrian's motivation category.
2. The autonomous driving system based on pedestrian motivation according to claim 1, characterized in that, The data acquisition module includes: The intersection recognition unit is used to determine whether an intelligent driving vehicle has entered an intersection without traffic lights, based on a geographic information system and the vehicle's current location. The environmental sensing unit is used to acquire information about the surrounding environment at traffic intersections.
3. The autonomous driving system based on pedestrian motivation according to claim 2, characterized in that, The pedestrian detection module includes: The video sequence processing unit is used to process the traffic intersection video sequence and output the traffic intersection inter-frame image. The pedestrian recognition unit is used to determine whether there are pedestrians crossing the intersection based on the inter-frame images of the traffic intersection. If there are pedestrians crossing, it outputs the inter-frame images of the pedestrians. The pedestrian recognition strategy includes video processing logic and pedestrian detection logic. The video processing logic is configured within the video sequence processing unit, and the pedestrian detection logic is configured within the pedestrian recognition unit.
4. The autonomous driving system based on pedestrian motivation according to claim 3, characterized in that, The video processing logic includes: Based on the video sequence of the traffic intersection captured by the camera equipment, the inter-frame images of the traffic intersection are acquired frame by frame, and differential denoising processing is performed on the inter-frame images of the traffic intersection. The differentially denoised traffic intersection frame image is adjusted based on the surrounding environment information to obtain an optimized intersection image. The surrounding environment information includes ambient light and ambient weather. Background subtraction is performed on the optimized intersection image, and the contour features of the optimized intersection image are calculated through morphological opening operations.
5. The autonomous driving system based on pedestrian motivation according to claim 4, characterized in that, The pedestrian detection logic includes: A pedestrian recognition network is constructed, and the contour features of the optimized intersection image are used as the input parameters of the pedestrian recognition network. The pedestrian recognition network is trained to determine whether there are pedestrians crossing the intersection. If there are pedestrians, the pedestrian inter-frame image is output. The pedestrian recognition network includes: The sampling layer is used to expand the planar dimension of the input parameters according to bilinear interpolation, transforming contour features of different levels and scales into a uniform scale; Convolutional layers are used to expand the channels of the sampling layer's output and fuse the outputs of the sampling layer using a 1x1 convolutional fusion channel; The detection layer is used to activate the output of the convolutional layer. The output of the activated convolutional layer is backpropagated according to the loss function. After the convergence condition is met, the inter-frame image of the pedestrian is output. The activation process includes planar dimension activation mapping and stereo dimension activation mapping.
6. The autonomous driving system based on pedestrian motivation according to claim 5, characterized in that, The pedestrian intent reasoning module includes: The pedestrian trajectory analysis unit is used to extract the trajectory information of pedestrians based on their positions in the inter-frame images of pedestrians. A pedestrian pose recognition unit is used to analyze the pedestrian's body pose based on the inter-frame images of the pedestrian; The motivation classification unit is used to classify the pedestrian's crossing motivation based on the trajectory information and body posture, and output the pedestrian's motivation category, wherein the motivation category includes preparing to cross the street, waiting to cross the street, and having no intention to cross the street; The motivation prediction strategy includes information analysis logic and information integration logic. The information analysis logic is used to process the pedestrian inter-frame images and extract pedestrian features from the pedestrian inter-frame images. The information integration logic is used to fuse the pedestrian features and process the fused pedestrian features through a neural network algorithm to output the pedestrian's motivation category. The information analysis logic is configured within the pedestrian trajectory analysis unit and the pedestrian posture recognition unit, and the information integration logic is configured within the motivation classification unit.
7. The autonomous driving system based on pedestrian motivation according to claim 6, characterized in that, The information analysis logic includes: The pixel displacement vector corresponding to the pedestrian pixel in the inter-frame image is calculated by optical flow method, and the trajectory information of the pedestrian is determined based on the pixel displacement vector. The inter-frame images of pedestrians are skeletonized by morphological operations to obtain pedestrian skeleton maps. The pedestrian skeleton maps are then learned and optimized by adaptive graph convolution to obtain pedestrian orientation features and pedestrian dynamic features. The pedestrian orientation features and pedestrian dynamic features are weighted and concatenated, and the pedestrian's body posture is calculated using the sigmoid activation function.
8. The autonomous driving system based on pedestrian motivation according to claim 7, characterized in that, The information integration logic includes: A pedestrian crossing intention prediction network is constructed, and the trajectory information and body posture are used as input parameters of the pedestrian crossing intention prediction network. The pedestrian crossing intention prediction network processes the input parameters and outputs the pedestrian's motivation category. The pedestrian crossing intention prediction network includes: The action classification layer calculates the similarity between the input parameters and the standard dataset of pedestrian crossing behavior by using a normalized embedded Gaussian function, and enhances and classifies the similarity by using an adjacency matrix and a mask to calculate the pedestrian motion vector. The intent reasoning layer constructs a pedestrian walking intent reasoning module, takes the pedestrian motion vector as the input parameter of the pedestrian walking intent reasoning module, transforms the input parameter into real-time pedestrian description features at the intersection through semantic conversion, and infers from the real-time pedestrian description features to obtain pedestrian behavior intent information; The prediction layer constructs a pedestrian future trajectory prediction module, which takes the pedestrian behavioral intention information as the input parameter of the pedestrian future trajectory prediction module, and performs inference on the input parameter through the weight matrix and bias matrix to output the category of pedestrian motivation judgment.
9. The autonomous driving system based on pedestrian motivation according to claim 8, characterized in that, The vehicle control module includes: The vehicle status acquisition unit is used to acquire the current status information of the intelligent driving vehicle in real time. The pedestrian motivation analysis unit is used to receive the pedestrian motivation category output by the pedestrian intention reasoning module, and further analyze and judge the pedestrian's intention to cross the street, and determine whether the pedestrian has the intention to cross immediately or just stay on the side of the road to wait. The vehicle speed adjustment unit dynamically adjusts the vehicle's speed based on the output of the pedestrian motivation analysis unit. An automatic braking unit is used to perform automatic braking in emergency situations.
10. An automated driving control method based on pedestrian motivation, characterized in that, The method includes: When an intelligent driving vehicle approaches a traffic intersection, it uses camera equipment to monitor the intersection in real time and obtain video sequences of the intersection. Based on the pedestrian recognition strategy at the intersection, pedestrian detection is performed on the video sequence of the traffic intersection to determine whether there are pedestrians crossing the intersection. If there are pedestrians, the inter-frame images of pedestrians at the intersection are acquired and the vehicle is controlled to slow down. A pedestrian motivation prediction network is constructed. The inter-frame images of pedestrians at the intersection are used as input parameters of the pedestrian motivation prediction network. The pedestrian motivation prediction network is used to train the input parameters to predict the pedestrians' intention to cross the street and output the category of the pedestrian motivation judgment. The intelligent driving vehicle is controlled based on the category determined by the pedestrian's motivation.
Citation Information
Cited By
Vehicle automatic driving control method fusing belief updating and drift diffusion model
CN121573012A
Vehicle automatic driving control method fusing belief update and drift diffusion model
CN121573012B