A road pedestrian traffic risk prediction method for vehicle active safety

By combining variable resolution image patch mapping and spatial suppression enhancement strategy modules, a pedestrian crossing risk prediction model is constructed, which solves the prediction error problem of existing methods in complex traffic scenarios and achieves efficient and accurate pedestrian crossing risk prediction.

CN122637129APending Publication Date: 2026-08-25CHANGAN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610985958.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-03
Publication Date
2026-08-25

AI Technical Summary

Technical Problem

Existing pedestrian behavior prediction methods suffer from problems such as strong information dependence, high computational complexity, insufficient target attention capability, and weak spatiotemporal modeling capability in complex traffic scenarios, especially with large prediction errors in occluded scenarios and distant targets.

Method used

A pedestrian traffic risk prediction model is constructed using a variable resolution image patch mapping module and a spatial suppression enhancement strategy module. Prediction is completed end-to-end within a single model. Supervised training is performed by combining a multimodal large language model and the parallel branches of the spatial suppression enhancement strategy module, and natural language prediction results are output.

Benefits of technology

It reduces system complexity and deployment costs, improves the ability to finely perceive small targets and coordinate coding of global scenes, enhances the accuracy and generalization of predictions, and is suitable for practical application scenarios with limited on-board computing resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122637129A_ABST
    Figure CN122637129A_ABST
Patent Text Reader

Abstract

A road pedestrian traffic risk prediction method for vehicle active safety, comprising: acquiring a video frame sequence and a first frame pedestrian bounding box; constructing a prediction model comprising a variable resolution image block mapping module, a multi-modal large language model and a spatial suppression enhancement strategy module connected in series; the spatial suppression enhancement strategy module comprises a positioning branch, a prediction branch, a classification branch and a generation branch connected in parallel; training with the video frame sequence and the first frame pedestrian bounding box as input and the crossing behavior label as output. In the inference stage, the generation branch directly outputs the natural language prediction result of "crossing" or "not crossing", discards the traditional multi-model cascade pipeline, does not require external equipment, only requires the original video frame sequence and the first frame pedestrian bounding box, and completes the prediction end to end in the model to directly output the natural language prediction result of "crossing" or "not crossing", greatly reduces the system complexity and deployment cost, and is suitable for practical application scenarios with limited vehicle-mounted computing resources.
Need to check novelty before this filing date? Find Prior Art

Claims

1. A method for constructing a road pedestrian crossing risk prediction model for vehicle active safety, characterized in that, Includes the following steps: S1, Obtain the dataset, which includes a video frame sequence and corresponding first-frame pedestrian bounding box annotations and street crossing behavior labels. Each frame of the video frame sequence is a panoramic input image. The number of observed frames in the video frame sequence is The prediction time domain is ; Set the coarse-grained block size and the fine-grained block size, where the coarse-grained block size is greater than the fine-grained block size; S2, Construct a pedestrian crossing risk prediction model, which includes a variable resolution image patch mapping module, a multimodal large language model, and a spatial suppression enhancement strategy module connected in series. The spatial suppression enhancement strategy module includes parallel localization branches, prediction branches, classification branches, and generation branches; S3, using the video frame sequence and the pedestrian bounding box of the first frame as input, and the pedestrian crossing behavior label as output, train the pedestrian crossing risk prediction model constructed in S2 to obtain the trained pedestrian crossing risk prediction model. The specific training steps include: S31, the video frame sequence and the first frame pedestrian bounding box are input into the pedestrian traffic risk prediction model. After forward propagation calculation by the variable resolution image patch mapping module, the multimodal large language model and the spatial suppression enhancement strategy module, a multimodal fusion representation is obtained. The multimodal fusion representation is then input into the spatial suppression enhancement strategy module, which outputs historical bounding box regression values ​​from the localization branch, future location prediction values ​​from the prediction branch, street crossing behavior classification prediction values ​​from the classification branch, and street crossing behavior prediction text from the generation branch. S32, calculate the Smooth L1 loss for the historical bounding box regression value output by the localization branch and the pedestrian bounding box annotations in the dataset to obtain the localization loss; calculate the Smooth L1 loss for the future location prediction value output by the prediction branch and the pedestrian bounding box annotations in the dataset to obtain the prediction loss; calculate the cross-entropy loss for the street crossing behavior classification prediction value output by the classification branch and the street crossing behavior label to obtain the classification loss; calculate the language modeling loss for the street crossing behavior prediction text output by the generation branch and the street crossing behavior label to obtain the generation loss. S33, the positioning loss, prediction loss, classification loss and generation loss are weighted and summed to obtain the total loss, and the parameters of the pedestrian crossing risk prediction model are updated through backpropagation with the minimization of the total loss as the optimization direction. S34. Repeat steps S3.1 to S3.3 until convergence, to obtain the trained pedestrian crossing risk prediction model.

2. The method for constructing a road pedestrian crossing risk prediction model for vehicle active safety as described in claim 1, characterized in that, The variable resolution image patch mapping module in S31 includes a cascaded geometric alignment stage and a content-aware stage. The specific processing procedure of the geometric alignment stage is as follows: S311, a pedestrian bounding box is given for the first frame of the video frame sequence. According to a fixed ratio For the pedestrian bounding box Expand the bounding box to obtain the expanded bounding box. and the expanded bounding box Normalized to a square region ; in, This represents the coordinates of the top-left corner of the pedestrian bounding box. and These represent the width and height of the pedestrian bounding box, respectively. S312, according to the square region crop pedestrian partial images from each frame of panoramic input images ; respectively partition the panoramic input images and pedestrian partial images using the coarse-grained block size and the fine-grained block size set in S1 to obtain a coarse-grained image block grid and a fine-grained image block grid, where the position index of each image block in the coarse-grained image block grid is , where 0 ≤ i < H, 0 ≤ j < W, and H and W are the numbers of image blocks in the height direction and the width direction of the coarse-grained image block grid respectively; then perform feature extraction on the coarse-grained image block grid and the fine-grained image block grid to obtain a coarse-grained feature map and a fine-grained feature map ; ; S313, in the coarse-grained feature map The above determines whether it falls into the square area. The set of position indices of all coarse-grained image patches within the set, and the calculation of each coarse-grained image patch in the set of position indices based on the size of the coarse-grained patch. In the panoramic input image The pixel boundaries are then projected onto the pedestrian local image. In the image space, relative coordinates are obtained, and these relative coordinates are converted into the fine-grained feature map according to the fine-grained block size. The range of image block indexes in the image; S314, from the coarse-grained feature map Extract the corresponding position of each coarse-grained image patch from the location index set. The coarse-grained features at the location, from the fine-grained feature map Extract all fine-grained features corresponding to the index range of the image patch, then aggregate all fine-grained features using average pooling to obtain aligned fine-grained features; replace the coarse-grained feature map with the aligned fine-grained features. Corresponding position The coarse-grained features at the location are used to obtain the pooling feature tensor. ; Construct a binary overlay mask M, where the mask value for the position corresponding to the coarse-grained image block in the location index set is set to 1, and the mask value for the position corresponding to the coarse-grained image block outside the location index set is set to 0.

3. The method for constructing a road pedestrian crossing risk prediction model for vehicle active safety as described in claim 2, characterized in that, The specific processing steps of the content-aware stage are as follows: Based on the pooling feature tensor Each position The attention weight map is calculated by taking the L2 norm of the feature vectors and mapping them using the Sigmoid activation function. The attention weight map Spatial dimensions and the coarse-grained feature map The space dimensions are consistent. Indicates position Attention weight map for fine-grained features; Determine whether each location belongs to the pedestrian area based on the binary overlay mask M: For The position of the attention weight map As weights, the pooling feature tensor With coarse-grained feature map Perform a weighted summation to obtain the characteristics of that position; for The location is directly taken from the coarse-grained feature map. As a characteristic of this location; Features from all locations together constitute the panoramic fusion feature. .

4. The method for constructing a road pedestrian traffic risk prediction model for vehicle active safety as described in claim 3, characterized in that, The multimodal large language model in S31 includes a visual encoder, an adapter, and a backbone encoder connected in series along the data flow direction. The specific processing procedure is as follows: The panoramic fusion feature The visual words are processed sequentially by the visual encoder and the adapter to obtain visual lexical units; the visual lexical units are concatenated with user prompt text lexical units to obtain a multimodal lexical sequence; the multimodal lexical sequence is input into the backbone encoder for multimodal encoding to obtain a multimodal fusion representation.

5. The method for constructing a road pedestrian crossing risk prediction model for vehicle active safety as described in claim 4, characterized in that, S311 by a fixed ratio For the pedestrian bounding box The expansion is specifically achieved by maintaining the pedestrian bounding box. The center position remains unchanged, and the width and height Expanding outwards in all directions The expanded bounding box is obtained by multiplying the bounding box by 1. ; in, The value is 1.

5.

6. The method for constructing a road pedestrian crossing risk prediction model for vehicle active safety as described in claim 4, characterized in that, In S313, each coarse-grained image block in the location index set is calculated based on the coarse-grained block size. In the panoramic input image The pixel boundaries in the image are defined as follows: ; ; ; ; in, The size of the coarse-grained block; Input the panoramic image width, Input the panoramic image The height is in pixels; , , and These are the current coarse-grained image patches in the panoramic input image. The left boundary horizontal coordinate, the upper boundary vertical coordinate, the right boundary horizontal coordinate, and the lower boundary vertical coordinate.

7. The method for constructing a road pedestrian crossing risk prediction model for vehicle active safety as described in claim 4, characterized in that, The total loss described in S33 is: in, For the generation loss, For the classification loss, For the positioning loss, The predicted loss; , , and These are the weighting coefficients for each loss.

8. A method for predicting pedestrian traffic risks on roads for active vehicle safety, characterized in that, Includes the following steps: The video frame sequence to be predicted and the first frame pedestrian bounding box are obtained, and the video frame sequence to be predicted and the first frame pedestrian bounding box are input into the pedestrian crossing risk prediction model constructed by the method described in any one of claims 1 to 7. After being processed by the variable resolution image patch mapping module, the multimodal large language model and the spatial suppression enhancement strategy module in sequence, the street crossing behavior prediction text is output by the generation branch to obtain the natural language prediction result of "crossing the street" or "not crossing the street".

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the road pedestrian crossing risk prediction method for vehicle active safety as described in claim 8.

10. A computer program product, characterized in that, It includes a computer program / instruction, which, when executed by a processor, implements the road pedestrian crossing risk prediction method for vehicle active safety as described in claim 8.