Text-guided cascaded network for pedestrian crossing intention prediction method and system

By using a text-guided cascaded network approach and leveraging vehicle-mounted cameras and a large language model to generate behavioral descriptions, the problem of inaccurate pedestrian intent expression in existing technologies is solved, achieving more accurate prediction of pedestrian crossing intentions.

CN120913180BActive Publication Date: 2026-01-02WUHAN UNIV OF TECH
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202511406677.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-29
Publication Date
2026-01-02
Estimated Expiration
2045-09-29

AI Technical Summary

Technical Problem

Existing multimodal methods are insufficient in capturing fine-grained pedestrian behavior, and suffer from high computational overhead, data noise, and information sparsity, resulting in inaccurate expression of pedestrian dynamic intentions.

Method used

A text-guided cascaded network approach is adopted, which uses an on-board camera to detect pedestrian coordinate sequences, combines a large language model to generate a description of pedestrian crossing behavior, extracts and aligns features through a coordinate sequence decoder and encoder, and uses a cosine similarity function to obtain the pedestrian intent prediction results.

Benefits of technology

It significantly reduces trajectory prediction error, improves the model's ability to express pedestrian dynamic intentions, takes into account both global motion trends and local subtle displacements, and enhances the robustness and generalization ability of feature extraction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120913180B_ABST
    Figure CN120913180B_ABST
Patent Text Reader

Abstract

The application provides a text-guided cascaded network pedestrian crossing intention prediction method and system, relates to the technical field of pedestrian behavior analysis in an intelligent transportation system, and comprises the following steps: collecting road video by using a vehicle-mounted camera, detecting pedestrians in the road video based on a YOLO algorithm, and extracting pedestrian coordinate sequences with continuous frames; sequentially performing dimension expansion and feature extraction on the pedestrian coordinate sequences to obtain predicted pedestrian coordinates corresponding to the pedestrian coordinate sequences; splicing the pedestrian coordinate sequences and the predicted pedestrian coordinates to obtain complete coordinate sequence features; generating a crossing behavior description corresponding to a structured prompt word based on a structured prompt word generated by a large language model and a pedestrian crossing video, encoding the crossing behavior description into behavior description features; and performing feature alignment on the complete coordinate sequence features and the behavior description features by using a cosine similarity function to obtain a pedestrian intention prediction result. The application helps to improve the expression ability of pedestrian dynamic intention.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of pedestrian behavior analysis in intelligent transportation systems, and in particular to a pedestrian crossing intention prediction method and system based on a text-guided cascade network. BACKGROUND

[0002] In recent years, autonomous driving technology and intelligent transportation systems have made significant progress, and the environmental perception, decision planning and path control of vehicles have been continuously improved. These advances have greatly improved road safety and traffic efficiency; however, predicting pedestrian behavior remains a major challenge in autonomous driving and intelligent transportation. However, existing multi-modal methods still have shortcomings in capturing fine-grained behavior, and often face problems such as excessive computational overhead, data noise, and information sparsity.

[0003] Chinese Patent No. CN111860269B discloses a multi-feature fusion series RNN structure and a pedestrian prediction method, which includes an information collection module, an information processing module, a series GRU module, a full connection layer module, an activation function module and a prediction module. The information collection module collects video images of pedestrians and the surrounding environment and the vehicle speed when the vehicle is driving in different roads and crowd density environments; the information processing module processes the collected data to generate a data set; each GRU in the series GRU module processes different information in the data set and the input of the hidden state of the previous stage GRU in series, and fuses and calculates the different information; the full connection layer module integrates the above multi-dimensional matrix to obtain a one-dimensional vector; the excitation function module processes the one-dimensional vector information; and the prediction module obtains the prediction result of the pedestrian trajectory. However, the above scheme only sends the video image frame and the vehicle speed signal into the GRU series network, and simply relies on the numerical fusion of visual and speed data, resulting in inaccurate expression of the dynamic intention of the pedestrian. Therefore, it is necessary to provide a pedestrian crossing intention prediction method and system based on a text-guided cascade network to improve the expression ability of the dynamic intention of the pedestrian. SUMMARY

[0004] Therefore, the present application provides a pedestrian crossing intention prediction method and system based on a text-guided cascade network, which significantly reduces the trajectory prediction error by using a coordinate sequence decoder to perform short-term prediction on the future trajectory of the pedestrian, and generates a crossing behavior description with the help of a large language model to make up for the semantic blind area of single coordinate information, thereby effectively improving the expression ability of the dynamic intention of the pedestrian.

[0005] The present application provides a pedestrian crossing intention prediction method based on a text-guided cascade network, which comprises:

[0006] A vehicle-mounted camera is used to collect a road video, and a pedestrian in the road video is detected based on a YOLO algorithm to extract a pedestrian coordinate sequence with continuous frames;

[0007] The pedestrian coordinate sequence is sequentially subjected to dimension expansion and feature extraction, and high-dimensional features extracted from the pedestrian coordinate sequence are input into a decoder to obtain predicted pedestrian coordinates corresponding to the pedestrian coordinate sequence;

[0008] The pedestrian coordinate sequence and the predicted pedestrian coordinates are spliced, and the spliced pedestrian coordinate sequence is input into an encoder to obtain complete coordinate sequence features;

[0009] Based on a structured prompt word generated by a large language model and a pedestrian crossing video, a crossing behavior description corresponding to the structured prompt word is generated, and the crossing behavior description is encoded into behavior description features;

[0010] The complete coordinate sequence features and the behavior description features are aligned by a cosine similarity function, and a pedestrian intention prediction result is obtained according to a loss function and a classification strategy corresponding to the aligned complete coordinate sequence features.

[0011] On the basis of the above technical solutions, preferably, the obtaining of the predicted pedestrian coordinates corresponding to the pedestrian coordinate sequence specifically includes:

[0012] The pedestrian coordinate sequence is subjected to dimension expansion based on a multilayer perception machine to obtain an upgraded pedestrian coordinate sequence;

[0013] The upgraded pedestrian coordinate sequence is subjected to feature extraction based on an encoder to obtain high-dimensional features corresponding to the pedestrian coordinate sequence;

[0014] The high-dimensional features are input into a decoder to predict predicted pedestrian coordinates corresponding to a moment when a pedestrian is about to take an action.

[0015] On the basis of the above technical solutions, preferably, the splicing of the pedestrian coordinate sequence and the predicted pedestrian coordinates, and the input of the spliced pedestrian coordinate sequence into an encoder to obtain complete coordinate sequence features specifically includes:

[0016] The predicted pedestrian coordinates and the pedestrian coordinate sequence are spliced to obtain a spliced predicted coordinate sequence;

[0017] The spliced predicted coordinate sequence is constrained based on an L2 regular loss function, and the spliced predicted coordinate sequence is sequentially subjected to dimension expansion and feature extraction again to obtain complete coordinate sequence features corresponding to the spliced predicted coordinate sequence.

[0018] Further preferably, the structured prompt word generated by the large language model specifically comprises:

[0019] With the driving scene setting as the context, a role-based prompt word is constructed based on the large language model, and a reasoning standard for judging the pedestrian crossing behavior is generated, wherein the reasoning standard comprises the pedestrian behavior and body language, the pedestrian position and the place, the traffic signal and the road condition, the speed and motion of the surrounding vehicles, and the environment and time factors;

[0020] The reasoning standard is analyzed and selected by semantics to extract an intention index and construct a structured prompt word.

[0021] Further preferably, the generation of the crossing behavior description corresponding to the structured prompt word specifically comprises:

[0022] The structured prompt word and the pedestrian crossing video are input into a video understanding model to extract a text feature vector and a visual feature sequence, respectively;

[0023] The text feature vector and the visual feature sequence are fused to obtain a pedestrian behavior description.

[0024] Further preferably, the pedestrian behavior and body language comprises one or more of the gaze direction, the body orientation, the walking speed, and the gesture action, the pedestrian position and the place comprises one or more of the zebra crossing or intersection, the proximity to the curb, and the middle of the road, the traffic signal and the road condition comprises one or more of the pedestrian signal, the vehicle signal, the traffic density, and the road width and isolation facilities, the speed and motion of the surrounding vehicles comprises one or more of the other vehicle deceleration, the other vehicle parking, and the traffic gap, and the environment and time factors comprises one or more of the school area and residential area, the bus station or parked vehicle, the night, the poor visibility, and the weather condition.

[0025] Further preferably, the encoder is an encoder stacked by 8 layers of BiMamba, and the decoder is a decoder composed of multiple layers of perception mechanism.

[0026] In a second aspect of the present application, a text-guided cascaded network pedestrian crossing intention prediction system is provided, which comprises a cascaded network module, a behavior description generation module, and a central perception classification module, wherein,

[0027] The cascade network module is used for collecting road video by using a vehicle-mounted camera, detecting pedestrians in the road video based on a YOLO algorithm, extracting a pedestrian coordinate sequence with continuous frames, sequentially performing dimension expansion and feature extraction on the pedestrian coordinate sequence, inputting high-dimensional features extracted from the pedestrian coordinate sequence into a decoder to obtain predicted pedestrian coordinates corresponding to the pedestrian coordinate sequence, splicing the pedestrian coordinate sequence and the predicted pedestrian coordinates, and inputting the spliced pedestrian coordinate sequence into an encoder to obtain complete coordinate sequence features.

[0028] The behavior description generation module is used for generating a crossing behavior description corresponding to the structured prompt word based on the structured prompt word generated by the large language model and the pedestrian crossing video, and encoding the crossing behavior description into behavior description features.

[0029] The center perception classification module is used for performing feature alignment on the complete coordinate sequence features and the behavior description features through a cosine similarity function, and obtaining a pedestrian intention prediction result according to a loss function and a classification strategy corresponding to the aligned complete coordinate sequence features.

[0030] In a third aspect of the present application, an electronic device is provided, which includes a processor, a memory, a user interface and a network interface, the memory is used to store instructions, the user interface and the network interface are used to communicate with other devices, and the processor is used to execute the instructions stored in the memory.

[0031] In a fourth aspect of the present application, a computer readable storage medium is provided, which stores a computer program, and the computer program is executed by a processor to implement the steps of a text-guided cascade network pedestrian crossing intention prediction method.

[0032] The text-guided cascade network pedestrian crossing intention prediction method and system provided by the present application have the following beneficial effects relative to the prior art:

[0033] (1) The future trajectory of the pedestrian is short-term predicted by using the coordinate sequence decoder, and then spliced with the historical trajectory to input into the encoder for whole segment spatiotemporal feature extraction, so as to fully capture the historical and future continuity and time dependence, which can take into account the global motion trend and local subtle displacement, significantly reduce the trajectory prediction error, and introduce semantic priori into the pure numerical trajectory data by means of the crossing behavior description generated by the large language model, make up for the semantic blind area of single coordinate information, map the spatiotemporal features and semantic features to a consistent distribution by using the classical similarity measure, compatible with multiple source information, suppress the interference of single modal noise on prediction, combine the coordinate prediction loss and the classification loss, the network can simultaneously learn discrete intention and continuous trajectory, improve the adaptability to complex scenes, and thus effectively improve the expression ability of the model to the dynamic intention of the pedestrian.

[0034] (2) By splicing the historical coordinate sequence with the predicted coordinates, the network can simultaneously obtain the continuous trajectory information of the actions that have occurred and the actions that will occur of the pedestrian, fully retain the temporal coherence and spatial consistency, and introduce an L2 regular loss to impose a smoothing constraint on the complete trajectory sequence, suppress the mutation and noise at the splicing place, ensure the smooth transition of the feature distribution, and further improve the robustness and generalization ability of feature extraction. At the same time, the complete coordinate sequence after the constraint is dimensionally expanded, non-linear interaction is introduced, and then the spatial and temporal dependence is deeply extracted by the encoder, which can capture more delicate trajectory details and global motion trends, and enhance the feature expression ability. BRIEF DESCRIPTION OF DRAWINGS

[0035] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings needed to be used in the embodiments or prior art description. Obviously, the drawings in the following description are only some embodiments of the present application, and for those skilled in the art, other drawings can also be obtained without creative labor on the basis of these drawings.

[0036] Figure 1 A flowchart of a text-guided cascaded network pedestrian crossing intention prediction method provided by the present application is shown in the figure.

[0037] Figure 2 A network model diagram provided by the present application is shown in the figure.

[0038] Figure 3 A structure diagram of a real-time task scheduling system provided by the present application is shown in the figure.

[0039] Figure 4 A structure diagram of an electronic device provided by the present application is shown in the figure.

[0040] Explanation of reference numerals: 1, cascaded network pedestrian crossing intention prediction system; 11, cascaded network module; 12, behavior description generation module; 13, central perception classification module; 2, electronic device; 21, processor; 22, communication bus; 23, user interface; 24, network interface; 25, memory. DETAILED DESCRIPTION

[0041] The technical solutions in the embodiments of the present application will be described clearly and completely below in conjunction with the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, not all. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.

[0042] The present application discloses a text-guided cascaded network pedestrian crossing intention prediction method, as shown inFigure 1 The steps of the method include S1-S5.

[0043] Step S1, using a vehicle-mounted camera to collect road video, and detecting pedestrians in the road video based on a YOLO algorithm to extract a pedestrian coordinate sequence with consecutive frames.

[0044] In this step, a high-definition camera (e.g., 1920x1080 @30 fps) is installed on the front windshield or directly in front of the vehicle, ensuring that the field of view covers the entire lane in the driving direction of the vehicle. The camera has a GPS / IMU synchronization module, which timestamps and records the location information for each frame of video. Real-time recording is performed in MP4 (H.264 encoding) or ROS bag format, with a frame rate of 25-30 fps and low latency storage to an edge computing unit or SD card.

[0045] Start the front / rear camera synchronization recording, and perform necessary distortion correction and distortion model parameter calibration (checkerboard calibration method). Use a lightweight real-time target detection network (such as YOLOv5-s or YOLOv7 Tiny) to fine-tune on a public dataset (COCO, Cityscapes) containing pedestrian classes. Deploy the model on a vehicle-mounted GPU / NPU to ensure a single-frame inference delay of ≤30 ms.

[0046] Step S2, sequentially performing dimension expansion and feature extraction on the pedestrian coordinate sequence, and inputting the high-dimensional features extracted from the pedestrian coordinate sequence into a decoder to obtain predicted pedestrian coordinates corresponding to the pedestrian coordinate sequence.

[0047] Further, the encoder is an 8-layer BiMamba stacked encoder, and the decoder is a multi-layer perceptron decoder.

[0048] In this step, steps S21-S23 are also included.

[0049] Step S21, performing dimension expansion on the pedestrian coordinate sequence based on a multi-layer perceptron to obtain an upgraded pedestrian coordinate sequence.

[0050] In this step, the observed 16-frame coordinate sequence Dimension expansion enhances the expression ability of features:

[0051]

[0052] where MLP represents a multi-layer perceptron, represents the feature representation after upgrading through the MLP.

[0053] Step S22, performing feature extraction on the upgraded pedestrian coordinate sequence based on an encoder to obtain high-dimensional features corresponding to the pedestrian coordinate sequence.

[0054] In this step, the preliminary feature extraction of the coordinate sequence is performed using an encoder Encoder stacked by 8 layers of BiMamba to obtain high-level features :

[0055]

[0056] wherein, represents high-level features, Encoder represents an encoder stacked by 8 layers of BiMamba, is a feature representation after dimensionality increase by MLP.

[0057] Step S23, input the high-dimensional features into the decoder to predict the predicted pedestrian coordinates corresponding to the action moment of the pedestrian.

[0058] In this step, based on the high-level features of the observed coordinate sequence obtained in step S22 , the Decoder decoder is used to predict the coordinates of the action moment of the pedestrian .

[0059] In this embodiment, the MLP maps the original coordinate sequence to a high dimension, not only expands the feature space, but also introduces nonlinear interaction, so that the subsequent network can capture more rich displacement change rules and micro motion trends, and alleviate the short-term motion details easily submerged by noise under pure linear mapping, improve the prediction sensitivity. The encoder extracts the long and short term dependence information in the historical trajectory through multiple layers of time series depth. The decoder focuses on the coordinate prediction of the action moment of the pedestrian, which can effectively focus on the turning point and significantly reduce the average / endpoint error. From coordinate dimensionality increase to spatiotemporal coding to action moment prediction, the whole process supports unified loss function or multi-task loss joint training. The network can minimize the coordinate prediction error while providing more discriminative feature representation for subsequent intention classification, thereby improving the overall prediction quality.

[0060] Step S3, splice the pedestrian coordinate sequence and the predicted pedestrian coordinate, and input the spliced pedestrian coordinate sequence into the encoder to obtain the complete coordinate sequence feature.

[0061] In this step, steps S31-S32 are also included.

[0062] Step S31, splice the predicted pedestrian coordinate and the pedestrian coordinate sequence to obtain a spliced predicted coordinate sequence.

[0063] In this step, the predicted coordinate is combined with the observed coordinate sequence to obtain a complete coordinate sequence of 17 frames :

[0064]

[0065]

[0066] wherein, t denotes the moment when the pedestrian is about to take action, Decoder denotes a decoder composed of a multi-layer perception machine (MLP), h denotes a high-level feature, x denotes a complete 17-frame coordinate sequence, denotes splicing, and T denotes a 16-frame coordinate sequence that has been observed.

[0067] In step S32, the spliced predicted coordinate sequence is constrained based on an L2 regular loss function, and the spliced predicted coordinate sequence is sequentially subjected to dimension expansion and feature extraction again to obtain a complete coordinate sequence feature corresponding to the spliced predicted coordinate sequence.

[0068] In this step, the obtained predicted coordinate is constrained by using an L2 regular loss to ensure the accuracy of the prediction result:

[0069]

[0070] wherein, L denotes a regression loss, and B denotes the number of samples in a training batch, t denotes the moment when the pedestrian is about to take action, x denotes a real coordinate of the moment when the pedestrian is about to take action, denotes the square of the Euclidean distance.

[0071] Further, after obtaining the complete coordinate sequence , steps S21 and S22 are repeated to obtain features by dimension expansion, and then the features are extracted by using an encoder Encoder to obtain a complete coordinate sequence feature :

[0072]

[0073]

[0074] wherein, h denotes a feature after dimension expansion, and MLP denotes a multi-layer perception machine, x denotes a complete 17-frame coordinate sequence, x denotes a complete coordinate sequence feature, and Encoder denotes an encoder stacked by 8 BiMamba.

[0075] In this step, the network can simultaneously obtain the continuous trajectory information of the actions that have occurred and the actions that will occur of the pedestrian by splicing the historical coordinate sequence with the predicted coordinates, fully retaining the temporal coherence and spatial consistency. The L2 regular loss is introduced to apply a smoothing constraint on the complete trajectory sequence, suppress the mutation and noise at the splicing place, ensure the smooth transition of the feature distribution, and further improve the robustness and generalization ability of feature extraction. The complete coordinate sequence after constraint is dimensionally expanded, and the nonlinear interaction is introduced before the spatial and temporal dependence is extracted by the encoder, which can capture more delicate trajectory details and global motion trends, enhancing the feature expression ability. The obtained complete coordinate sequence features have higher discriminability and consistency when aligning with the behavior semantic features in the subsequent, making the cosine similarity matching more accurate, significantly improving the accuracy and confidence of pedestrian intention classification.

[0076] Step S4, based on the structured prompt words generated by the large language model and the pedestrian crossing video, a crossing behavior description corresponding to the structured prompt words is generated, and the crossing behavior description is encoded into a behavior description feature.

[0077] In this step, the structured prompt words generated by the large language model specifically include:

[0078] Taking the driving scene setting as the context, the role-based prompt words are constructed based on the large language model, and the reasoning standard for judging the pedestrian crossing behavior is generated, wherein the reasoning standard includes pedestrian behavior and body language, pedestrian position and location, traffic signal and road condition, speed and motion of surrounding vehicles, and environmental and time factors; through semantic analysis and feature selection on the reasoning standard, intention indicators are extracted and structured prompt words are constructed.

[0079] The pedestrian behavior and body language include one or more of gaze direction, body orientation, walking speed, and hand gestures, the pedestrian position and location include one or more of zebra crossing or intersection, close to the curb, and middle of the road, the traffic signal and road condition include one or more of pedestrian signal, vehicle signal, traffic density, and road width and isolation facilities, the speed and motion of surrounding vehicles include one or more of other vehicle deceleration, other vehicle parking, and traffic gap, and the environmental and time factors include one or more of school area and residential area, bus station or parked vehicle, night, poor visibility, and weather conditions.

[0080] In this embodiment, a real-world in-vehicle driving scenario is used as the context for the prompts, incorporating information such as road conditions, traffic lights, weather, and time. This ensures that the LLM (Local Mode Module) generates behavioral reasoning standards that are highly consistent with the actual application environment, improving semantic relevance and prediction reliability. By integrating various factors such as body language (lifting a foot, turning a head, gestures), pedestrian location and geographical elements (zebra crossings, sidewalks, curbs), traffic signals and road conditions, surrounding vehicle speeds and trajectories, ambient lighting, and time, fine-grained reasoning standards are formed, significantly improving the accuracy and robustness of judging pedestrian crossing intentions. Various reasoning standards and corresponding intention indicators are clearly listed, and the prompts themselves possess a clear logical hierarchy and traceability. When traffic rules, scene elements, or deployment cities change, only the prompt template needs to be adjusted or indicators added or deleted, quickly adapting to the new environment. The behavioral descriptions output by the structured prompts have high semantic concentration. When aligned with coordinate sequence features in the vector space, they significantly enhance the discriminative advantage of cosine similarity, accelerating convergence and improving the classification accuracy after multimodal fusion.

[0081] In the steps above, the driving scenario is set as the context, and role-based prompts are constructed to give GPT the identity of "the driver making a decision," thereby generating inference criteria for judging pedestrian crossing behavior. Please refer to the question-and-answer process below:

[0082] Question: If you are a professional driver, what factors would you use to determine whether a pedestrian is ready to cross the road?

[0083] Answer: As a professional driver, I assess the likelihood of a pedestrian crossing the road based on several factors, including visual cues, environmental context, and traffic conditions. Here are the key factors I consider:

[0084] Pedestrian behavior and body language: gaze direction, body orientation, walking speed, and hand gestures;

[0085] Pedestrian location and position: zebra crossing or intersection, near the curb, middle of the road;

[0086] Traffic signals and road conditions: pedestrian signals, vehicle signals, traffic density, road width and barriers;

[0087] The speed and movement of surrounding vehicles: other vehicles slowing down or stopping, traffic gaps;

[0088] Environmental and time factors: school area vs. residential area, bus stops or parked vehicles, nighttime / poor visibility, weather conditions.

[0089] Furthermore, a description of the street crossing behavior corresponding to the structured prompt words is generated, specifically including:

[0090] The structured prompt words and the pedestrian crossing video are input into a video understanding model to extract a text feature vector and a visual feature sequence respectively; the text feature vector and the visual feature sequence are fused to obtain a pedestrian behavior description.

[0091] In this embodiment, by semantic analysis and feature selection, key intention indicators are extracted, and a structured prompt is constructed by combining additional constraints:

[0092] Prompt: Observe the position of the pedestrian relative to the road, movement speed and direction, line of sight and head movement, and any gestures. Respond only when these behaviors are observable. If there are multiple people, only observe the first person you see. The answer should not exceed 77 words. Do not include words such as "video", "image", "picture", "frame" that are unrelated to behavior in the answer.

[0093] Further, the obtained prompt words and the pedestrian crossing video are input into a video understanding large model MiniCPM to obtain a pedestrian behavior description:

[0094] Description: The pedestrian is crossing the parking lot from left to right. The head is turning in the direction of the line of sight, as if he is talking or looking at something outside the picture. Due to the limited image quality, the gestures are not clearly visible.

[0095] In step S5, the complete coordinate sequence feature and the behavior description feature are aligned by a cosine similarity function, and the pedestrian intention prediction result is obtained according to the loss function corresponding to the aligned complete coordinate sequence feature and the classification strategy.

[0096] In this step, the complete coordinate sequence feature E and the behavior description feature T are aligned using cosine similarity to obtain the alignment loss

[0097]

[0098]

[0099] wherein COS(E, T) represents the cosine similarity between the complete coordinate sequence feature E and the behavior description feature T, and respectively represent the L2 norm of the coordinate sequence feature E and the behavior description feature T, and the alignment loss of the coordinate sequence feature E and the behavior description feature T.

[0100] Further, the aligned coordinate sequence feature is input into a center perception module to obtain a prediction result, and a center perception loss is used for constraint:

[0101]

[0102]

[0103]

[0104]

[0105]

[0106] where, represents the predicted class, represents the class index c that maximizes the maximum among all classes, W is the weight matrix of the classification layer, E represents the feature representation of the coordinate sequence, b represents the bias vector of the classification layer, represents the classification loss, represents the center loss, N represents the number of samples in a training batch, C represents the total number of classes, represents the predicted probability that the ith sample belongs to the cth class, represents the logits output (unnormalized value of neural network output) of the ith sample to the cth class, is the feature representation of sample i, represents the feature center of the class Y to which sample i belongs.

[0107] The final loss function is defined as follows:

[0108]

[0109] where, represents the total classification loss, represents the classification loss, represents the center loss, represents the coordinate prediction regression loss, represents the loss of aligning the coordinate sequence features with the behavior description features, and α, β, γ respectively represent the hyperparameters corresponding to the classification loss, the center loss and the coordinate prediction regression loss.

[0110] Please refer to Figure 2 , Figure 2A text-guided cascaded network pedestrian crossing intention prediction process is demonstrated. The "driver decision-making" role context input ChatGPT, and ChatGPT outputs structured prompts such as "observe the position of the pedestrian relative to the road, moving speed and direction, line of sight and head movement, and any gestures...". The above prompts and pedestrian crossing videos are simultaneously input into a multi-modal video understanding model MiniCPM, which generates natural language behavior descriptions (for example: "the pedestrian is crossing the parking lot from left to right. The head is turned in the direction of the line of sight, as if he is talking or looking at something outside the picture. Due to the limited image quality, the gestures are not very clear and visible.") through cross-modal fusion and self-attention.

[0111] The video is decomposed into consecutive frames T1...T N The trajectory encoder extracts high-dimensional features, and the decoder predicts the next coordinate. The original and predicted coordinates are spliced and input into the second stage encoder to obtain the complete coordinate sequence feature. The structured prompt is input into the text encoder (such as Long-CLIP text branch), and the text feature vector is extracted. The trajectory features and text features are aligned / fused through cosine similarity or cross-modal self-attention. The central perception classification module is input, and the pedestrian crossing intention prediction result (such as "not crossing the road") is output.

[0112] By introducing a text-guided mechanism in the pedestrian crossing intention prediction task, combining a large language model (LLM) to capture fine-grained pedestrian behavior features, and implementing a cascaded network to extend the coordinate sequence, the model's ability to express dynamic intentions is effectively improved. Specifically, first, a pre-trained large language model is used to generate multi-dimensional intention reasoning features, and high-quality prompts are constructed through manual screening. Then, a video understanding large model is introduced, and the prompts are used to generate detailed behavior descriptions for each crossing sample, thereby assisting the model in identifying complex behaviors such as looking and making phone calls. In parallel, a cascaded network is designed to predict and extend the observed coordinates to generate a more complete coordinate sequence, and the coordinate sequence and text behavior description are aligned to model the collaborative understanding of multi-modal information. Finally, a central perception classification module is proposed to improve the model's fine-grained discrimination ability in the coordinate sequence representation space. Experimental results on the public datasets JAAD and PIE show that the proposed method significantly outperforms existing methods in the intention prediction task, verifying its effectiveness and robustness in fine-grained behavior modeling and complex scene understanding.

[0113] In this embodiment, the future trajectory of the pedestrian is short-term predicted by using the coordinate sequence decoder, and then spliced with the historical trajectory to input the encoder for whole spatiotemporal feature extraction, which fully captures the continuity of history and future and the time dependence, can balance the global motion trend and the local subtle displacement, significantly reduces the trajectory prediction error, and generates the crossing behavior description with the help of the large language model, introduces the semantic prior into the pure numerical trajectory data, supplements the semantic blind area of single coordinate information, maps the spatiotemporal features and semantic features to a consistent distribution by using the classical similarity measure, compatible with multi-source information, suppresses the interference of single modal noise on prediction, combines the coordinate prediction loss and the classification loss, the network can learn discrete intention and continuous trajectory at the same time, improve the adaptability to complex scenes, and then effectively improve the expression ability of the model to the dynamic intention of the pedestrian.

[0114] Based on the above method, the embodiment of the application discloses a text-guided cascaded network pedestrian crossing intention prediction system, referring to Figure 3 , the cascaded network pedestrian crossing intention prediction system 1 includes a cascaded network module 11, a behavior description generation module 12, and a central perception classification module 13, wherein,

[0115] The cascaded network module 11 is used to collect road video using a vehicle-mounted camera, and detect pedestrians in the road video based on the YOLO algorithm to extract pedestrian coordinate sequences with continuous frames. The pedestrian coordinate sequences are sequentially dimensionally expanded and feature extracted, and the high-dimensional features extracted from the pedestrian coordinate sequences are input into a decoder to obtain predicted pedestrian coordinates corresponding to the pedestrian coordinate sequences. The pedestrian coordinate sequences and the predicted pedestrian coordinates are spliced, and the spliced pedestrian coordinate sequences are input into an encoder to obtain complete coordinate sequence features.

[0116] The behavior description generation module 12 is used to generate a crossing behavior description corresponding to the structured prompt word based on the structured prompt word generated by the large language model and the pedestrian crossing video, and encode the crossing behavior description into behavior description features.

[0117] The central perception classification module 13 is used to perform feature alignment on the complete coordinate sequence features and the behavior description features by using a cosine similarity function, and obtain a pedestrian intention prediction result according to a loss function and a classification strategy corresponding to the aligned complete coordinate sequence features.

[0118] It can be understood that the cascaded network module is an encoder-decoder architecture for extracting and extending coordinate features. The module first predicts the coordinates of the previous frame of the pedestrian action using the observed coordinate sequence, and then adds the observed coordinate sequence to form a complete coordinate sequence and extract coordinate sequence features.

[0119] The behavior description generation module is used to extract more fine-grained behavior characteristics of the pedestrian crossing. The module first uses a GPT large language model to generate reasoning criteria for judging the behavior of the pedestrian crossing, then takes key intention indicators, and combines additional constraints to construct structured prompts. The prompts and pedestrian crossing videos are input into the video understanding large language model MiniCPM to obtain pedestrian behavior descriptions. Finally, a pre-trained text encoder is used to extract features.

[0120] The center perception classification module combines center loss and classification loss to enhance the model's ability to distinguish coordinate sequences with more fine granularity.

[0121] In one example, the cascaded network module 11 is used to perform dimension expansion on the pedestrian coordinate sequence based on a multi-layer perception machine to obtain an upgraded pedestrian coordinate sequence; perform feature extraction on the upgraded pedestrian coordinate sequence based on an encoder to obtain high-dimensional features corresponding to the pedestrian coordinate sequence; and input the high-dimensional features into a decoder to predict the predicted pedestrian coordinates corresponding to the moment when the pedestrian is about to take action.

[0122] In one example, the cascaded network module 11 is used to splice the predicted pedestrian coordinates and the pedestrian coordinate sequence to obtain a spliced predicted coordinate sequence, constrain the spliced predicted coordinate sequence based on an L2 regular loss function, and perform dimension expansion and feature extraction on the spliced predicted coordinate sequence again to obtain complete coordinate sequence features corresponding to the spliced predicted coordinate sequence.

[0123] In one example, the structured prompts generated by the large language model include:

[0124] The driving scenario setting is used as the context, the role-based prompts are constructed based on the large language model, and the reasoning criteria for judging the behavior of the pedestrian crossing are generated, wherein the reasoning criteria include pedestrian behavior and body language, pedestrian position and location, traffic signals and road conditions, speed and motion of surrounding vehicles, and environmental and time factors.

[0125] The reasoning criteria are analyzed and selected by semantics to extract intention indicators and construct structured prompts.

[0126] In one example, the behavior description generation module 12 is used to input the structured prompts into a pre-trained text encoder to make the text encoder output a text feature vector; divide the N frames of video from the pedestrian crossing video into the visual encoder to extract the visual feature sequence corresponding to each frame in the pedestrian crossing video; input the text feature vector and the visual feature sequence into the video understanding model at the same time to obtain the output fused behavior fusion feature representation; and input the behavior fusion feature representation into the text decoder corresponding to the video understanding model to generate the crossing behavior description corresponding to the structured prompts.

[0127] In one example, the structured prompt word is input into a pre-trained text encoder to cause the text encoder to output a text feature vector;

[0128] N frames of video divided from the pedestrian crossing video are respectively input into a visual encoder to extract a visual feature sequence corresponding to each frame in the pedestrian crossing video;

[0129] The text feature vector and the visual feature sequence are simultaneously input into a video understanding model to obtain an output fused behavior fusion feature representation;

[0130] The behavior fusion feature representation is input into a text decoder corresponding to the video understanding model to generate a crossing behavior description corresponding to the structured prompt word.

[0131] In one example, the encoder is an encoder stacked by 8 layers of BiMamba, and the decoder is a decoder composed of multiple layers of perception mechanism.

[0132] Please refer to Figure 4 , an embodiment of the present application provides a structural schematic diagram of an electronic device. As shown in Figure 4 , the electronic device 2 can include at least one processor 21, at least one network interface 24, a user interface 23, a memory 25, and at least one communication bus 22.

[0133] The communication bus 22 is used to realize the connection and communication between the components.

[0134] The user interface 23 can include a display screen (Display), a camera (Camera), and optionally the user interface 23 can also include a standard wired interface and a wireless interface.

[0135] The network interface 24 can optionally include a standard wired interface and a wireless interface (such as a WI-FI interface).

[0136] The processor 21 can include one or more processing cores. The processor 21 connects various parts within the server through various interfaces and lines, performs various functions of the server and processes data by running or executing instructions, programs, code sets or instruction sets stored in the memory 25, and calling data stored in the memory 25. Alternatively, the processor 21 can be implemented in at least one of a hardware form of a digital signal processing (DSP), a field-programmable gate array (FPGA), and a programmable logic array (PLA). The processor 21 can be integrated with a combination of one or more of a central processing unit (CPU), a graphics processing unit (GPU), and a modem. Among them, the CPU mainly processes the operating system, user interface, and application programs; the GPU is responsible for rendering and drawing the content required to be displayed on the display screen; and the modem is used for processing wireless communication. It can be understood that the above-mentioned modem can also not be integrated into the processor 21, but can be realized by a separate chip.

[0137] The memory 25 can include a random access memory (RAM) and a read-only memory (ROM). Alternatively, the memory 25 includes a non-transitory computer-readable storage medium. The memory 25 can be used to store instructions, programs, codes, code sets or instruction sets. The memory 25 can include a program storage area and a data storage area, wherein the program storage area can store instructions for implementing an operating system, instructions for at least one function (such as a touch function, a sound playing function, an image playing function, etc.), instructions for implementing the above-mentioned various method embodiments, etc.; the data storage area can store data involved in the above-mentioned various method embodiments, etc. The memory 25 can alternatively be at least one storage device located away from the aforementioned processor 21. As shown in the figure, the memory 25 as a computer storage medium can include an operating system, a network communication module, a user interface module, and an application program of a text-guided cascaded network pedestrian crossing intention prediction method. Figure 4

[0138] In Figure 4 ​In the electronic device 2 shown, the user interface 23 is mainly used to provide an interface for the user to input, and obtain data input by the user; and the processor 21 can be used to invoke an application program of a text-guided cascaded network pedestrian crossing intention prediction method stored in the memory 25, which, when executed by one or more processors, causes the electronic device to perform one or more methods in the above embodiments.

[0139] A computer-readable storage medium stores instructions that, when executed by one or more processors, cause a computer to perform one or more methods in the above embodiments.

[0140] It should be noted that, for the foregoing method embodiments, in order to simply describe, they are all described as a series of action combinations, but those skilled in the art should know that the present application is not limited to the order of the actions described, because according to the present application, certain steps can be performed in other orders or at the same time. Secondly, those skilled in the art should know that the embodiments described in the specification all belong to preferred embodiments, and the actions and modules involved are not necessarily required by the present application.

[0141] In the above embodiments, the description of each embodiment has its own focus, and the parts not described in detail in a certain embodiment can be referred to the related description of other embodiments.

[0142] In the several embodiments provided by the present application, it should be understood that the disclosed device can be implemented in other ways. For example, the device embodiments described above are only schematic. The division of the units is only a logical function division. There can be another division manner for actual implementation, for example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the displayed or discussed mutual couplings or direct couplings or communication connections between different units, can be indirect couplings or communication connections through some interfaces, devices or units, and can be electrical or other forms.

[0143] The units described as separate components can or can not be physically separate, and the components displayed as units can or can not be physical units, that is, they can be located in one place, or can be distributed on a plurality of network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the embodiment scheme.

[0144] In addition, each functional unit in the embodiments of the present application can be integrated in one processing unit, or each unit can exist physically, or two or more units can be integrated in one unit. The above integrated unit can be realized in the form of hardware, or in the form of software functional units.

[0145] The integrated unit, if implemented in the form of a software function unit and sold or used as an independent product, can be stored in a computer readable memory. Based on such understanding, the technical solutions of the present application essentially or say the part that contributes to the prior art or the whole or part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a memory and includes a plurality of instructions for causing a computer device (which can be a personal computer, a server or a network device, etc.) to execute all or part of the steps of the embodiments of the present application. The aforementioned memory includes: a U disk, a mobile hard disk, a magnetic disk or an optical disk and various program code storage media.

[0146] The above is only an exemplary embodiment of the present disclosure, and cannot limit the scope of the present disclosure. That is, any equivalent changes and modifications made in accordance with the teachings of the present disclosure are still within the scope of the present disclosure. Other embodiments of the present disclosure will be readily apparent to those skilled in the art upon considering the specification and practicing the true disclosure. The present application is intended to cover any variations, uses or adaptations of the present disclosure that follow the general principles of the present disclosure and include common knowledge or conventional techniques in the art that are not described in the present disclosure.

[0147] The above is only a preferred embodiment of the present application and does not limit the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principles of the present application shall be included in the protection scope of the present application.

Claims

1. A method for pedestrian crossing intention prediction based on text-guided cascaded network, characterized in that, The method comprises: acquiring a road video using a vehicle-mounted camera and detecting pedestrians in the road video based on a YOLO algorithm to extract a pedestrian coordinate sequence with consecutive frames; sequentially performing dimension expansion and feature extraction on the pedestrian coordinate sequence, inputting high-dimensional features extracted by the pedestrian coordinate sequence into a decoder to obtain predicted pedestrian coordinates corresponding to the pedestrian coordinate sequence; splicing the pedestrian coordinate sequence and the predicted pedestrian coordinates and inputting the spliced pedestrian coordinate sequence into an encoder to obtain complete coordinate sequence features; generating a crossing behavior description corresponding to a structured prompt word generated by a large language model based on the structured prompt word and a pedestrian crossing video, and encoding the crossing behavior description into behavior description features; the structured prompt word generated by the large language model specifically comprises: taking a driving scene setting as a context, constructing a role-based prompt word based on a large language model, and generating an inference standard for judging pedestrian crossing behavior, wherein the inference standard comprises pedestrian behavior and body language, pedestrian position and location, traffic signals and road conditions, speed and motion of surrounding vehicles, and environment and time factors; extracting intention indicators and constructing a structured prompt word through semantic analysis and feature selection on the inference standard; aligning the complete coordinate sequence features and the behavior description features through a cosine similarity function, and obtaining pedestrian intention prediction results according to a loss function and a classification strategy corresponding to the aligned complete coordinate sequence features.

2. The text-guided cascading network-based pedestrian crossing intention prediction method of claim 1, wherein, the obtaining of the predicted pedestrian coordinates corresponding to the pedestrian coordinate sequence specifically comprises: dimension expansion on the pedestrian coordinate sequence based on a multilayer perceptron to obtain an upgraded pedestrian coordinate sequence; feature extraction on the upgraded pedestrian coordinate sequence based on an encoder to obtain high-dimensional features corresponding to the pedestrian coordinate sequence; inputting the high-dimensional features into a decoder to predict the predicted pedestrian coordinates corresponding to the moment when the pedestrian is about to take action.

3. The text-guided cascading network-based pedestrian crossing intention prediction method of claim 1, wherein, the splicing of the pedestrian coordinate sequence and the predicted pedestrian coordinates and the inputting of the spliced pedestrian coordinate sequence into an encoder to obtain complete coordinate sequence features specifically comprises: splicing the predicted pedestrian coordinates and the pedestrian coordinate sequence to obtain a spliced predicted coordinate sequence; constraining the spliced predicted coordinate sequence based on an L2 regular loss function, and sequentially performing dimension expansion and feature extraction on the spliced predicted coordinate sequence to obtain complete coordinate sequence features corresponding to the spliced predicted coordinate sequence.

4. The text-guided cascading network-based pedestrian crossing intention prediction method of claim 1, wherein, the generation of a crossing behavior description corresponding to the structured prompt word specifically comprises: inputting the structured prompt word and the pedestrian crossing video into a video understanding model to extract a text feature vector and a visual feature sequence, respectively; performing behavior feature fusion on the text feature vector and the visual feature sequence to obtain a pedestrian behavior description.

5. The text-guided cascading network-based pedestrian crossing intention prediction method of claim 1, wherein, The pedestrian behavior and body language includes one or more of gaze direction, body orientation, walking speed, and hand gesture, the pedestrian position and location includes one or more of zebra crossing or intersection, close to the curb, and middle of the road, the traffic signal and road condition includes one or more of pedestrian signal, vehicle signal, traffic density, and road width and isolation facilities, the speed and motion of surrounding vehicles includes one or more of other vehicle deceleration, other vehicle parking, and traffic gap, and the environment and time factor includes one or more of school area and residential area, bus station or parked vehicle, night, poor visibility, and weather condition.

6. The text-guided cascading network-based pedestrian crossing intention prediction method of claim 1, wherein, The encoder is an encoder stacked by 8 layers of BiMamba, and the decoder is a decoder composed of multiple layers of perception mechanism. 7.A text-guided cascaded network system for pedestrian crossing intention prediction, characterized in that, The cascade network pedestrian crossing intention prediction system (1) comprises a cascade network module (11), a behavior description generation module (12), and a central perception classification module (13), wherein, The cascade network module (11) is used to collect road video using a vehicle-mounted camera, detect pedestrians in the road video based on a YOLO algorithm, extract a pedestrian coordinate sequence with consecutive frames, sequentially perform dimension expansion and feature extraction on the pedestrian coordinate sequence, input high-dimensional features extracted from the pedestrian coordinate sequence into a decoder to obtain predicted pedestrian coordinates corresponding to the pedestrian coordinate sequence, splice the pedestrian coordinate sequence and the predicted pedestrian coordinates, and input the spliced pedestrian coordinate sequence into an encoder to obtain complete coordinate sequence features; The behavior description generation module (12) is used to generate a crossing behavior description corresponding to the structured prompt word based on the structured prompt word generated by the large language model and the pedestrian crossing video, and encode the crossing behavior description into behavior description features; The structured prompt word generated by the large language model specifically comprises: Taking the driving scene setting as the context, constructing a role-based prompt word based on a large language model, and generating reasoning standards for judging pedestrian crossing behavior, wherein the reasoning standards include pedestrian behavior and body language, pedestrian position and location, traffic signal and road condition, speed and motion of surrounding vehicles, and environment and time factor; Through semantic analysis and feature selection on the reasoning standards, an intention index is extracted and a structured prompt word is constructed; The central perception classification module (13) is used to perform feature alignment on the complete coordinate sequence features and the behavior description features through a cosine similarity function, and obtain pedestrian intention prediction results according to the loss function and classification strategy corresponding to the aligned complete coordinate sequence features.

8. An electronic device, comprising: The electronic device (2) comprises a processor (21), a memory (25), a user interface (23), and a network interface (24), the memory (25) is used to store instructions, the user interface (23) and the network interface (24) are used to communicate with other devices, and the processor (21) is used to execute the instructions stored in the memory (25) to enable the electronic device (2) to perform the method of any one of claims 1-6.

9. A computer readable storage medium having stored thereon a computer program, characterized in that, The computer program, which is executed by a processor, implements the steps of the method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • A multi-feature fusion cascade RNN structure and pedestrian prediction method

    CN111860269B

  • Pedestrian crossing intention prediction method and device, electronic equipment and readable storage medium

    CN113807298A

  • Method for predicting intention of traffic participant

    CN116665147A