Pedestrian crossing prediction method and system based on virtual-real multi-modal knowledge migration
Through virtual and real multimodal knowledge transfer technology, combined with teacher model pre-training, style conversion and knowledge distillation, multimodal features are integrated to predict pedestrian crossing intentions, solving the accuracy and generalization ability of pedestrian crossing detection in harsh environments, and significantly improving prediction accuracy.
Patent Information
- Application Number
- CN202510077537.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-17
- Publication Date
- 2025-05-16
AI Technical Summary
The prior art is difficult to effectively detect pedestrian crossing behavior under severe weather and lighting conditions, and the labeling data is limited, resulting in insufficient generalization capabilities of the model.
Using the virtual and real multimodal knowledge transfer method, the teacher model is pre-trained by the pedestrian frame in the synthetic data, combined with style converter, distribution approximator and knowledge distillation technology, the pedestrian frame data characteristics, style conversion characteristics and shared characteristics are integrated, and the gated unit is fused to predict the pedestrian's intention signal through crossing.
It significantly improves the accuracy of pedestrian crossing prediction, enhances the generalization ability and cross-domain adaptability of the model, and solves the problem of the differences in the cross-domain distribution of knowledge in different fields.
Smart Images

Figure CN120014566A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of traffic safety, and in particular to a pedestrian crossing prediction method and system based on virtual-real multimodal knowledge transfer. Background Art
[0002] The problem of pedestrian crossing prediction (PCP) has received increasing attention, and more and more scholars have studied the related contents of pedestrian posture, bounding box, vehicle speed and semantic segmentation map. However, it is challenging to label these clues in practice, especially the bad weather and lighting conditions pose a problem for PCP. Therefore, the key to improving the generalization ability of pedestrian crossing behavior detection models is to enhance sample diversity. Given the limited labeled data of pedestrian crossing behavior in actual scenarios, the current research trend is to generate synthetic datasets with dynamic changes, enrich training data by introducing domain adaptation technology, and then optimize the performance of the model in different environments, and improve the prediction performance by adopting a domain adaptation framework. However, there are significant cross-domain distribution differences in knowledge in different fields. This feature requires low accuracy when dealing with personalized content provision tasks. Summary of the invention
[0003] The purpose of the present invention is to provide a pedestrian crossing prediction method and system with virtual-real multimodal knowledge transfer to solve the above-mentioned problems.
[0004] To achieve the above object, the present invention adopts the following technical solutions: In a first aspect, the present invention provides a pedestrian crossing prediction method based on virtual-real multimodal knowledge transfer, comprising: By using the pedestrian boxes in the synthetic data generated by the model to pre-train the teacher model in the knowledge extractor, the pedestrian box data features at the future time p are obtained; Through the style converter, the visual features of the RGB frame of the synthetic data under various conditions are converted into the corresponding real RGB image to obtain the style conversion features; Integrate the depth map of synthetic data, the semantic segmentation map of synthetic data and the real RGB map to obtain shared feature embedding; The pedestrian frame data features, style conversion features and shared features are integrated into a learnable gating unit for fusion to predict the pedestrian crossing intention signal.
[0005] Furthermore, the teacher model in the knowledge extractor is pre-trained by using the pedestrian frame in the synthetic data to obtain the pedestrian frame data features at the future time p, including: Using synthetic data sets, the bounding boxes of pedestrians from time 0 to T are provided as input to the teacher model Transformer network, so as to obtain the pedestrian box information of synthetic data from time T to T+p, complete the pre-training stage and fix the model parameters; The teacher model is used to guide the pedestrian boxes in the real data, and the student model, namely the ResNet+LSTM network, is used to predict and obtain the pedestrian box information of the real data at time T~T+p.
[0006] Furthermore, the style converter converts the visual features of the RGB frame of the synthetic data under various conditions into corresponding real RGB images, including: Remove irrelevant background noise in the global image by cropping the rectangular area around the pedestrian bounding box and scaling it; The processed RGB frames are used to generate a set of style-transferred images by applying the adaptive instance normalization AdaIN method. Subsequently, these images are encoded by inputting the spatiotemporal backbone network model to predict the pedestrian's crossing intention.
[0007] Furthermore, the synthetic depth map, the synthetic semantic segmentation map and the real RGB map are integrated to obtain shared feature embedding, including: By cropping the rectangular area around the pedestrian bounding box and scaling it; For the rectangular area around the pedestrian box, the input data is encoded through the same backbone network, and the DisA network is used to perform bidirectional approximation to identify the shared feature distribution.
[0008] Furthermore, the pedestrian frame data features, style conversion features and shared features are integrated into a learnable gating unit for fusion, including: The pedestrian frame data features, style conversion features and shared features are stacked to obtain the input vector F; Pass the input vector F through the linear layer and the normalization layer to obtain the gating weight w of feature fusion; The vector fusion of the gating operation is achieved by weighted summation of vectors F and w.
[0009] Furthermore, the pedestrian crossing intention signal is predicted, including: The gated fusion vector is passed through a linear layer and a Gumbel-softmax function to obtain the probability score of the pedestrian crossing intention.
[0010] In a second aspect, the present invention provides a pedestrian crossing prediction system with virtual-real multimodal knowledge transfer, comprising: A pedestrian frame data feature acquisition module is used to pre-train the teacher model in the knowledge extractor by using the pedestrian frames in the synthetic data to obtain the pedestrian frame data features at the future time p; A style conversion module, which is used to convert the visual features of the RGB frames of the synthetic data under various conditions into corresponding real RGB images through a style converter; The shared feature acquisition module is used to integrate the synthetic depth map, the synthetic semantic segmentation map and the real RGB map to obtain the shared feature embedding; The prediction output module is used to integrate the pedestrian box data features, style conversion features and shared features into a learnable gating unit for fusion and predict the pedestrian crossing intention signal.
[0011] Furthermore, the teacher model in the knowledge extractor is pre-trained by using the pedestrian frame in the synthetic data to obtain the pedestrian frame data features at the future time p, including: Using synthetic data sets, the bounding boxes of pedestrians from time 0 to T are provided as input to the teacher model Transformer network, so as to obtain the pedestrian box information of synthetic data from time T to T+p, complete the pre-training stage and fix the model parameters; The teacher model is used to guide the pedestrian boxes in the real data, and the student model, namely the ResNet+LSTM network, is used to predict and obtain the pedestrian box information of the real data at time T~T+p.
[0012] In a third aspect, the present invention provides a computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the pedestrian crossing prediction method of virtual-real multimodal knowledge transfer when executing the computer program.
[0013] In a fourth aspect, the present invention provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps of a pedestrian crossing prediction method for virtual-real multimodal knowledge transfer.
[0014] Compared with the prior art, the present invention has the following technical effects: The present invention proposes a pedestrian crossing prediction method based on virtual-real multimodal knowledge transfer. By integrating advanced technologies such as style transfer, distribution approximation and knowledge distillation, it realizes the effective transfer of cross-domain knowledge between different categories and significantly improves the accuracy of pedestrian crossing prediction.
[0015] By pre-training the teacher model (Transformer network) with pedestrian boxes in synthetic data, the teacher model can learn the laws and patterns of pedestrian movement. This helps to effectively guide the student model (ResNet+LSTM network) to predict pedestrian boxes in subsequent steps.
[0016] The pre-training process improves the generalization ability of the model, allowing it to adapt faster and predict accurately when faced with real data.
[0017] The teacher model guides the pedestrian boxes in the real data and uses knowledge distillation to bring the prediction results of the student model closer to the prediction results of the teacher model. Knowledge distillation can make full use of the knowledge learned by the teacher model to improve the prediction performance of the student model while keeping the student model lightweight and efficient.
[0018] The style converter converts the visual features of the RGB frames of the synthetic data under various conditions into the corresponding real RGB frames. This step effectively narrows the domain gap between synthetic data and real data and improves the generalization ability of the model.
[0019] The AdaIN (Adaptive Instance Normalization) method is applied to generate a collection of style-transferred images, which are encoded by inputting the spatiotemporal backbone model to predict pedestrians' crossing intentions. The AdaIN method can quickly and effectively achieve style transfer while keeping the image content unchanged, providing more realistic and richer visual features for subsequent intention prediction.
[0020] The synthetic depth map, synthetic semantic segmentation map and real RGB map are integrated through the distribution approximator to obtain shared feature embedding. This step realizes the effective fusion of multimodal information and improves the prediction accuracy of the model by approximating the distribution of real data.
[0021] For the rectangular area around the pedestrian box, the DisA network is used to perform bidirectional approximation to identify the shared feature distribution. The DisA network can more accurately capture the key features of pedestrian motion and provide more reliable feature support for subsequent intention prediction.
[0022] The pedestrian frame features, style conversion features, and shared feature vectors are integrated into a learnable gating unit for fusion. The gating unit can adaptively adjust the weights of different features to achieve effective feature fusion, thereby improving the prediction accuracy of the model. The fused feature input is used to predict the pedestrian crossing intention signal. By integrating multimodal information and performing effective feature fusion, the model can more accurately predict the pedestrian's crossing intention, providing important technical support for autonomous driving, intelligent transportation and other fields.
[0023] This technical solution integrates advanced technologies such as style transfer, distribution approximation and knowledge distillation to achieve effective transfer of cross-domain knowledge between different categories, significantly improving the accuracy of pedestrian crossing prediction. Specifically, the prediction performance of the student model is improved through pre-training and knowledge distillation of the teacher model; the domain gap between synthetic data and real data is narrowed through style transfer and AdaIN methods; the effective fusion and feature extraction of multimodal information are achieved through distribution approximators and DisA networks; and the accuracy of the model's intention prediction is improved through learnable gating units and fused feature inputs. These technical effects work together to give this technical solution significant advantages and application prospects in the field of pedestrian crossing prediction. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] Figure 1 It is a flow chart of the pedestrian crossing prediction method of virtual-real multimodal knowledge transfer in the present invention; Figure 2 Schematic diagram of the structure of the pedestrian crossing prediction model for virtual-real multimodal knowledge transfer in the present invention. DETAILED DESCRIPTION
[0025] The present invention is further described below in conjunction with the accompanying drawings: Example 1, please refer to Figure 1 The present invention provides a pedestrian crossing prediction method based on virtual-real multimodal knowledge transfer, comprising: By using the pedestrian boxes in the synthetic data to pre-train the teacher model in the knowledge extractor, the pedestrian box data features at the future time p are obtained; Through the style converter, the visual features of the RGB frames of the synthetic data under various conditions are converted into corresponding real RGB images; Integrate the synthetic depth map, synthetic semantic segmentation map and real RGB map to obtain shared feature embedding; The pedestrian frame data features, style conversion features and shared features are integrated into a learnable gating unit for fusion to predict the pedestrian crossing intention signal.
[0026] This paper proposes a pedestrian crossing prediction method with virtual-real multimodal knowledge transfer, aiming to design an appropriate adaptation mechanism for cross-domain knowledge between different categories, and to achieve effective knowledge transfer in specific scenarios with the help of gated knowledge fusion. The framework integrating style transfer, distribution approximation and knowledge distillation is adopted to efficiently process multivariate information such as vision, semantics, depth and bounding box, so as to significantly improve the accuracy of pedestrian crossing prediction.
[0027] Embodiment 2, the present invention provides a pedestrian crossing prediction method based on virtual-real multimodal knowledge transfer, comprising: Step S1, pre-training the teacher model in the knowledge extractor by using the pedestrian boxes in the synthetic data, aiming to effectively guide the student model to predict the pedestrian boxes; Step S2, converting the visual features of the RGB frame of the synthetic data under various conditions into the corresponding real RGB frame through a style converter; Step S3, integrating the synthetic depth map, the synthetic semantic segmentation map and the real RGB map through a distribution approximator to obtain a shared feature embedding; Step S4, integrating pedestrian frame features, style conversion features and shared feature vectors into a learnable gating unit for fusion; Step S5, fusion feature input is used to predict pedestrian crossing intention signal.
[0028] Wherein step S1 comprises the following steps: Step S11, using a synthetic data set, providing the pedestrian's bounding box from time 0 to T as input to the teacher model Transformer network, so as to obtain the pedestrian box information of the synthetic data from time T to T+p, complete the pre-training stage and fix the model parameters; In step S12, the teacher model is used to guide the pedestrian frames in the real data, and the student model, i.e., the ResNet+LSTM network, is used to predict and obtain the pedestrian frame information of the real data at time T~T+p.
[0029] Step S2 includes the following steps: Step S21, by cropping the rectangular area around the pedestrian bounding box and scaling it appropriately, a large amount of irrelevant background noise in the global image is eliminated; In step S22, the processed RGB frames generate a set of style transfer images by applying the AdaIN method, and then these images are encoded by inputting the spatiotemporal backbone model to predict the pedestrian's crossing intention.
[0030] Step S3 includes the following steps: Step S31, removing irrelevant background noise in the global image by cropping and scaling the rectangular area around the pedestrian boundary box; Step S32: for the rectangular area around the pedestrian box, the input data is encoded through the same backbone network, and the DisA network is used to perform bidirectional approximation to identify the shared feature distribution.
[0031] Step S4 includes the following steps: Step S41, stacking pedestrian frame data features, style conversion features and shared features to obtain an input vector F; Step S42, the input vector F is passed through a linear layer and a normalization layer to obtain a gating weight w for feature fusion; Step S43, using the weighted sum of vectors F and w to implement vector fusion of the gated operation; Step S5 includes the following steps: Step S51, the gated fusion vector is passed through a linear layer and a Gumbel-softmax function to obtain a pedestrian crossing intention probability score.
[0032] Example 3 Continuous 16 frames of video are input as observations. The bounding boxes of pedestrian objects in frames 0 to 16 in the synthetic dataset are used as input. Through the three Transformer networks in the teacher model, each Transformer layer integrates 8 self-attention modules to extract the pedestrian box information from frames 16 to 32, and finally generate a vector containing 64-dimensional features. After pre-training the teacher model and freezing its parameters, the pedestrian boxes generated by real data are used as guidance, and learning is performed through the student model (ResNet18 and two LSTM networks). The hidden state dimension of each LSTM layer is 100, and finally effective knowledge distillation of 16 to 32 pedestrian boxes in real data is achieved.
[0033] In the style converter, a local image around the pedestrian box is selected as input. The local area is adjusted to 112×112 size to eliminate a large amount of irrelevant background interference in the global image. The format of the input data is (batchsize, 16, 3, 112, 112), where batchsize is set to 2. The style transfer image set is generated by applying the AdaIN method to the processed RGB frame, and the style transfer feature vector is obtained by encoding using the spatiotemporal backbone model to achieve image style transfer.
[0034] The distribution approximator and style converter use the same data structure as input. The input local image area is divided into patches of size 16×16, which are converted into visual tokens in the backbone model. The output dimension of the backbone model is set to a 64-dimensional feature vector. The 16-frame observation contains 784 tokens, which are bidirectionally approximated by the DisA network to generate shared feature embeddings.
[0035] The pedestrian box features, style conversion features, and shared feature vectors extracted by the three modules are integrated into a learnable gated unit to achieve feature fusion. Subsequently, the final pedestrian crossing intention prediction signal is generated through three linear layers and a Gumbel-softmax function.
[0036] In the present invention, the training process of the teacher model and the student model adopts the Adam optimizer, and the learning rate and decay rate are configured to be 0.8. The decay step size is set to 10, that is, after each training cycle, the learning rate is adjusted to 80% of its original value. The teacher model is trained for 10 cycles. For the JAAD dataset, the training cycle is set to 40, and for the PIE dataset, it is adjusted to 20 cycles. This adjustment is based on the consideration that the sample size of the PIE dataset is large and convergence has been achieved within 20 cycles. In order to ensure a fair evaluation, the training cycle of JAAD is consistent with that in PCPA. To avoid overfitting, we introduced a 50% random neuron deactivation rate. This study was implemented using Python 3.7 and PyTorch 1.7, and was executed on a system equipped with an NVIDIA GeForce RTX 2080 Ti GPU and 64GB of memory.
[0037] In the present invention, we constructed a new synthetic PCP dataset named Syn-PCP-3181. The Syn-PCP-3181 dataset consists of 3181 video sequences, of which 1,226 involve crossing behavior and 1955 do not. Each video sequence focuses on a pedestrian, involving a total of 3181 different pedestrians. The dataset contains 489,740 frames, of which each group of crossing sequences has 240 frames and each group of non-crossing sequences has 100 frames. For actual data, we used the JAAD and PIE datasets for analysis. The JAAD project brings together data from many countries around the world, and finally collected 346 video sequences, a total of 75,000 frames of video, involving 2,786 pedestrians. Among the samples, a total of 495 plus 191 samples involve pedestrian crossing behavior. The remaining 2,100 pedestrians are samples of not crossing the road. PIE collected six hours of continuous daytime data in Toronto, including 55 sequences, 293,000 frames of images and information about 1,834 pedestrians. Compared with other advanced models, our model achieved comparable performance on all indicators of the three datasets.
[0038] In yet another embodiment of the present invention, a pedestrian crossing prediction system for virtual-real multimodal knowledge transfer is provided, which can be used to implement the above-mentioned pedestrian crossing prediction method for virtual-real multimodal knowledge transfer. Specifically, the system includes: A pedestrian frame data feature acquisition module is used to pre-train the teacher model in the knowledge extractor by using the pedestrian frames in the synthetic data to obtain the pedestrian frame data features at the future time p; A style conversion module, which is used to convert the visual features of the RGB frames of the synthetic data under various conditions into corresponding real RGB images through a style converter; The shared feature acquisition module is used to integrate the synthetic depth map, the synthetic semantic segmentation map and the real RGB map to obtain the shared feature embedding; The prediction output module is used to integrate the pedestrian box data features, style conversion features and shared features into a learnable gating unit for fusion and predict the pedestrian crossing intention signal.
[0039] The division of modules in the embodiments of the present invention is schematic and is only a logical function division. There may be other division methods in actual implementation. In addition, each functional module in each embodiment of the present invention may be integrated into one processor, or may exist physically separately, or two or more modules may be integrated into one module. The above-mentioned integrated modules may be implemented in the form of hardware or in the form of software functional modules.
[0040] In another embodiment of the present invention, a computer device is provided, the computer device including a processor and a memory, the memory is used to store a computer program, the computer program includes program instructions, and the processor is used to execute the program instructions stored in the computer storage medium. The processor may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components, etc. It is the computing core and control core of the terminal, which is suitable for implementing one or more instructions, and is specifically suitable for loading and executing one or more instructions in the computer storage medium to implement the corresponding method flow or corresponding function; the processor described in the embodiment of the present invention can be used for the operation of a pedestrian crossing prediction method for virtual-real multimodal knowledge transfer.
[0041] In another embodiment of the present invention, the present invention also provides a storage medium, specifically a computer-readable storage medium (Memory), which is a memory device in a computer device for storing programs and data. It is understandable that the computer-readable storage medium here can include both built-in storage media in a computer device and, of course, extended storage media supported by the computer device. The computer-readable storage medium provides a storage space, which stores the operating system of the terminal. In addition, one or more instructions suitable for being loaded and executed by a processor are also stored in the storage space, and these instructions can be one or more computer programs (including program codes). It should be noted that the computer-readable storage medium here can be a high-speed RAM memory or a non-volatile memory, such as at least one disk memory. The processor can load and execute one or more instructions stored in the computer-readable storage medium to implement the corresponding steps of the pedestrian crossing prediction method for virtual-real multimodal knowledge transfer in the above embodiment.
[0042] It will be appreciated by those skilled in the art that embodiments of the present invention may be provided as methods, systems, or computer program products. Therefore, the present invention may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0043] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0044] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0045] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
[0046] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the above embodiments, ordinary technicians in the relevant field should understand that the specific implementation methods of the present invention can still be modified or replaced by equivalents. Any modification or equivalent replacement that does not depart from the spirit and scope of the present invention should be covered within the scope of protection of the claims of the present invention.
Claims
1. A pedestrian crossing prediction method based on virtual-real multimodal knowledge transfer, characterized in that: include: By using the pedestrian boxes in the synthetic data generated by the model to pre-train the teacher model in the knowledge extractor, the pedestrian box data features at the future time p are obtained; Through the style converter, the visual features of the RGB frame of the synthetic data under various conditions are converted into the corresponding real RGB image to obtain the style conversion features; Integrate the depth map of synthetic data, the semantic segmentation map of synthetic data and the real RGB map to obtain shared feature embedding; The pedestrian frame data features, style conversion features and shared features are integrated into a learnable gating unit for fusion to predict the pedestrian crossing intention signal.
2. According to the method of claim 1, the pedestrian crossing prediction method based on virtual-real multimodal knowledge transfer is characterized in that: The method of pre-training the teacher model in the knowledge extractor by using the pedestrian frames in the synthetic data to obtain the pedestrian frame data features at the future time p includes: Using synthetic data sets, the bounding boxes of pedestrians from time 0 to T are provided as input to the teacher model Transformer network, so as to obtain the pedestrian box information of synthetic data from time T to T+p, complete the pre-training stage and fix the model parameters; The teacher model is used to guide the pedestrian boxes in the real data, and the student model, namely the ResNet+LSTM network, is used to predict and obtain the pedestrian box information of the real data at time T~T+p.
3. The pedestrian crossing prediction method based on virtual-real multimodal knowledge transfer according to claim 1 is characterized in that: The style converter converts the visual features of the RGB frame of the synthetic data under various conditions into the corresponding real RGB image, including: Remove irrelevant background noise in the global image by cropping the rectangular area around the pedestrian bounding box and scaling it; The processed RGB frames are used to generate a set of style-transferred images by applying the adaptive instance normalization AdaIN method. Subsequently, these images are encoded by inputting the spatiotemporal backbone network model to predict the pedestrian's crossing intention.
4. The pedestrian crossing prediction method based on virtual-real multimodal knowledge transfer according to claim 1 is characterized in that: The synthesized depth map, the synthesized semantic segmentation map and the real RGB map are integrated to obtain shared feature embedding, including: By cropping the rectangular area around the pedestrian bounding box and scaling it; For the rectangular area around the pedestrian box, the input data is encoded through the same backbone network, and the DisA network is used to perform bidirectional approximation to identify the shared feature distribution.
5. The pedestrian crossing prediction method based on virtual-real multimodal knowledge transfer according to claim 1 is characterized in that: Integrate pedestrian frame data features, style conversion features, and shared features into a learnable gating unit for fusion, including: The pedestrian frame data features, style conversion features and shared features are stacked to obtain the input vector F; Pass the input vector F through the linear layer and the normalization layer to obtain the gating weight w of feature fusion; The vector fusion of the gating operation is achieved by weighted summation of vectors F and w.
6. The pedestrian crossing prediction method based on virtual-real multimodal knowledge transfer according to claim 1 is characterized in that: Predict pedestrian crossing intention signals, including: The gated fusion vector is passed through a linear layer and a Gumbel-softmax function to obtain the probability score of the pedestrian crossing intention.
7. A pedestrian crossing prediction system with virtual-real multimodal knowledge transfer, characterized in that: include: A pedestrian frame data feature acquisition module is used to pre-train the teacher model in the knowledge extractor by using the pedestrian frames in the synthetic data to obtain the pedestrian frame data features at the future time p; A style conversion module, which is used to convert the visual features of the RGB frames of the synthetic data under various conditions into corresponding real RGB images through a style converter; The shared feature acquisition module is used to integrate the synthetic depth map, the synthetic semantic segmentation map and the real RGB map to obtain the shared feature embedding; The prediction output module is used to integrate the pedestrian box data features, style conversion features and shared features into a learnable gating unit for fusion and predict the pedestrian crossing intention signal.
8. The pedestrian crossing prediction system of virtual-real multimodal knowledge transfer according to claim 7 is characterized in that: The method of pre-training the teacher model in the knowledge extractor by using the pedestrian frames in the synthetic data to obtain the pedestrian frame data features at the future time p includes: Using synthetic data sets, the bounding boxes of pedestrians from time 0 to T are provided as input to the teacher model Transformer network, so as to obtain the pedestrian box information of synthetic data from time T to T+p, complete the pre-training stage and fix the model parameters; The teacher model is used to guide the pedestrian boxes in the real data, and the student model, namely the ResNet+LSTM network, is used to predict and obtain the pedestrian box information of the real data at time T~T+p.
9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the steps of the pedestrian crossing prediction method of virtual-real multimodal knowledge transfer as described in any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of a pedestrian crossing prediction method for virtual-real multimodal knowledge transfer as described in any one of claims 1 to 7 are implemented.