Traffic accident risk prediction method and system based on urban streetscape
By performing multi-dimensional feature extraction and dual-channel prediction model analysis on street view images, the problem of inaccurate identification of traffic accident risks in urban street views has been solved, achieving high-precision risk prediction and street view optimization, and promoting the intelligent development of traffic safety.
Patent Information
- Application Number
- CN202511621645.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-07
- Publication Date
- 2026-01-20
AI Technical Summary
Existing technologies struggle to fully uncover the impact of complex environmental features on traffic accident risks in urban streetscapes, leading to inaccurate traffic accident risk identification and a lack of interpretable decision-making basis. Current research primarily relies on passive risk identification, which is insufficient to effectively reduce accident risks.
By extracting multi-dimensional features from street view images, including the fusion of semantic features, depth features, and entropy features, and using a dual-channel prediction model that combines visual feature extraction and quantitative feature analysis, traffic accident risk prediction is performed, generating high-precision prediction results and optimizing the street view environment.
It achieves high-precision traffic accident risk prediction, provides visualized risk identification and street view optimization solutions, and provides scientific guidance for autonomous driving technology and urban road planning, thereby reducing the risk of traffic accidents.
Smart Images

Figure CN121365873A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of image processing, in particular to a traffic accident risk prediction method and system based on urban street view. BACKGROUND
[0002] Urban traffic safety has become one of the major challenges in the process of global urbanization. Frequent traffic accidents pose a serious threat to the safety of life and property of residents and restrict the sustainable development of cities. Existing research on the mechanism of urban street view and traffic accidents still has many limitations. First, relying solely on semantic segmentation methods cannot fully exploit the influence of discrete features (such as street view image depth) in complex street view environments. Second, most studies have failed to explore the contribution mechanism of street view features to accident risk, making it difficult to provide city planners with interpretable decision-making basis. Existing research mainly focuses on passive traffic accident risk identification, resulting in a high risk of traffic accidents. SUMMARY
[0003] The embodiments of the present application aim to provide a traffic accident risk prediction method and system based on urban street view, which can reduce the risk of traffic accidents.
[0004] The technical solution of the present application is implemented as follows: In a first aspect, the embodiments of the present application provide a traffic accident risk prediction method based on urban street view, which comprises: obtaining a street view image and vehicle driving data of a current driving area; extracting multi-dimensional features from the street view image to obtain multi-modal features; wherein the multi-modal features include semantic features, depth features and entropy features; fusing the semantic features, the depth features and the entropy features to obtain multi-dimensional fusion features; using a pre-determined double-channel prediction model to predict traffic accident risk based on the multi-dimensional fusion features and the vehicle driving data, and obtaining a prediction result; and optimizing the urban street view corresponding to the street view image according to the prediction result.
[0005] In the above solution, the multi-dimensional feature extraction from the street view image to obtain the multi-modal features comprises: performing pixel-level semantic segmentation on the street view image to obtain the semantic features; extracting depth information based on the street view image to obtain the depth features; determining the entropy features by calculating the visual entropy of the street view image through a multi-scale image calculation module; determining the multi-modal features based on the semantic features, the depth features and the entropy features.
[0006] In the scheme, the pixel-level semantic segmentation of the street view image is performed to obtain the semantic feature, including: performing pixel-level semantic segmentation on the street view image to obtain a structured label map; each pixel in the structured label map corresponds to a semantic label; based on the structured label map, calculating the proportion of multiple different street view elements; based on the proportion of multiple different street view elements and the semantic label corresponding to each pixel in the structured label map, determining the semantic feature.
[0007] In the scheme, the depth information extraction based on the street view image is performed to obtain the depth feature, including: based on the street view image, generating a depth map corresponding to the street view image; the depth map represents the three-dimensional structure information of the urban street view; performing depth information extraction on the depth map to obtain average depth, depth variance, average gradient and gradient variance; the average depth represents the overall spatial openness of the scene; the depth variance represents the hierarchy and complexity of the object distribution in the scene; the average gradient represents the overall intensity of the scene geometric mutation; the gradient variance represents the distribution heterogeneity of the scene geometric mutation; based on the depth map, the average depth, the depth variance, the average gradient and the gradient variance, determining the depth feature.
[0008] In the scheme, the dual-channel prediction model includes a visual feature extraction module and a quantitative feature analysis module; the traffic accident risk prediction of the multi-dimensional fusion feature and the vehicle driving data through the pre-determined dual-channel prediction model to obtain a prediction result, including: the traffic accident risk prediction of the multi-dimensional fusion feature and the vehicle driving data through the visual feature extraction module in the dual-channel prediction model to obtain a first predicted accident probability; the traffic accident risk prediction of the multi-dimensional fusion feature and the vehicle driving data through the quantitative feature analysis module in the dual-channel prediction model to obtain a second predicted accident probability; based on the first predicted accident probability and the second predicted accident probability, determining the prediction result.
[0009] In the scheme, the traffic accident risk prediction of the multi-dimensional fusion feature and the vehicle driving data through the visual feature extraction module in the dual-channel prediction model to obtain a first predicted accident probability, including: extracting, by the visual feature extraction module in the double-channel prediction model, a depth map from the multi-dimensional fusion feature; Taking the depth map as a spatial attention guide signal, processing the multi-dimensional fusion feature and the spatial distribution and time sequence change of the vehicle driving data by a spatial-temporal attention submodule in the visual feature extraction module to obtain spatial distribution information and time sequence information; Based on the spatial distribution information and the time sequence information, calculating a spatial-temporal correlation weight by a three-dimensional attention matrix; Based on the spatial-temporal correlation weight, performing spatial-temporal weighting processing on the multi-dimensional fusion feature and the vehicle driving data to obtain a spatial-temporal weighted feature; Performing traffic accident risk prediction on the spatial-temporal weighted feature to obtain the first predicted accident probability.
[0010] In the above scheme, the traffic accident risk prediction on the multi-dimensional fusion feature and the vehicle driving data by the quantitative feature analysis module in the double-channel prediction model to obtain a second predicted accident probability comprises: Performing feature analysis on the multi-dimensional fusion feature by the quantitative feature analysis module in the double-channel prediction model to determine semantic features of a plurality of different street scene elements, four depth features, and an entropy feature; Performing risk analysis on the semantic features of the plurality of different street scene elements to determine respective risk indexes of the plurality of different street scene elements; Based on the four depth features, the respective risk indexes of the plurality of different street scene elements, and the entropy feature, performing traffic accident risk prediction to obtain the second predicted accident probability.
[0011] In a second aspect, an embodiment of the present application provides a traffic accident risk prediction system based on urban street scenes, which comprises an acquisition module, an extraction module, a fusion module, and a prediction module, wherein, The acquisition module is configured to acquire a street scene image and vehicle driving data of a current driving area; The extraction module is configured to perform multi-dimensional feature extraction on the street scene image to obtain multi-modal features; wherein the multi-modal features comprise semantic features, depth features, and an entropy feature; The fusion module is configured to perform multi-modal feature fusion on the semantic features, the depth features, and the entropy feature to obtain multi-dimensional fusion features; The prediction module is configured to perform traffic accident risk prediction on the multi-dimensional fusion features and the vehicle driving data by a predetermined double-channel prediction model to obtain a prediction result; and to optimize an urban street scene corresponding to the street scene image according to the prediction result.
[0012] In a third aspect, the embodiments of the present application provide a traffic accident risk prediction device based on urban street view, comprising: a processor and a memory; wherein, The memory is configured to store a computer program. The processor is configured to call and run the computer program from the memory to execute the method according to the first aspect.
[0013] In a fourth aspect, the embodiments of the present application provide a computer readable storage medium storing executable instructions for causing a processor to execute the method according to the first aspect.
[0014] The embodiments of the present application provide a traffic accident risk prediction method and system based on urban street view. The method comprises: obtaining a street view image and vehicle driving data of a current driving area; performing multi-dimensional feature extraction on the street view image to obtain multi-modal features; wherein the multi-modal features comprise semantic features, depth features and entropy features; performing multi-modal feature fusion on the semantic features, the depth features and the entropy features to obtain multi-dimensional fusion features; performing traffic accident risk prediction on the multi-dimensional fusion features and the vehicle driving data through a pre-determined double-channel prediction model to obtain a prediction result; and optimizing the urban street view corresponding to the street view image according to the prediction result. In the above solution, the street view image is extracted in multiple dimensions, and the urban road traffic accident risk is predicted based on the multi-dimensional features of the street view to determine the prediction result. The street view environment is optimized according to the prediction result, thereby reducing the risk of traffic accidents. BRIEF DESCRIPTION OF DRAWINGS
[0015] The accompanying drawings, which are incorporated in and constitute a part of the specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor on the basis of these drawings.
[0016] The flowchart shown in the accompanying drawings is only an exemplary illustration, and is not necessarily required to include all contents and operations / steps, nor is it necessarily required to be executed in the order described. For example, some operations / steps can be further divided, and some operations / steps can be combined or partially combined, so that the actual execution order can be changed according to the actual situation.
[0017] Figure 1 An optional flowchart of a traffic accident risk prediction method based on urban street view provided by the embodiments of the present application is shown in the accompanying drawings. Figure 2A structural schematic diagram of a traffic accident risk prediction system based on urban street view provided by an embodiment of the present application is shown in FIG. 1. Figure 3 A structural schematic diagram of a traffic accident risk prediction device based on urban street view provided by an embodiment of the present application is shown in FIG. 2. DETAILED DESCRIPTION
[0018] In order to make the objects, technical solutions and advantages of the embodiments of the present application clearer, the specific technical solutions of the present application will be further described in detail below with reference to the accompanying drawings of the embodiments of the present application. The following embodiments are used to illustrate the present application, but are not used to limit the scope of the present application.
[0019] Unless otherwise defined, all technical and scientific terms used in the present application have the same meanings as those commonly understood by one skilled in the art to which the present application belongs. The terms used in the present application are only for the purpose of describing the embodiments of the present application, and are not intended to limit the present application.
[0020] In the following description, the terms “some embodiments”, “the embodiment”, “the embodiments of the present application” and the like describe a subset of all possible embodiments, but it can be understood that “some embodiments” can be the same subset or different subsets of all possible embodiments, and can be combined with each other without conflict.
[0021] If the application file appears similar description of “first / second”, the following description is added, in the following description, the terms “first\second\third” only distinguish similar objects, and do not represent the specific order of the objects. It can be understood that “first\second\third” can be interchanged in specific order or sequence as allowed, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein.
[0022] The embodiments of the present application provide a traffic accident risk prediction method based on urban street view, Figure 1 An optional flowchart of a traffic accident risk prediction method based on urban street view provided by an embodiment of the present application will be described with reference to the steps shown in FIG. 3. Figure 1
[0023] S101, acquire street view images and vehicle driving data of a current driving area.
[0024] In some embodiments of the present application, the street view image is an important medium for recording urban space, including building facades, street layout, public facilities (such as street lamps, mailboxes), traffic signs and other physical space elements, which are the most basic components of the street view image. For example, through the street view, the urban morphological characteristics such as building height and street width can be quantitatively analyzed. Dynamic scenes such as pedestrians, vehicles, and markets are covered, such as street vendors, night market crowds, and traffic flow, which can reflect the rhythm of urban life and the state of social economy; natural elements such as greenery, water bodies, and weather phenomena (such as rain and fog) are included, as well as graffiti, billboards, and neon lights, which together shape the visual characteristics of the city. For example, the combination of neon lights and graffiti in the tunnel can form a unique underground space image.
[0025] In some embodiments of the present application, the vehicle driving data includes vehicle driving records, specifically including vehicle driving speed, vehicle state and vehicle surrounding environment under continuous driving time, and the like.
[0026] In some embodiments of the present application, the traffic accident risk prediction method based on urban street view is suitable for the urban traffic accident risk prediction scene.
[0027] In some embodiments of the present application, the traffic accident risk prediction method based on urban street view is suitable for the traffic accident risk prediction system based on urban street view.
[0028] In some embodiments of the present application, the street view image of the vehicle in the current driving area is obtained through the image acquisition device, and the vehicle driving data of the vehicle is obtained through the recorder of the vehicle.
[0029] S102, multi-dimensional feature extraction is performed on the street view image to obtain multi-modal features; wherein the multi-modal features include semantic features, depth features and entropy features.
[0030] In some embodiments of the present application, the street view image is subjected to pixel-level semantic segmentation to obtain semantic features; based on the street view image, depth information is extracted to obtain depth features; the visual entropy of the street view image is calculated through a multi-scale image calculation module to determine the entropy features; and the multi-modal features are determined based on the semantic features, the depth features and the entropy features.
[0031] S103, multi-modal feature fusion is performed on the semantic features, the depth features and the entropy features to obtain multi-dimensional fusion features.
[0032] In some embodiments of the present application, the semantic features, the depth features and the entropy features are input into a feature fusion model for multi-modal feature fusion to obtain multi-dimensional fusion features.
[0033] S104, predicting a traffic accident risk by the pre-determined double-channel prediction model based on the multi-dimensional fusion feature and the vehicle driving data to obtain a prediction result; and optimizing the urban street corresponding to the street view image according to the prediction result.
[0034] In some embodiments of the present application, the double-channel prediction model comprises a visual feature extraction module and a quantitative feature analysis module. In some embodiments of the present application, the visual feature extraction module in the double-channel prediction model is used to predict a traffic accident risk based on the multi-dimensional fusion feature and the vehicle driving data to obtain a first predicted accident probability; the quantitative feature analysis module in the double-channel prediction model is used to predict a traffic accident risk based on the multi-dimensional fusion feature and the vehicle driving data to obtain a second predicted accident probability; and the prediction result is determined based on the first predicted accident probability and the second predicted accident probability.
[0035] In some embodiments of the present application, when the prediction result indicates that a traffic accident is likely to occur, the first predicted accident probability and the second predicted accident probability are used to analyze high-risk street view factors so as to adjust the high-risk street view factors and optimize the urban street corresponding to the street view image.
[0036] It can be understood that the double-channel prediction model based on the precise correlation between micro street view visual features and accident risks can realize high-precision risk identification and visualization and further propose an operable street view optimization scheme, which not only provides more refined visual risk warning for autonomous driving technology, but also provides scientific and quantitative optimization guidance for urban road planning and design, thereby promoting the refinement and intelligent development of traffic safety management at the street view scale. Since the multi-dimensional features of the street view image are extracted, the urban road traffic accident risk is predicted based on the multi-dimensional features of the street view, the prediction result is determined, and the street view environment is optimized according to the prediction result, thereby reducing the risk of traffic accidents.
[0037] In some embodiments of the present application, S102 can be implemented by S201-S204 as follows: S201, performing pixel-level semantic segmentation on the street view image to obtain semantic features.
[0038] In some embodiments of the present application, the street view image is subjected to pixel-level semantic segmentation to obtain a structured label map; each pixel in the structured label map corresponds to a semantic label; the proportion of a plurality of different street view elements is calculated based on the structured label map; and the semantic features are determined based on the proportion of the plurality of different street view elements and the semantic label corresponding to each pixel in the structured label map.
[0039] S202, extracting depth information based on the street view image to obtain depth features.
[0040] In some embodiments of the present application, a depth map corresponding to the street view image is generated based on the street view image; the depth map represents three-dimensional structure information of the urban street view; depth information extraction is performed on the depth map to obtain average depth, depth variance, average gradient, and gradient variance; the average depth represents the overall spatial openness of the scene; the depth variance represents the hierarchy and complexity of the object distribution in the scene; the average gradient represents the overall intensity of the geometric mutation of the scene; and the gradient variance represents the distribution heterogeneity of the geometric mutation of the scene; and the depth feature is determined based on the depth map, the average depth, the depth variance, the average gradient, and the gradient variance.
[0041] In S203, the visual entropy of the street view image is calculated by the multi-scale image calculation module to determine the entropy feature.
[0042] In some embodiments of the present application, the visual entropy (Visual Entropy) is an image information measurement method combining information entropy theory and human visual characteristics (HVS), and the core is to quantify the human perceptible information complexity in the image. The visual entropy introduces a human visual sensitivity weight based on the information entropy, for example, a higher entropy value is given to a high-frequency texture area.
[0043] In some embodiments of the present application, the calculation of the visual entropy of the street view image includes grayscale, probability statistics, and entropy value calculation; the grayscale is to convert a color image to a grayscale image to simplify the calculation; the probability statistics is to calculate the probability of each grayscale value through a histogram; and the entropy value calculation is to apply an entropy formula, but the weight needs to be adjusted in combination with a visual model.
[0044] It should be noted that the street view image includes various driving influencing factors such as roads, buildings, pedestrians, vehicles, vegetation, and traffic signals; the roads, buildings, pedestrians, vehicles, vegetation, and traffic signals have different influences on the influencing factors of traffic accidents, and different weights are given to the roads, buildings, pedestrians, vehicles, vegetation, and traffic signals for calculating the visual entropy by applying the entropy formula.
[0045] In S204, the multi-modal feature is determined based on the semantic feature, the depth feature, and the entropy feature.
[0046] It can be understood that the multi-modal feature is determined by extracting the semantic feature, the depth feature, and the entropy feature based on the street view image, which is beneficial to subsequent traffic accident risk prediction by the multi-modal feature.
[0047] In some embodiments of the present application, S104 can be implemented by S301-S303 as follows: In S301, the traffic accident risk prediction is performed on the multi-dimensional fusion feature and the vehicle driving data by the visual feature extraction module in the dual-channel prediction model to obtain a first predicted accident probability.
[0048] In some embodiments of the present application, the depth map is extracted from the multi-dimensional fusion feature by a visual feature extraction module in the dual-channel prediction model; the depth map is taken as a spatial attention guide signal, and the spatial distribution and time sequence change of the multi-dimensional fusion feature and the vehicle driving data are processed by a spatial-temporal attention submodule in the visual feature extraction module to obtain spatial distribution information and time sequence information; based on the spatial distribution information and the time sequence information, a spatial-temporal correlation weight is calculated by a three-dimensional attention matrix; based on the spatial-temporal correlation weight, the multi-dimensional fusion feature and the vehicle driving data are subjected to spatial-temporal weighting processing to obtain a spatial-temporal weighted feature; and the spatial-temporal weighted feature is subjected to traffic accident risk prediction to obtain a first predicted accident probability.
[0049] In S302, the multi-dimensional fusion feature and the vehicle driving data are subjected to traffic accident risk prediction by a quantitative feature analysis module in the dual-channel prediction model to obtain a second predicted accident probability.
[0050] In some embodiments of the present application, the multi-dimensional fusion feature is subjected to feature analysis by a quantitative feature analysis module in the dual-channel prediction model to determine semantic features, depth features and entropy features of a plurality of different street scene elements; the semantic features of the plurality of different street scene elements are subjected to risk analysis to determine respective risk indexes of the plurality of different street scene elements; and based on the depth features, the respective risk indexes of the plurality of different street scene elements and the entropy features, traffic accident risk prediction is performed to obtain the second predicted accident probability.
[0051] In S303, a prediction result is determined based on the first predicted accident probability and the second predicted accident probability.
[0052] In some embodiments of the present application, if the predicted accident probability of at least one of the first predicted accident probability and the second predicted accident probability is greater than an accident probability threshold, it is indicated that the prediction result is that a traffic accident is likely to occur.
[0053] In some embodiments of the present application, if the first predicted accident probability is greater than the accident probability threshold, it is indicated that the prediction result is that a traffic accident is likely to occur; if the second predicted accident probability is greater than the accident probability threshold, it is indicated that the prediction result is that a traffic accident is likely to occur; and if both the first predicted accident probability and the second predicted accident probability are greater than the accident probability threshold, it is indicated that the prediction result is that a traffic accident is likely to occur.
[0054] It can be understood that the multi-dimensional fusion feature and the vehicle driving data are respectively subjected to traffic accident risk prediction by the visual feature extraction module and the quantitative feature analysis module in the dual-channel prediction model, and the features from different dimensions are analyzed, which can avoid inaccurate prediction of a single feature and improve the accuracy of traffic accident risk prediction.
[0055] The research process of the present application includes four main stages: data collection and preprocessing, model construction and training, result interpretation and optimization, and verification and evaluation.
[0056] In the data collection phase, the urban road traffic accident data is first screened, and the corresponding street view images are obtained as the model training data source. Then, the pre-trained model is used to extract the semantic information and depth information of the street view images. In the model construction and training phase, the CNN is used to associate the street view images with the probability of traffic accidents, and a visual information-based risk assessment model is constructed. Since CNN relies on visual perception, it can capture the complex patterns implied in street view images, but it has certain limitations in the quantitative modeling of explicit features. To make up for this deficiency, 11 quantitative street features extracted from semantic segmentation and depth estimation are further integrated, and an ANN regression model is used for training and optimization. ANN provides a relatively objective risk assessment by processing these quantitative data, enhancing the model's ability to analyze the accident mechanism. After model training, strict evaluation is carried out using MSE, MAE, RMSE and self-defined indicators, and SHAP analysis method is used to quantify the contribution of each explanatory variable (such as tree density, street canyon enclosure degree, sky proportion, etc.) to the model prediction results, revealing the influence mechanism of street elements on accident risk. Finally, based on the SHAP analysis results, image generation technologies such as Stable Diffusion and ControlNet are used to optimize the street view images, and the CNN model is re-inputted for verification to evaluate the improvement effect of street view optimization on accident risk prediction. Through this research framework, we not only improve the accuracy of urban road traffic accident risk prediction, but also provide a scientific basis for street view optimization, aiming to achieve a safer traffic environment.
[0057] The motor vehicle collision data used in this application is derived from the motor vehicle collision dataset provided by the XX City Open Data Portal (NYC Open Data), covering all motor vehicle collision events in XX City from 2016 to 2022, totaling 2,139,792 records. The data includes detailed information such as accident time, location, coordinates, number of deaths, accident cause, weather conditions and road conditions, providing a high-quality data basis for studying urban traffic accident risk. To ensure the relevance and accuracy of the data, this application screens and cleans the data as follows: First, the time range is limited to 2020-2022 (statistical year table) to capture the spatial distribution characteristics of recent accidents. Then, only accidents occurring between 6 am and 6 pm are retained to avoid the interference of low visibility at night on risk assessment. Finally, records containing complete latitude and longitude information are selected, and missing values and accident cause data unrelated to the urban built environment are excluded. After the above preliminary screening, a subset containing 598,654 valid records is obtained, which provides high-quality input for subsequent analysis and modeling.
[0058] To ensure the accuracy of spatial analysis, the geographic coordinates of all accident records are converted from the WGS84 coordinate system (EPSG: 4326) to the Long Island coordinate system (EPSG: 2263), which is more suitable for local spatial analysis in XX City. The converted coordinates provide geometric consistency for high-precision spatial calculations. During the data preprocessing stage, a 10-meter radius buffer zone is established for each accident point to quantify its spatial distribution characteristics. The selection of this buffer zone is based on the following considerations: (1) A 10-meter radius can capture the local characteristics of urban street scenes, such as intersections, lane distribution, and the impact of pedestrian activity areas. (2) Buffer zone definition helps identify high-accident density areas and accident clustering patterns. The specific steps include: (1) Generating a buffer zone based on geographic coordinates to capture the spatial range of each accident point. (2) Using the sjoin method in the geopandas library to detect and remove duplicate accident points. (3) Counting the number of accidents within each buffer zone to generate spatial density-based modeling data.
[0059] To further improve the usability of the data and the stability of the model training, the number of accidents data is log-normalized. The normalized data distribution is closer to the normal distribution, reducing the influence of extreme points and improving the convergence speed and accuracy of machine learning algorithms. Through the above strict data processing and preprocessing steps, the present application constructs a high-quality and balanced traffic accident data set, laying a solid foundation for subsequent model training and prediction analysis. This process ensures the integrity and representativeness of the data, while enhancing the reliability of the analysis results.
[0060] Street View data collection and processing: To build the association model between street view data and traffic accident risk, the application batched the street view images corresponding to the accident latitude and longitude coordinates through the API interface of the Google Street View platform. These images not only visually present the visual environment of urban space, but also provide rich input data for quantitative analysis. To further extract high-level semantic information from street view images, a pre-trained semantic segmentation model (ade20k-resnet50dilated-ppm_deepsup) based on the ADE20K dataset was used. This model is widely used in deep learning models for urban environment analysis, combining the ResNet-50 network architecture, dilated convolutions, and pyramid pooling modules (PPM), and using deep supervision. With its rich training data and powerful feature learning ability, ADE20K performs outstanding segmentation accuracy and generalization ability in complex scenarios, especially suitable for analyzing the multi-dimensional element distribution of urban street view images, and can identify and segment multiple semantic categories in the image (such as roads, buildings, pedestrians, vehicles, vegetation, traffic signals, etc.). Through semantic segmentation, the street view image is converted into a structured label map, and each pixel is assigned a corresponding semantic label, forming input data containing fine semantic information. In addition, the model also generates a quantitative table recording the pixel proportion of each street view element in the image. This quantitative table provides reliable index data support for in-depth understanding of the potential relationship between street view environment and traffic accidents. Finally, based on the street view images processed by the ADE20K model, not only does it provide structured input for subsequent machine learning models, but also provides accurate quantitative basis for studying the relative importance of different street view elements (such as green rate and building enclosure degree), thereby laying a data foundation for revealing the impact mechanism of street view environment on traffic accident risk.
[0061] In this application, in order to enrich the feature space of street view data, a deep learning model for depth estimation based on monocular image-MIDAS (Monocular Depth Estimation via a Self-supervised Network) is used to extract high-precision depth information from street view images. Its DPT-Swin2-Large-384 pre-training model version combines the feature extraction capabilities of Transformer architecture and Swin Transformer, especially suitable for capturing multi-scale spatial information in complex street scene. Through the processing of the MIDAS model, the corresponding depth map is generated from each street view image, and the value of each pixel represents the relative distance of the scene point to the observer. These depth maps provide three-dimensional structural information of urban street scenes, providing spatial feature data for subsequent analysis. In order to quantify the role of depth information in traffic accident prediction, two key statistical features are extracted from each depth map: mean depth and depth variance. The mean depth represents the average value of all pixel depth values in the street view image, suitable for measuring the overall spatial openness and object distribution density, and can reflect the overall distance relationship between objects in the scene and the observer. For example: a lower mean depth may correspond to a narrow street or a high-density building environment, and this spatial characteristic may limit the driver's visibility, and a higher mean depth usually indicates an open road or a wider driving space. Depth variance represents the degree of variation of depth values in the street view image, which can capture local changes and complexity in the scene, and can reflect the spatial complexity and structural changes of the scene. Higher depth variance usually indicates a complex road layout, such as multi-level intersections or larger spatial obstacles (such as roadblocks or traffic facilities). In addition, higher depth variance may also indicate significant height changes, such as high and low buildings or rapidly changing pedestrian and vehicle distribution. These characteristics may pose challenges to the driver's visual perception and judgment ability. By introducing the MIDAS depth estimation model and extracting key depth features, this application provides a new spatial analysis perspective for traffic accident risk prediction. Depth information not only enriches the feature space of street view data, but also provides important support for revealing the potential relationship between urban street scenes and traffic safety.
[0062] The method of the present application can adapt to complex and variable urban road environments, provide real-time and quantitative risk assessment and optimization feedback, and provide strong support for road safety simulation and optimization adjustment. It is worth emphasizing that the goal of the present application is not to completely eliminate traffic accidents, but to provide scientific and data-driven decision support tools for urban planners and designers, to promote road safety optimization, and to promote the development of urban traffic environment towards a safer and more sustainable direction.
[0063] The urban street view-based traffic accident risk prediction method of the present application further comprises: S1, multi-modal data collaborative collection and feature engineering.
[0064] ① Accident-street view spatio-temporal alignment database construction.
[0065] Data collection: Based on the XX city open platform, traffic accident data (original record 2,139,792) from 2016 to 2022 was obtained, and valid samples were screened by spatio-temporal constraint conditions (time window: 6:00-18:00; spatial accuracy: WGS84 to EPSG:2263 coordinate system). Through 10-meter buffer de-duplication and long-tail distribution correction (SMOTE oversampling + random undersampling), a balanced dataset (N=7,963) was finally constructed. Street view image acquisition: Google Street View API was called to collect multi-view street view images (pitch angle ± 15°, horizontal rotation angle 0° / 90° / 180° / 270°) centered on the accident point coordinates, resolution ≥2048×1024, and a spatio-temporal matching street view-accident pair was constructed.
[0066] ② Development of three-dimensional feature analysis system.
[0067] Semantic feature extraction: ADE20K pre-trained model (ResNet-50dilated+PPM_deepsup) was used for pixel-level semantic segmentation to calculate the proportion of 11 types of street view elements (such as SCE=building enclosure, SVI=sky visibility), depth feature calculation: based on MiDaS model (DPT-Swin2-Large-384) to generate depth map, extract average depth (MD, representing the overall spatial openness of the scene), depth variance (DV, reflecting the hierarchy and complexity of object distribution in the scene), average gradient (MG, quantifying the overall intensity of geometric mutations in the scene) and gradient variance (GV, revealing the distribution heterogeneity of geometric mutations).
[0068] Visual entropy analysis: Design multi-scale image entropy calculation module (window size 32×32 / 64×64 / 128×128), quantify the disorder of street view elements (IEI=weighted mean of entropy value).
[0069] S2, mixed model architecture design and explainability enhancement.
[0070] ① CNN-ANN dual-channel prediction model Visual Feature Extraction Branch (CNN): Backbone Network: EfficientNet-B4 pre-trained model (ImageNet weight initialization), improved design: embedding a deep attention module (Depth-Aware Attention), using the MiDaS depth map as a spatial attention guidance signal, output layer: regression head predicts accident probability, loss function uses Huber Loss (δ=1.0) to balance the influence of outliers. Quantized Feature Analysis Branch (ANN): Input Layer: 11-dimensional semantic features + 4-dimensional depth features + 1-dimensional entropy features (total 16 dimensions), hidden layer: 3 fully connected layers (64-32-16 neurons), activation function is Swish, batch normalization processing regularization strategy: Dropout (p=0.2) + L2 regularization (λ=1e-4).
[0071] ② Hierarchical interpretability analysis system An improved Grad-CAM++ (introducing deep channel weights) was used to locate high-risk sensitive areas. SHAP analysis was employed to quantify the contribution ranking of factors and nonlinear thresholds. Dynamic simulation: An interactive attribution web UI platform was developed to support visualization of risk transmission paths under parameter adjustments (e.g., the impact curve of increasing the greening rate from 20% to 30% on the accident incidence rate).
[0072] Based on the traffic accident risk prediction method based on urban street view described in the above embodiments, this application also provides a traffic accident risk prediction system based on urban street view, such as... Figure 2 As shown, Figure 2 This application provides a schematic diagram of the structure of a traffic accident risk prediction system based on urban street view. The traffic accident risk prediction system 2 based on urban street view includes: an acquisition module 201, an extraction module 202, a fusion module 203, and a prediction module 204. The acquisition module 201 is used to acquire street view images and vehicle driving data of the current driving area; The extraction module 202 is used to extract multi-dimensional features from the street view image to obtain multimodal features; wherein, the multimodal features include semantic features, depth features and entropy features; The fusion module 203 is used to perform multimodal feature fusion on the semantic features, the deep features and the entropy features to obtain multidimensional fused features; The prediction module 204 is used to predict traffic accident risks based on the multidimensional fusion features and the vehicle driving data using a pre-determined dual-channel prediction model, and to obtain prediction results; and to optimize the urban street scene corresponding to the street scene image based on the prediction results.
[0073] In some embodiments of the present application, the extraction module 202 is further configured to perform pixel-level semantic segmentation on the street view image to obtain the semantic feature; perform depth information extraction based on the street view image to obtain the depth feature; calculate the visual entropy of the street view image through a multi-scale image calculation module to determine the entropy feature; and determine the multi-modal feature based on the semantic feature, the depth feature, and the entropy feature.
[0074] In some embodiments of the present application, the extraction module 202 is further configured to perform pixel-level semantic segmentation on the street view image to obtain a structured label map; each pixel in the structured label map corresponds to a semantic label; calculate a proportion of a plurality of different street view elements based on the structured label map; and determine the semantic feature based on the proportion of the plurality of different street view elements and the semantic label corresponding to each pixel in the structured label map.
[0075] In some embodiments of the present application, the extraction module 202 is further configured to generate a depth map corresponding to the street view image based on the street view image; the depth map represents three-dimensional structural information of the urban street view; perform depth information extraction on the depth map to obtain an average depth, a depth variance, an average gradient, and a gradient variance; the average depth represents the overall spatial openness of the scene; the depth variance represents the hierarchy and complexity of the object distribution in the scene; the average gradient represents the overall intensity of the geometric mutation of the scene; and the gradient variance represents the distribution heterogeneity of the geometric mutation of the scene; and determine the depth feature based on the depth map, the average depth, the depth variance, the average gradient, and the gradient variance.
[0076] In some embodiments of the present application, the dual-channel prediction model includes a visual feature extraction module and a quantitative feature analysis module. The prediction module 204 is further configured to perform traffic accident risk prediction on the multi-dimensional fusion feature and the vehicle driving data through the visual feature extraction module in the dual-channel prediction model to obtain a first predicted accident probability; perform traffic accident risk prediction on the multi-dimensional fusion feature and the vehicle driving data through the quantitative feature analysis module in the dual-channel prediction model to obtain a second predicted accident probability; and determine the prediction result based on the first predicted accident probability and the second predicted accident probability.
[0077] In some embodiments of the present application, the prediction module 204 is further configured to extract a depth map from the multi-dimensional fusion feature through the visual feature extraction module in the dual-channel prediction model; use the depth map as a spatial attention guide signal, process the spatial distribution and time sequence change of the multi-dimensional fusion feature and the vehicle driving data through a spatio-temporal attention submodule in the visual feature extraction module to obtain spatial distribution information and time sequence information; calculate spatio-temporal correlation weights based on the spatial distribution information and the time sequence information through a three-dimensional attention matrix; perform spatio-temporal weighting processing on the multi-dimensional fusion feature and the vehicle driving data based on the spatio-temporal correlation weights to obtain spatio-temporal weighted features; and perform traffic accident risk prediction on the spatio-temporal weighted features to obtain the first predicted accident probability.
[0078] In some embodiments of the present application, the prediction module 204 is further configured to perform feature analysis on the multi-dimensional fusion feature through the quantitative feature analysis module in the dual-channel prediction model to determine semantic features, depth features, and entropy features of a plurality of different street scene elements; perform risk analysis on the semantic features of the plurality of different street scene elements to determine respective risk indexes of the plurality of different street scene elements; and perform traffic accident risk prediction based on the depth features, the respective risk indexes of the plurality of different street scene elements, and the entropy features to obtain the second predicted accident probability.
[0079] Based on the above-mentioned embodiment of the city street view-based traffic accident risk prediction method, the present application further provides a city street view-based traffic accident risk prediction device, as shown in Figure 3 Figure 3 FIG. 3 is a structural schematic diagram of a city street view-based traffic accident risk prediction device provided by an embodiment of the present application. The city street view-based traffic accident risk prediction device 3 includes a processor 301 and a memory 302. The memory 302 is configured to store a computer program, and the processor 301 is configured to call and run the computer program from the memory to perform the city street view-based traffic accident risk prediction method as described in the above-mentioned embodiments.
[0080] In the embodiments of the present application, the processor 301 can be at least one of an Application Specific Integrated Circuit (ASIC), a Digital Signal Processor (DSP), a Digital Signal Processing Device (DSPD), a Programmable Logic Device (PLD), a Field Programmable Gate Array (FPGA), a Central Processing Unit (CPU), a controller, a microcontroller, or a microprocessor. It can be understood that, for different devices, the electronic device used to implement the functions of the processor can also be other devices, and the embodiments of the present application are not limited in this regard.
[0081] The embodiments of the present application provide a computer readable storage medium storing a computer program, which is used to implement the traffic accident risk prediction method based on urban street view when executed by a processor.
[0082] For example, the program instructions corresponding to the traffic accident risk prediction method based on urban street view in the embodiments of the present application can be stored on a storage medium such as an optical disc, a hard disk, a U disk, etc. When the program instructions corresponding to the traffic accident risk prediction method based on urban street view in the embodiments of the present application are read by an electronic device or executed, the traffic accident risk prediction method based on urban street view as described in any of the above embodiments can be implemented.
[0083] In addition, each functional module in the embodiments of the present application can be integrated in one processing unit, or each unit can exist physically independently, or two or more units can be integrated in one unit. The integrated unit can be implemented in the form of hardware or in the form of a software functional module.
[0084] If the integrated unit is implemented in the form of a software function module and is not sold or used as an independent product, it can be stored in a computer readable storage medium based on such understanding. The technical solutions of the embodiments can essentially or contribute to the prior art or all or part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a plurality of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor to execute all or part of the steps of the embodiments. The aforementioned storage medium includes a U disk, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and various media that can store program codes.
[0085] It should be understood that the "one embodiment" or "an embodiment" or "some embodiments" mentioned throughout the specification means that the specific features, structures or characteristics related to the embodiments are included in at least one embodiment of the present application. Therefore, "in one embodiment" or "in an embodiment" or "in some embodiments" appearing throughout the specification does not necessarily refer to the same embodiment. In addition, these specific features, structures or characteristics can be combined in one or more embodiments in any suitable manner. It should be understood that in various embodiments of the present application, the size of the sequence number of the above processes does not mean the order of execution, and the execution order of the processes should be determined according to its function and inherent logic, and should not constitute any limitation on the implementation process of the embodiments of the present application. The above sequence number of the embodiments of the present application is only for description, not representing the advantages and disadvantages of the embodiments. The above description of each embodiment tends to emphasize the differences between each embodiment, and the same or similar parts can be referred to each other. For the sake of brevity, the present application will not be described again.
[0086] The above modules described as separate components can or can not be physically separated, and the components displayed as modules can or can not be physical modules; they can be located in one place or distributed on multiple network units; part or all of the modules can be selected according to actual needs to achieve the purpose of the embodiments.
[0087] In addition, each functional module in each embodiment of the present application can be integrated in one processing unit, or each module can be a separate unit, or two or more modules can be integrated in one unit. The integrated module can be realized in the form of hardware or hardware plus software function unit.
[0088] Those skilled in the art can understand that all or part of the steps of the foregoing method embodiments can be completed by relevant hardware of program instructions, and the foregoing program can be stored in a computer readable storage medium. When the program is executed, the program executes the steps of the foregoing method embodiments. The foregoing storage medium includes a mobile storage device, a read only memory (ROM), a magnetic disc or an optical disc, and various media that can store program codes.
[0089] The methods disclosed in the several method embodiments provided by the embodiments of the present application can be combined arbitrarily without conflict to obtain new method embodiments.
[0090] The features disclosed in the several product embodiments provided by the embodiments of the present application can be combined arbitrarily without conflict to obtain new product embodiments.
[0091] The features disclosed in the several method or device embodiments provided by the embodiments of the present application can be combined arbitrarily without conflict to obtain new method or device embodiments.
[0092] The foregoing is only a manner of implementing the embodiments of the present application, but the protection scope of the embodiments of the present application is not limited to this. Any person skilled in the art can easily think of changes or replacements within the technical range disclosed by the present application, which should be covered in the protection scope of the embodiments of the present application. Therefore, the protection scope of the embodiments of the present application should be subject to the protection scope of the claims.
Claims
1. A method for predicting traffic accident risk based on urban street view, characterized in that, The method comprises: acquiring a street view image and vehicle driving data of a current driving area; performing multi-dimensional feature extraction on the street view image to obtain multi-modal features; wherein the multi-modal features comprise semantic features, depth features and entropy features; performing multi-modal feature fusion on the semantic features, the depth features and the entropy features to obtain multi-dimensional fusion features; performing traffic accident risk prediction on the multi-dimensional fusion features and the vehicle driving data through a pre-determined double-channel prediction model to obtain a prediction result; and optimizing a city street view corresponding to the street view image according to the prediction result.
2. The method of claim 1, wherein, The multi-dimensional feature extraction on the street view image to obtain multi-modal features comprises: performing pixel-level semantic segmentation on the street view image to obtain the semantic features; extracting depth information based on the street view image to obtain the depth features; calculating the visual entropy of the street view image through a multi-scale image calculation module to determine the entropy features; determining the multi-modal features based on the semantic features, the depth features and the entropy features.
3. The method of claim 2, wherein, The pixel-level semantic segmentation on the street view image to obtain the semantic features comprises: performing pixel-level semantic segmentation on the street view image to obtain a structured label map; wherein each pixel in the structured label map corresponds to a semantic label; calculating proportions of a plurality of different street view elements based on the structured label map; determining the semantic features based on the proportions of the plurality of different street view elements and the semantic label corresponding to each pixel in the structured label map.
4. The method of claim 2, wherein, The extraction of depth information based on the street view image to obtain the depth features comprises: generating a depth map corresponding to the street view image based on the street view image; the depth map represents three-dimensional structural information of a city street view; extracting depth information from the depth map to obtain an average depth, a depth variance, an average gradient and a gradient variance; wherein the average depth represents the overall spatial openness of a scene; the depth variance represents the hierarchy and complexity of object distribution in the scene; the average gradient represents the overall intensity of geometric mutation of the scene; and the gradient variance represents the distribution heterogeneity of geometric mutation of the scene; determining the depth features based on the depth map, the average depth, the depth variance, the average gradient and the gradient variance.
5. The method of claim 1, wherein, The double-channel prediction model comprises a visual feature extraction module and a quantitative feature analysis module; The traffic accident risk prediction on the multi-dimensional fusion features and the vehicle driving data through the pre-determined double-channel prediction model to obtain a prediction result comprises: performing traffic accident risk prediction on the multi-dimensional fusion features and the vehicle driving data through the visual feature extraction module in the double-channel prediction model to obtain a first predicted accident probability; performing traffic accident risk prediction on the multi-dimensional fusion features and the vehicle driving data through the quantitative feature analysis module in the double-channel prediction model to obtain a second predicted accident probability; determining the prediction result based on the first predicted accident probability and the second predicted accident probability.
6. The method of claim 5, wherein, The traffic accident risk prediction based on the urban street view system comprises an acquisition module, an extraction module, a fusion module and a prediction module, wherein, The acquisition module is configured to acquire street view images and vehicle driving data of a current driving area; The extraction module is configured to perform multi-dimensional feature extraction on the street view images to obtain multi-modal features; wherein the multi-modal features comprise semantic features, depth features and entropy features; The fusion module is configured to perform multi-modal feature fusion on the semantic features, the depth features and the entropy features to obtain multi-dimensional fusion features; The prediction module is configured to perform traffic accident risk prediction on the multi-dimensional fusion features and the vehicle driving data by using a pre-determined double-channel prediction model to obtain a prediction result; and to optimize the urban street view corresponding to the street view images according to the prediction result. including:
7. The method of claim 5, wherein, a processor and a memory, wherein, the memory is configured to store a computer program; the processor is configured to call and run the computer program from the memory to execute the method of any one of claims 1 to 7. executable instructions are stored for causing the processor to execute when the method of any one of claims 1 to 7 is implemented. 8.A city street view based traffic accident risk prediction system, characterized in that, 9. An urban street view-based traffic accident risk prediction device characterized by comprising: 10. A computer-readable storage medium, characterized in that,