A street view recognition-based urban traffic three-dimensional model vegetation element automatic generation method, system, terminal and storage medium
By constructing a vegetation element knowledge graph and street view recognition technology, efficient and intelligent 3D scene generation is achieved, solving the problem of missing vegetation information in high-precision maps, improving the restoration accuracy and simulation efficiency of 3D scenes, supporting large-scale rapid construction and possessing reusability.
Patent Information
- Application Number
- CN202511501434.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-21
- Publication Date
- 2026-02-06
- Estimated Expiration
- 2045-10-21
AI Technical Summary
Existing high-precision maps lack road and vegetation information, resulting in insufficient 3D scene reproduction, low simulation efficiency and poor reusability. They rely on manual operation, which is inefficient, inconsistent in style, difficult to replicate on a large scale, and inconsistent material naming and parameters, leading to numerous compatibility issues.
A knowledge graph of vegetation elements is constructed. Pixel-level semantic segmentation and transfer learning are used to identify vegetation categories through street view recognition. Cosine similarity vector retrieval is used to match vegetation models. Batch import of 3D models is achieved through reference object correction and knowledge graph fusion to meet ecological constraints.
It achieves efficient and intelligent 3D scene generation, improves the fidelity and simulation efficiency of 3D scenes, supports large-scale rapid construction and has reusability, and reduces labor costs.
Smart Images

Figure CN120976448B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of urban simulation and three-dimensional modeling, and particularly relates to a method and system for automatically generating vegetation elements of a three-dimensional model of urban traffic based on street view recognition, a terminal, and a computer readable storage medium. BACKGROUND
[0002] With the increasing complexity of urban traffic systems and the popularization of technologies such as autonomous driving and vehicle-road cooperation, a high-fidelity three-dimensional simulation environment has become an essential tool for research and development, testing, and decision evaluation. The demand for algorithm verification, strategy evaluation, and emergency plan testing in a controllable and safe virtual environment is growing. At the same time, urban planning and digital management also require visual and analyzable three-dimensional scenes as a carrier for policy making and public communication. Therefore, the generation of three-dimensional scenes with high restoration, scalability, and standardization has certain social value and market demand.
[0003] Currently, the existing technology mainly has the following deficiencies:
[0004] (1) Data element missing: existing high-definition maps usually do not contain road vegetation and other information, which is crucial for visual simulation and scene aesthetics.
[0005] (2) Low simulation efficiency and poor reusability: current scene modeling can only complete batch generation of simple geometry, and relies heavily on manual drawing and manual adjustment. A large number of models need to be imported and corrected one by one, which is difficult to achieve precise matching at the detail level, and the workload is huge and difficult to maintain consistency. At the same time, it is also not conducive to the formation of reusable component library and standardized pipeline.
[0006] Therefore, the existing technology still needs to be improved and developed. SUMMARY
[0007] The main purpose of the present application is to provide a method and system for automatically generating vegetation elements of a three-dimensional model of urban traffic based on street view recognition, a terminal, and a computer readable storage medium, which aims to solve the problems of element missing, high labor cost, difficult material integration, and insufficient automation capability in the prior art during three-dimensional scene generation.
[0008] To achieve the above purpose, the present application provides a method for automatically generating vegetation elements of a three-dimensional model of urban traffic based on street view recognition, which comprises the following steps:
[0009] Construct a vegetation element knowledge graph for a target area, define vegetation types and typical attributes, directionally collect three-dimensional models and multi-view images, and store the unified metadata in a material library;
[0010] The street view image is pixel-level semantic segmentation, and the vegetation category is identified based on the transfer learning and attention mechanism, and the ROI and feature vector are output;
[0011] Based on the ontology mapping, the candidate set is screened, the cosine similarity vector retrieval method is used to select the most matched vegetation model for the local feature vector of the identified ROI and the pre-stored feature vector in the material library, and the mapping is recorded;
[0012] Based on the reference correction, knowledge graph fusion and automatic identification or optimization strategy, the reliable scale of the vegetation ROI is corrected, the quantity or position is inferred, and the batch three-dimensional model meeting the ecological constraints is imported and placed into the target simulation platform or visualization platform.
[0013] Optionally, the street view recognition-based urban traffic three-dimensional model vegetation element automatic generation method, wherein the vegetation element knowledge graph facing the target area is constructed, the vegetation type and typical attribute are defined, the three-dimensional model and multi-view image are collected in a directional manner, the metadata is unified and stored in the material library, and specifically includes:
[0014] Determine the target area, construct the vegetation element knowledge graph facing the target area, determine the common vegetation category hierarchy and attribute of the target area, and take the vegetation element knowledge graph as a semantic reference ontology for collection and retrieval;
[0015] Based on the category and attribute in the vegetation element knowledge graph, three-dimensional materials and multi-view images related to vegetation are collected and screened in a batch manner using a directional crawler and a map or data interface, and the source, license information and multi-view information are recorded;
[0016] Format and metadata standardization processing are performed on the collected three-dimensional materials, the naming specification is unified, the metadata is generated, and the material library is organized according to the semantic directory for retrieval and version management;
[0017] Format verification and automatic repair are performed on the original three-dimensional model and the map, the format is converted in batches, the problems are checked and automatically repaired, and the samples that do not meet the quality standards are labeled.
[0018] Optionally, the street view recognition-based urban traffic three-dimensional model vegetation element automatic generation method, wherein the street view image is pixel-level semantic segmentation, the vegetation category is identified based on the transfer learning and attention mechanism, and the ROI and feature vector are output, specifically including:
[0019] The street view image is pixel-level semantic segmentation to locate the vegetation area ROI, and the region mask and local feature are output to provide input samples for fine-grained identification;
[0020] Based on segmentation, a fine-grained classifier is trained to identify the ROI vegetation category;
[0021] The fine-grained classifier outputs a semantic label, a local confidence score, and a corresponding local visual feature vector for each identified object.
[0022] Optionally, the method for automatic generation of vegetation elements of a city traffic three-dimensional model based on street view recognition, wherein the fine-grained classifier is based on a transfer learning strategy, uses a deep backbone network pre-trained on a large-scale dataset as a basis, and combines an attention mechanism to fine-tune to a local vegetation dataset to enhance sensitivity to texture and morphological features.
[0023] Optionally, the method for automatic generation of vegetation elements of a city traffic three-dimensional model based on street view recognition, wherein the candidate set screening based on ontology mapping adopts a cosine similarity vector retrieval method to select the most matched vegetation model for the local feature vector of the identified ROI and the pre-stored feature vector in the material library, and records the mapping, specifically including:
[0024] According to the correspondence between the predefined categories and subcategories in the ontology, candidate resources of the same category are screened out in the material library, and preliminary filtering is performed according to the size range and version constraints;
[0025] Similarity calculation is performed on the depth feature vector extracted from the identified region and the candidate material feature vector, and the K candidates in the ordering result are taken as the final mapping candidate list in descending order of similarity;
[0026] The K candidates are executed based on a rule-based priority determination, and the final determined material ID and the corresponding texture file path of the map, normal and roughness are output.
[0027] Optionally, the method for automatic generation of vegetation elements of a city traffic three-dimensional model based on street view recognition, wherein the reference object correction, knowledge graph fusion, and automatic discrimination or optimization strategy are used to correct the reliable scale, quantity or position inference of the vegetation ROI, and batch three-dimensional models satisfying ecological constraints are imported and placed into a target simulation platform or visualization platform, specifically including:
[0028] For each vegetation ROI, the size of each vegetation ROI in the real space is determined through the map scale information, and the scaling factor is obtained by correcting the estimation result according to the typical scale of the reference object.
[0029] For each ROI, the observation features are automatically extracted based on the segmentation mask, and the planting type of the ROI is determined by combining the knowledge graph prior, to determine the number and size interval of the placed vegetation.
[0030] Based on the vegetation model and size or quantity estimation, scale and coordinate transformation is performed, the affine coordinate transformation required by the simulation platform is constructed, and batch import and instantiation are performed through the simulation interface, while the map and material parameters are batch assigned through the map interface.
[0031] After the matched three-dimensional model is instantiated and imported into the target simulation platform or visualization platform, an automatic acceptance check is performed, and an abnormal instance is marked in metadata and triggers a manual or remapping process, and finally a complete mapping link is retained.
[0032] Optionally, the automatic acceptance check includes collision body consistency check, missing texture check, and occlusion or visibility check.
[0033] In addition, to achieve the above object, the present application also provides a city traffic three-dimensional model vegetation element automatic generation system based on street view recognition, wherein the city traffic three-dimensional model vegetation element automatic generation system based on street view recognition comprises:
[0034] A knowledge graph and material library construction module is configured to construct a vegetation element knowledge graph for a target area, define vegetation types and typical attributes, directionally collect three-dimensional models and multi-view images, and store unified metadata in a material library.
[0035] A vegetation recognition module is configured to perform pixel-level semantic segmentation on street view images, identify vegetation categories based on transfer learning and attention mechanism, and output ROI and feature vectors.
[0036] A material matching module is configured to filter a candidate set based on ontology mapping, select the most matched vegetation model for the local feature vector of the identified ROI and the pre-stored feature vector in the material library by using a cosine similarity vector retrieval method, and record the mapping.
[0037] A model import module is configured to correct based on a reference object, fuse with a knowledge graph, and automatically determine or optimize strategies, correct the reliable scale of the vegetation ROI, infer the quantity or position, and import and place batch three-dimensional models that meet ecological constraints into a target simulation platform or visualization platform.
[0038] In addition, to achieve the above object, the present application also provides a terminal, wherein the terminal comprises a memory, a processor, and a city traffic three-dimensional model vegetation element automatic generation program based on street view recognition stored on the memory and executable on the processor, and the city traffic three-dimensional model vegetation element automatic generation program based on street view recognition, when executed by the processor, implements the steps of the city traffic three-dimensional model vegetation element automatic generation method based on street view recognition.
[0039] In addition, in order to achieve the above object, the present application also provides a computer readable storage medium, wherein the computer readable storage medium stores a street view recognition based automatic generation program of urban traffic three-dimensional model vegetation elements, and the street view recognition based automatic generation program of urban traffic three-dimensional model vegetation elements realizes the steps of the street view recognition based automatic generation method of urban traffic three-dimensional model vegetation elements when executed by a processor.
[0040] In the present application, a vegetation element knowledge graph facing a target area is constructed, vegetation types and typical attributes are defined, three-dimensional models and multi-view images are collected in a directional manner, metadata is unified and stored in a material library; pixel-level semantic segmentation is performed on street view images, vegetation categories are recognized based on transfer learning and attention mechanism, and ROI and feature vectors are output; candidate sets are screened based on ontology mapping, the most matched vegetation model is selected for the local feature vector of the recognized ROI and the pre-stored feature vector in the material library by using a cosine similarity vector retrieval method, and mapping is recorded; based on reference correction, knowledge graph fusion, automatic discrimination or optimization strategy, the reliable scale of the vegetation ROI is corrected, the quantity or position of the vegetation ROI is inferred, and batch three-dimensional models meeting the ecological constraints are imported and placed into a target simulation platform or a visualization platform. Through the integrated process of material collection standardization, semantic recognition, vector matching and programmed generation, the present application realizes an integrated pipeline from street view images to three-dimensional scenes, and provides efficient, intelligent and standardized technical support for traffic simulation and urban visualization. BRIEF DESCRIPTION OF DRAWINGS
[0041] Figure 1 is a flowchart of a preferred embodiment of the street view recognition based automatic generation method of urban traffic three-dimensional model vegetation elements of the present application;
[0042] Figure 2 is a schematic diagram of the process of generating urban traffic three-dimensional model vegetation elements in the preferred embodiment of the street view recognition based automatic generation method of urban traffic three-dimensional model vegetation elements of the present application;
[0043] Figure 3 is a schematic diagram of vegetation materials in the preferred embodiment of the street view recognition based automatic generation method of urban traffic three-dimensional model vegetation elements of the present application;
[0044] Figure 4 is a generation effect diagram in the preferred embodiment of the street view recognition based automatic generation method of urban traffic three-dimensional model vegetation elements of the present application;
[0045] Figure 5 is a structure diagram of a preferred embodiment of the street view recognition based automatic generation system of urban traffic three-dimensional model vegetation elements of the present application;
[0046] Figure 6Structure diagram of a preferred embodiment of the terminal of the present application. DETAILED DESCRIPTION
[0047] In order to make the objectives, technical solutions, and advantages of the present application clearer and more apparent, the present application will be further described in detail below with reference to the accompanying drawings and examples. It should be understood that the specific examples described herein are merely intended to explain the present application and are not intended to limit the present application.
[0048] With the development of urbanization and intelligent transportation, high-precision traffic simulation and three-dimensional scenes have become an important infrastructure in many fields such as automatic driving verification, traffic management, urban planning, and emergency drills. However, the existing high-precision maps and modeling processes generally lack key elements such as vegetation types, resulting in insufficient restoration of three-dimensional scenes and difficulty in meeting high-fidelity simulation needs. At the same time, traditional modeling is highly dependent on manual operation, which is low in efficiency, inconsistent in style, difficult to scale, and inconsistent in naming and parameters from different sources, bringing a large number of compatibility problems to batch processing and scripted integration. Therefore, a standardized and intelligent automatic pipeline from street view or map data to three-dimensional scenes is needed to support large-scale, fast, and reusable traffic three-dimensional scene construction.
[0049] In view of one or more of the above problems, the present application constructs a vegetation element knowledge graph for a target area, defines vegetation categories and their typical attributes, directionally collects three-dimensional models and multi-view images, and stores unified metadata in a database; performs pixel-level semantic segmentation on street view images and identifies vegetation categories based on transfer learning and attention mechanisms, and outputs ROI and feature vectors; filters a candidate set based on ontology mapping, selects the most matching vegetation model for the local feature vector of the identified ROI from the pre-stored feature vectors in the material library using a cosine similarity vector retrieval method, and records the mapping; and through reference object correction, knowledge graph fusion, and automatic discrimination or optimization strategies, reliable scale correction, quantity or position inference, and batch three-dimensional model import and placement that meet ecological constraints for the vegetation ROI are achieved.
[0050] The method for automatically generating vegetation elements of urban traffic three-dimensional models based on street view recognition according to the preferred embodiment of the present application, as shown in Figure 1 and Figure 2 includes the following steps:
[0051] Step S10, construct a vegetation element knowledge graph for a target area, define vegetation types and typical attributes, directionally collect three-dimensional models and multi-view images, and store unified metadata in a material library.
[0052] Specifically, a target area is determined, a vegetation element knowledge graph facing the target area is constructed, common vegetation category hierarchies and attributes of the target area are determined, and the vegetation element knowledge graph is taken as a semantic reference ontology for subsequent collection and retrieval. Based on the categories and attributes in the vegetation element knowledge graph, directional crawlers and map or data interface are used to collect and filter three-dimensional materials and multi-view images related to vegetation in batches, and the sources, license information and multi-view information are recorded. The collected three-dimensional materials are formatted and metadata standardized, the naming specification is unified, metadata is generated, and the material library is organized according to the semantic directory for retrieval and version management. Format verification and automatic repair are performed on the original three-dimensional models (referring to the original source version of the collected three-dimensional model file, which has not been formatted) and maps (in three-dimensional modeling, a map refers to a two-dimensional image mapped to the surface of a three-dimensional model, used to represent the color, detail, material and other characteristics of the surface). Batch format conversion, problem checking and automatic repair (such as automatic repair of normal direction, non-manifold geometry, UV overlap, missing map link and other common problems) are performed, and samples that do not meet the quality standards are labeled (such as being labeled as “to be manually reviewed”).
[0053] Specifically, an ontology engineering method is used to construct a vegetation element knowledge graph, the hierarchical relationship of common vegetation categories in a certain area (class→genus→species or class→subclass, etc.) is annotated by experts and sorted out from literature or existing databases, and key attributes are defined for each node: typical crown range, typical height range, seasonal characteristics, typical leaf shape or leaf color, typical growth density, ecological constraints, etc. The knowledge graph is implemented in an exchangeable ontology format (OWL or RDF), and the ontology is maintained in a version and annotated in a construction tool (such as Protégé), so as to use the ontology as a semantic reference for crawler priority and subsequent retrieval.
[0054] Based on the above vegetation element knowledge graph, a directional crawler framework is designed and deployed to collect materials and street view images in batches. The crawler uses the vegetation categories and attribute keywords listed in the ontology as retrieval seeds to perform semantic priority crawling on target material sites and multi-view preview pages, etc. For example, based on the research on Unreal Engine general formats (FBX, OBJ or GLB) on the Aigei material resource website, a directional crawler framework is designed. When the crawler is executed, it automatically captures the “name”, “download link”, “license” and multi-angle preview information of the model page, and downloads the mainstream format files.
[0055] Further, the metadata of all original models and maps is standardized. Referring to the material library metadata framework, the corresponding JSON description file is uniformly generated, and the fields cover id, name, source website, license, category, subcategory, number of triangular faces, map resolution, model size (length x width x height), upload time, version number, keyword label, etc., which is convenient for subsequent retrieval and tracing. All model files are renamed according to the category_subcategory_source_version.fbx specification.
[0056] In the format verification and automatic repair stage, with the help of Blender Python API, non-FBX format (3DS or DAE) is batch converted to FBX (Filmbox), and the model is checked for normal direction, non-manifold geometry, UV overlap (multiple faces occupy the same area in the map coordinate system, which may cause problems such as map misplacement, pattern stretching, etc.), missing map link, etc. For slight geometric errors, the normal is automatically reconstructed, and the blank UV is filled. Serious error cases are marked as “to be manually reviewed”.
[0057] After completing the file verification, the original materials are quality screened, and samples with blur, distortion, overexposure or incomplete composition are removed to ensure that the images in the library have clear crown and branch details. This process ensures that all models are presented under the same lighting and background conditions, ensuring visual consistency in subsequent scenes. The keyword label of each material after screening is mapped to the ontology to generate semantic fields consistent with the target layer elements.
[0058] At the same time, street view map data is collected, and the segmentation information of the target area is extracted from the map. The sampling points along the two sides of the road are obtained at a set interval, and the multi-view images or panoramic snapshots of the corresponding view are obtained based on the sampling points calling the street view image interface.
[0059] The map-related meta information of each captured street view image is recorded, including latitude and longitude, road ID, tile_id or tile coordinates, map zoom level, map resolution, camera orientation or pitch angle, shooting time, etc. After preprocessing, a JSON description file is uniformly generated.
[0060] Step S20, pixel-level semantic segmentation is performed on the street view image, the vegetation category is identified based on transfer learning and attention mechanism, and the ROI and feature vector are output.
[0061] Specifically, pixel-level semantic segmentation is adopted for street view images to locate the vegetation region ROI, and the region mask and local features are output (the region mask is positioning information in the form of a binary image for marking the position of the ROI, and each pixel indicates whether the pixel belongs to vegetation; the local features are feature vectors, which are description information extracted for the ROI region, and are deep features from the intermediate layer output of the neural network, including texture, morphology, color, etc.), providing input samples for fine-grained recognition; based on segmentation, a fine-grained classifier is trained to identify the ROI vegetation category.
[0062] Specifically, a vegetation image dataset related to the target region is constructed, unified preprocessing and data augmentation are performed, and the dataset is divided into training set, validation set and test set in proportion to evaluate Top-1 (Top-1 accuracy is used to evaluate the recognition accuracy of the model for the vegetation category, Top-1 refers to whether the class with the highest prediction probability is the real class, and for example, Top-5 refers to whether the real class is within the top 5 prediction probabilities) accuracy and robustness; the fine-grained classifier is based on a transfer learning strategy, using a deep backbone network pre-trained on a large-scale dataset as the basis, combined with an attention mechanism to fine-tune to the local vegetation dataset to enhance the sensitivity to texture and morphological features, improve cross-domain feature alignment and recognition accuracy under small sample; the fine-grained classifier outputs the semantic label, local confidence score, and corresponding local visual feature vector of each recognition object.
[0063] Specifically, the street view RGB image to be processed is preprocessed according to a unified specification: first, center or equal ratio scaling is performed according to the aspect ratio, and then normalized to the ImageNet mean and standard deviation; the image used for pixel-level semantic segmentation is uniformly adjusted to 1024x2048 to balance segmentation accuracy and computational overhead; the ROI cropped image used for fine-grained classification is uniformly scaled to 224x224 pixels in subsequent steps to adapt to the input requirements of the fine-grained classifier. The preprocessing also records the geographical reference of the original image for subsequent coordinate mapping.
[0064] For localized vegetation types, the image dataset required for fine-grained recognition is supplemented. The preprocessed images are divided into training set, validation set and test set in the ratio of 8:1:1. The training set is used for network training and iterative optimization, the test set is used to evaluate the Top-1 Accuracy (classification evaluation index, the proportion of the class with the highest prediction probability equal to the real class) of the final model in the 20-class local vegetation recognition task, and the validation set is used to verify the performance and generalization ability of the network.
[0065] Based on the DeepLabV3+ native architecture, the pixel-level semantic segmentation is performed for key elements in the street view map image, and feature fusion is performed for the semantic segmentation process of street elements in complex scenes. By accurately defining the ROI (Region of Interest) of each type of element, an input is provided for the downstream recognition and detection task. The Backbone (backbone network, which refers to the main convolutional network used to extract image features) selects ResNet-101 pre-trained on ImageNet.
[0066] The semantic segmentation network of DeepLabV3+ first performs a series of downsampling and dilated convolution operations on the input RGB street view image to obtain rich high-level semantic features. Subsequently, five parallel features (dilation rates = 1, 6, 12, 18, 24) are generated through the Atrous Spatial Pyramid Pooling (ASPP) module, then reduced in dimension through 1x1 convolution and fused, and after fusion with shallow features in the Decoder, upsampled to the original image size through bilinear interpolation, and the probability distribution of each pixel is output.
[0067] The pixel-level cross-entropy (Cross-Entropy) is used as the loss function, and the learning rate decay strategy is adopted. The optimizer uses SGD+Momentum (momentum = 0.9, weight_decay = 1e-4), where SGD (Stochastic Gradient Descent) represents the basic deep learning optimization algorithm, momentum is the momentum mechanism, and weight_decay is the weight decay, that is, when using the stochastic gradient descent, a momentum of 0.9 is introduced to accelerate convergence and reduce oscillation, and a decay of 1e-4 is added to the weight parameter to prevent overfitting.
[0068] The trained DeepLabV3+ semantic segmentation model is tested. Load the DeepLabV3+ pre-trained weights, perform single-image full-resolution forward inference, and obtain the classification pixel probability map. Post-processing such as DenseCRF (Dense Conditional Random Field, a post-processing method based on pixel and color or position similarity, used to refine the segmentation boundary) and connected component filtering is performed to refine the edge, remove noise, and improve the coherence of the segmentation boundary. The segmentation effect is monitored by mIoU (Mean Intersection over Union, average intersection over union) and pixel accuracy (PA, Pixel accuracy).
[0069] After completing the DeepLabV3+ based semantic segmentation, the ROI of the vegetation elements is located, and these areas are cut, and the semantic labels, region masks and local features are output to provide input samples for subsequent fine-grained identification.
[0070] Further, a lightweight vegetation class detection model is trained to output the attribute labels and spatial positions corresponding to each element. The determination of the identification basic network of the lightweight vegetation class detection model selects deep convolutional neural networks DenseNet, Inception, ResNeXt and MobileNet with different architectures, respectively performs full network training on the image training data set, and uses the trained model to infer the test data set. The above model effects are evaluated by the identification accuracy Top1-ACC (Top1 accuracy), and the pre-trained model is fine-tuned and predicted by using the self-built data set, and the identification effects of the network model migrated on ImageNet are compared. The ResNet50 with the highest Top1-ACC accuracy is obtained as the backbone network of vegetation identification.
[0071] The training effect of the network parameters on the local data set may be limited by the size of the sample amount, so based on large data sets such as ImageNet, the structure and parameters obtained by pre-training are migrated to the local data set through transfer learning fine-tuning.
[0072] Further, on the basis of the basic network model, an attention mechanism is added to improve the identification accuracy. Training and testing are performed on the data set, and the CBAM (Convolutional Block Attention Module, a lightweight attention mechanism that sequentially performs channel attention and spatial attention on the feature map) attention mechanism is adopted, which fully considers the combination of channel domain attention and spatial domain attention, and increases the average value pooling and convolution operation. The optimized network selects the most matched material candidate.
[0073] For each vegetation ROI, two types of parallel feature vectors are extracted: one from the intermediate layer of the semantic segmentation network; the other from the penultimate layer of the fine-grained classifier. The extracted original floating point vectors are L2 normalized and PCA is performed to reduce storage and retrieval overhead; at the same time, a set of simple color or texture statistics are generated as auxiliary features. All features are written in a unified format to the JSON metadata field corresponding to the ROI, and are simultaneously written into the vector database for efficient approximate nearest neighbor retrieval.
[0074] Step S30, screening the candidate set based on ontology mapping, using cosine similarity vector retrieval method to select the most matched vegetation model for the local feature vector of the identified ROI and the pre-stored feature vector in the material library, and record the mapping.
[0075] Specifically, according to the correspondence between the categories and subcategories predefined in the ontology, the same type of candidate resources are screened out in the material library, and preliminary filtering is performed according to the size range and version constraints; similarity calculation is performed on the depth feature vector extracted from the identified region and the candidate material feature vector, and the K top candidates in the ordering result are taken as the final mapping candidate list (i.e. taking the Top-K candidates as the final mapping candidate list, that is, taking the top K with the highest similarity in the ordering result as the candidate set); the K candidates (i.e. Top-K candidates) are executed based on rule-based priority determination, and the finally determined material ID and the corresponding map, normal and roughness material file path are output.
[0076] Specifically, based on the vegetation knowledge graph constructed in step S10, the species or category label output in step S20 is mapped to the semantic category of the material library. In the mapping process, all matching material sets are retrieved in the material library according to the large category or subcategory (such as "banyan tree or large tree"). Then, according to the basic constraints (license compatibility, maximum or minimum size range of the model, etc.), the set is preliminarily filtered to obtain the candidate set. This step ensures that vector retrieval is only performed in a limited space related to semantics, thereby improving retrieval efficiency and matching semantic consistency.
[0077] For each identified ROI, take the local feature vector u provided in step S20, and perform L2 normalization on u to obtain u'. Each candidate material in the material library has pre-calculated and stored its visual or geometric feature vector v when it is stored in the library, and also performs L2 normalization to obtain v'.
[0078] The cosine similarity between each v' and u' is calculated in the candidate set:
[0079] ; (1)
[0080] The retrieval returns Top-K candidates ordered from high to low similarity. The similarity score, material data, etc. are retained in the returned results. Priority rules are applied to the Top-K candidates to determine the final candidate, and the priority rules can be set as a weighted combination of similarity, size matching degree, and license. The size matching degree is obtained by comparing the relative error between the model size marked by the material and the estimated size of the ROI, and the smaller the error, the higher the score; the license score is binary / graded, compatible=1, incompatible=0. The highest comprehensive score is the preferred match; if the similarity is higher than the acceptance threshold, it is automatically accepted; if the similarity is between the rejection threshold and the acceptance threshold, it is marked as a semi-automatic candidate and enters the artificial audit queue; below the rejection threshold, the fallback strategy is triggered to expand the semantic filtering range.
[0081] Step S40, based on the reference correction, knowledge graph fusion and automatic discrimination or optimization strategy, the reliable scale of the vegetation ROI is corrected, the quantity or position is inferred, and the batch three-dimensional model meeting the ecological constraints is imported and placed into the target simulation platform or visualization platform.
[0082] Specifically, for each vegetation ROI, the size of each vegetation ROI in the real space is determined through the map scale information, the scaling factor is obtained by correcting the estimation result according to the typical scale of the reference object; the observation features such as the minimum bounding rectangle length-width ratio of each ROI automatically extracted based on the segmentation mask are combined with the knowledge graph prior to determine the ROI planting type (single or multiple or low shrub), and the vegetation quantity and size interval are determined; based on the vegetation model and size or quantity estimation, scale and coordinate transformation is performed, affine coordinate transformation required by the simulation platform is constructed and batch import and instantiation are performed through the simulation interface, at the same time, batch assignment of textures and material parameters is performed through the mapping interface; after the matched three-dimensional model is instantiated and imported into the target simulation platform or visualization platform, automatic acceptance inspection is performed, including collision body consistency check, texture missing check, occlusion or visibility check, abnormal instances are marked in the metadata and trigger artificial or remapping process, and finally the complete mapping link is retained. As shown in Figure 3 As shown in Figure 4 As shown in
[0083] After determining the three-dimensional vegetation model, the scale ratio is calculated, the scaling factor is generated by proportion conversion according to the original model size of the material and the target size interval of the ROI.
[0084] Specifically, read the global scale s (m / px) of the street map, and combine the reference object for scale correction. In the image where the ROI is located, detect road elements with typical known width, such as lane width, double yellow line spacing, and sidewalk plate, etc. These elements can be obtained through semantic segmentation and rule detection, for example, the average width of the lane can be obtained by morphological detection of the lane line area. Assuming that the average width of the detected lane is pixels, and the knowledge graph gives the standard width of the lane in this area as , then calculate the local scale:
[0085] ; (2)
[0086] Fuse the global scale and , and use Bayesian update based on confidence:
[0087] ; (3)
[0088] wherein represents the square of the variance of the uncertainty of the global scale, represents the square of the uncertainty of the local scale, and the uncertainty is determined according to the reliability of the reference object; represents the posterior scale after fusion.
[0089] After obtaining the scale, the physical scale is calculated:
[0090] ; (4)
[0091] wherein represents the pixel area of the ROI, represents the width of the pixel bounding box, represents the actual physical area, represents the actual physical width.
[0092] Further, for each ROI, the planting type is determined according to the actual area, length-width ratio, and the number of connected domains, distribution sparsity, and other morphological characteristics of the vegetation to determine whether the vegetation distribution of the ROI is single tree, row of street trees, or shrub belt.
[0093] Specifically, if the pixel density inside the ROI is high, the outline is single, and the area is close to the typical single tree crown area, it is determined as single or clump vegetation; if the ROI is long and distributed along the edge of the road, it is determined as multiple row of street trees; if the area of the ROI is less than the single tree threshold and the height is insufficient, it is determined as low shrub.
[0094] If it is single or clump, the corrected scale Map the ROI centroid position to the actual plane coordinate; if it is a multi-plant or cluster case, extract the main direction line of the ROI, that is, the road direction of the minimum circumscribed rectangle main axis, and sort it according to the total length Typical plant distance in knowledge graph Estimate the number of plants
[0095] ; (5)
[0096] Among them, Indicates rounding.
[0097] And generate each plant position coordinate equidistantly or with random disturbance between the starting point and the ending point.
[0098] If it is a shrub belt, it is estimated according to the typical area density in the knowledge graph, and random points are scattered in the ROI and the minimum distance is avoided by constraint to avoid excessive aggregation.
[0099] The placement scheme must meet the physical and ecological constraints, including minimum distance limit, collision detection, and visual balance and minimum occlusion. Each placement checks the overlap rate with the placed instances, and tries to fine-tune the grid body that does not meet the requirements.
[0100] Further, after determining the placement scheme, the model is imported into the scene. Taking CARLA scene import as an example, the CARLA Python API interface is called to create Actors in batches. Load the custom blueprint converted from FBX or OBJ, and call world.spawn_actor(blueprint, transform) to instantiate the model for each identified region, wherein the transform parameter includes position, size, orientation, and coordinate transformation module, which maps the inferred two-dimensional plane coordinate to the CARLA local coordinate system , using affine transformation:
[0101] ; (6)
[0102] Among them, Indicates the scaling factor, , indicates the rotation matrix, is the offset of the reference origin of the simulation platform, is the terrain grid translation.
[0103] Material and map are assigned in batches through the set_attribute('texture', asset_path) interface of the blueprint, which automatically applies the matched high-resolution map to the three-dimensional model (i.e. the three-dimensional model instantiated in the simulation platform).
[0104] Through the above method, batch import and initial placement of models in a three-dimensional scene can be quickly completed, and the construction efficiency and controllability of a city traffic digital twin scene are greatly improved.
[0105] The present application aims at the problems of existing map details missing and low modeling efficiency, batch collects materials, vegetation recognition based on deep semantic segmentation and fine-grained classification, material matching of mapping + vector retrieval, import scheme generation, and realizes an integrated pipeline from street view images to three-dimensional scenes.
[0106] Further, as shown in Figure 5 Based on the above street view recognition-based automatic generation method of urban traffic three-dimensional model vegetation elements, the present application also correspondingly provides a street view recognition-based automatic generation system of urban traffic three-dimensional model vegetation elements, wherein the street view recognition-based automatic generation system of urban traffic three-dimensional model vegetation elements comprises:
[0107] A knowledge graph and material library construction module 51 is used to construct a vegetation element knowledge graph for a target area, define vegetation types and typical attributes, directionally collect three-dimensional models and multi-view images, and store unified metadata in a material library;
[0108] A vegetation recognition module 52 is used to perform pixel-level semantic segmentation on street view images, identify vegetation categories based on transfer learning and attention mechanisms, and output ROIs and feature vectors;
[0109] A material matching module 53 is used to filter a candidate set based on ontology mapping, use a cosine similarity vector retrieval method to select the most matched vegetation model for the local feature vector of the identified ROI and the pre-stored feature vector in the material library, and record the mapping;
[0110] A model import module 54 is used to correct based on a reference, fuse with a knowledge graph, automatically identify or optimize strategies, correct the reliable scale of vegetation ROIs, infer the quantity or position, and batch import three-dimensional models that meet ecological constraints into a target simulation platform or a visualization platform.
[0111] Further, as shown in Figure 6 Based on the above street view recognition-based automatic generation method and system of urban traffic three-dimensional model vegetation elements, the present application also correspondingly provides a terminal, which comprises a processor 10, a memory 20, and a display 30. Figure 6 Only part of the components of the terminal are shown, but it should be understood that all the shown components are not required to be implemented, and more or fewer components can be alternatively implemented.
[0112] The memory 20 can be an internal storage unit of the terminal in some embodiments, such as a hard disk or a memory of the terminal. The memory 20 can also be an external storage device of the terminal in other embodiments, such as a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. equipped on the terminal. Further, the memory 20 can include both the internal storage unit and the external storage device of the terminal. The memory 20 is used to store application software and various data installed on the terminal, such as program codes of the terminal, etc. The memory 20 can also be used to temporarily store data that has been output or will be output. In an embodiment, the memory 20 stores a program for automatically generating a vegetation element of a three-dimensional model of urban traffic based on street view recognition 40, which can be executed by the processor 10 to implement the method for automatically generating a vegetation element of a three-dimensional model of urban traffic based on street view recognition in the present application.
[0113] The processor 10 can be a Central Processing Unit (CPU), a microprocessor or other data processing chip in some embodiments, which is used to run program codes or process data stored in the memory 20, such as to execute the method for automatically generating a vegetation element of a three-dimensional model of urban traffic based on street view recognition, etc.
[0114] The display 30 can be an LED display, a liquid crystal display, a touch liquid crystal display, an OLED (Organic Light-Emitting Diode) touch, etc. in some embodiments. The display 30 is used to display information of the terminal and to display a visualized user interface. The processor 10, the memory 20 and the display 30 of the terminal communicate with each other through a system bus.
[0115] In an embodiment, the processor 10 implements the steps of the method for automatically generating a vegetation element of a three-dimensional model of urban traffic based on street view recognition as described above when executing the program for automatically generating a vegetation element of a three-dimensional model of urban traffic based on street view recognition 40 in the memory 20.
[0116] The present application also provides a computer readable storage medium, wherein the computer readable storage medium stores a program for automatically generating a vegetation element of a three-dimensional model of urban traffic based on street view recognition, which implements the steps of the method for automatically generating a vegetation element of a three-dimensional model of urban traffic based on street view recognition as described above when executed by a processor.
[0117] In summary, the present application provides a method, system, terminal and computer readable storage medium for automatically generating vegetation elements of a city traffic three-dimensional model based on street view recognition, the method comprising: constructing a vegetation element knowledge graph for a target area, defining vegetation types and typical attributes, directionally collecting three-dimensional models and multi-view images, and storing metadata in a material library; performing pixel-level semantic segmentation on street view images, identifying vegetation categories based on transfer learning and attention mechanisms, and outputting ROIs and feature vectors; filtering a candidate set based on ontology mapping, using a cosine similarity vector retrieval method to select the most matching vegetation model for the local feature vector of the identified ROI and the pre-stored feature vector in the material library, and recording the mapping; based on reference correction, knowledge graph fusion, automatic discrimination or optimization strategy, correcting the reliable scale of the vegetation ROI, inferring the quantity or position, and importing the batch three-dimensional models that meet the ecological constraints into the target simulation platform or visualization platform. The present application realizes the integration pipeline from street view images to three-dimensional scenes through the integrated process of material collection standardization, semantic recognition, vector matching and programmed generation, and provides efficient, intelligent and standardized technical support for traffic simulation and city visualization.
[0118] It should be noted that in this document, the terms "comprising", "containing", or any other variant thereof are intended to cover a non-exclusive inclusion, such that a process, method, article or terminal that comprises a list of elements does not only include those elements, but also other elements not explicitly listed, or other elements inherent to such a process, method, article or terminal. Without more limitations, the element defined by the phrase "comprising a" does not exclude the presence of additional identical elements in the process, method, article or terminal that includes the element.
[0119] Of course, those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing relevant hardware (such as a processor, a controller, etc.) through a computer program, and the program can be stored in a computer readable computer readable storage medium, and the program can include the processes of the above-mentioned method embodiments when executed. The computer readable storage medium can be a memory, a disk, an optical disk, etc.
[0120] It should be understood that the application is not limited to the above examples, and those skilled in the art can make improvements or changes according to the above description, and all such improvements and changes shall fall within the scope of protection of the claims of the present application.
Claims
1. A method for automatic generation of vegetation elements of a three-dimensional model of urban traffic based on street view recognition, characterized in that, The method comprises the following steps: Construct a vegetation element knowledge graph for a target area, define vegetation types and typical attributes, and collect three-dimensional models and multi-view images in a targeted manner, and store the unified metadata in a material library; Perform pixel-level semantic segmentation on the street view images, identify vegetation categories based on transfer learning and attention mechanisms, and output ROIs and feature vectors; Filter candidate sets based on ontology mapping, use the cosine similarity vector retrieval method to select the most matching vegetation model for the local feature vector of the identified ROI from the pre-stored feature vectors in the material library, and record the mapping; Based on the reference correction, knowledge graph fusion, and automatic discrimination or optimization strategy, correct the reliable scale of the vegetation ROI, infer the quantity or position, and import the batch of three-dimensional models that meet the ecological constraints into the target simulation platform or visualization platform; The method comprises the following steps: Determine the target area, construct a vegetation element knowledge graph for the target area, determine the common vegetation category hierarchy and attributes of the target area, and use the vegetation element knowledge graph as a semantic reference ontology for collection and retrieval; Based on the categories and attributes in the vegetation element knowledge graph, use a targeted crawler and map or data interface to batch collect and filter vegetation-related three-dimensional materials and multi-view images, and record the source, license information, and multi-view information; Format and metadata standardize the collected three-dimensional materials, unify the naming conventions, generate metadata, and organize the material library according to the semantic directory for retrieval and version management; Perform format checking and automatic repair on the original three-dimensional models and textures, batch convert formats, check and automatically repair existing problems, and label samples that do not meet the quality standards; The method comprises the following steps: For each vegetation ROI, determine its size in the real space through the map scale information, correct the estimation result based on the typical scale of the reference object, and obtain the scaling factor; For each ROI, automatically extract the observed features based on the segmentation mask, determine the planted type of the ROI based on the knowledge graph prior, and determine the number and size interval of the placed vegetation; Based on the vegetation model and size or quantity estimation, perform scale and coordinate transformation, construct the affine coordinate transformation required by the simulation platform, and batch import and instantiate through the simulation interface, and simultaneously batch assign textures and material parameters through the texture interface; After the matched three-dimensional models are instantiated and imported into the target simulation platform or visualization platform, perform automatic acceptance inspection, label abnormal instances in the metadata, and trigger the manual or remapping process, and finally retain the complete mapping link.
2. The method of claim 1, wherein the method further comprises: determining a height of the vegetation element based on the height of the building and the height of the street light; and determining a width of the vegetation element based on the width of the building and the width of the street light. The street view image is subjected to pixel-level semantic segmentation, vegetation categories are identified based on a transfer learning and an attention mechanism, and an ROI and a feature vector are output, specifically including: The street view image is subjected to pixel-level semantic segmentation, vegetation categories are identified based on a transfer learning and an attention mechanism, and an ROI and a feature vector are output, specifically including: The street view image is subjected to pixel-level semantic segmentation, vegetation categories are identified based on a transfer learning and an attention mechanism, and an ROI and a feature vector are output, specifically including: The fine-grained classifier outputs a semantic label, a local confidence score, and a corresponding local visual feature vector of each identified object.
3. The method of claim 1, wherein the method further comprises: determining a height of the vegetation element based on the height of the building and the height of the street light; and determining a width of the vegetation element based on the width of the building and the width of the street light. The fine-grained classifier is based on a transfer learning strategy, uses a deep backbone network pre-trained on a large-scale dataset as a basis, and combines an attention mechanism to fine-tune to a local vegetation dataset to enhance the sensitivity to texture and morphological features.
4. The method of claim 1, wherein the method further comprises: determining a height of the vegetation element based on the height of the building and the height of the street light; and determining a width of the vegetation element based on the width of the building and the width of the street light. The candidate set is filtered based on ontology mapping, a cosine similarity vector retrieval method is used to select the most matched vegetation model from the local feature vector of the identified ROI and the pre-stored feature vector in the material library, and the mapping is recorded, specifically including: According to the correspondence between the predefined categories and subcategories in the ontology, candidate resources of the same category are filtered out from the material library, and preliminary filtering is performed according to the size range and version constraints; Similarity calculation is performed on the deep feature vector extracted from the identified area and the candidate material feature vector, and the K candidates in the sorting result are taken as the final mapping candidate list in descending order of similarity; The K candidates are subjected to rule-based priority determination, and the final determined material ID and the corresponding texture file paths of the map, normal and roughness are output.
5. The method of claim 1, wherein the method further comprises: determining a height of the vegetation element based on the height of the building and the height of the street light; and determining a width of the vegetation element based on the width of the building and the width of the street light. The automatic acceptance check includes collision body consistency check, map missing check, and occlusion or visibility check.
6. A system for automatic generation of vegetation elements of a three-dimensional model of urban traffic based on street recognition, characterized in that, The street view recognition-based urban traffic three-dimensional model vegetation element automatic generation system is used to implement the street view recognition-based urban traffic three-dimensional model vegetation element automatic generation method of any one of claims 1-5, and the street view recognition-based urban traffic three-dimensional model vegetation element automatic generation system comprises: A knowledge graph and material library construction module is configured to construct a vegetation element knowledge graph for a target area, define vegetation types and typical attributes, and collect three-dimensional models and multi-view images in a targeted manner, and store unified metadata in a material library; A vegetation identification module is configured to perform pixel-level semantic segmentation on street view images, identify vegetation categories based on a transfer learning and an attention mechanism, and output an ROI and a feature vector; A material matching module is configured to filter a candidate set based on ontology mapping, and use a cosine similarity vector retrieval method to select the most matched vegetation model from the local feature vector of the identified ROI and the pre-stored feature vector in the material library, and record the mapping; A model import module is configured to correct based on a reference, fuse a knowledge graph, and automatically determine or optimize strategies, correct the reliable scale of a vegetation ROI, infer the quantity or position of the vegetation ROI, and import a batch of three-dimensional models that meet ecological constraints to a target simulation platform or a visualization platform.
7. A terminal, characterized by comprising: The terminal comprises a memory, a processor, and a street view recognition based automatic generation program of urban traffic three-dimensional model vegetation elements stored on the memory and capable of running on the processor, and the street view recognition based automatic generation program of urban traffic three-dimensional model vegetation elements, when executed by the processor, implements the steps of the street view recognition based automatic generation method of urban traffic three-dimensional model vegetation elements according to any one of claims 1-5.
8. A computer-readable storage medium, characterized in that, The computer readable storage medium stores a street view recognition based automatic generation program of urban traffic three-dimensional model vegetation elements, and the street view recognition based automatic generation program of urban traffic three-dimensional model vegetation elements, when executed by the processor, implements the steps of the street view recognition based automatic generation method of urban traffic three-dimensional model vegetation elements according to any one of claims 1-5.
Citation Information
Patent Citations
Three-dimensional parametric modeling method for beam bridge
CN110263376A
Scene graph generation method and system supporting historical and cultural block scene
CN118334414A