Urban traffic three-dimensional model vegetation element automatic generation method and system based on streetscape recognition, terminal and storage medium
By constructing a knowledge graph of vegetation elements and pixel-level semantic segmentation, combined with cosine similarity vector retrieval and reference object correction, the problem of missing vegetation elements in 3D scenes was solved, enabling efficient and intelligent automated import of vegetation models and improving the fidelity and modeling efficiency of 3D scenes.
Patent Information
- Application Number
- CN202511501434.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-21
- Publication Date
- 2025-11-18
- Estimated Expiration
- 2045-10-21
AI Technical Summary
Existing technologies lack vegetation elements when generating 3D scenes, resulting in insufficient fidelity, low modeling efficiency, inconsistent styles, and difficulty in standardization and reuse.
By constructing a knowledge graph of vegetation elements, pixel-level semantic segmentation and transfer learning are used to identify vegetation categories. Combined with cosine similarity vector retrieval and reference object correction, the automated import and ecological constraints of vegetation models are achieved.
It enables efficient and intelligent automatic generation of vegetation elements in 3D scenes, improving the fidelity and modeling efficiency of 3D scenes, and supporting large-scale, rapid and reusable construction of traffic 3D scenes.
Smart Images

Figure CN120976448A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of urban simulation and 3D modeling technology, and in particular to a method, system, terminal, and computer-readable storage medium for automatically generating vegetation elements in urban traffic 3D models based on street view recognition. Background Technology
[0002] With the increasing complexity of urban transportation systems and the promotion of technologies such as autonomous driving and vehicle-to-infrastructure (V2I) communication, high-fidelity 3D simulation environments have become essential tools for research and development, testing, and decision-making evaluation. The demand for algorithm verification, strategy evaluation, and emergency response plan testing in a controllable and safe virtual environment is constantly growing. Simultaneously, urban planning and digital management require visualized and analyzable 3D scenes as a medium for policy formulation and public communication. Therefore, the ability to generate highly realistic, scalable, and standardized 3D scenes has significant social value and market demand.
[0003] Currently, existing technologies have the following main shortcomings: (1) Missing data elements: Existing high-precision maps usually do not include information such as roads and vegetation. These street scene elements are crucial for visual simulation and scene aesthetics.
[0004] (2) Low simulation efficiency and poor reusability: Current scene modeling can only complete the batch generation of simple geometry, which relies heavily on manual drawing and manual adjustment. A large number of models need to be imported and corrected one by one, making it difficult to achieve accurate matching at the level of detail. The workload is huge and it is difficult to maintain consistency of types. At the same time, it is not conducive to the formation of a reusable component library and standardized pipeline.
[0005] Therefore, existing technologies still need to be improved and developed. Summary of the Invention
[0006] The main objective of this invention is to provide a method, system, terminal, and computer-readable storage medium for automatically generating vegetation elements in a 3D urban traffic model based on street view recognition. This aims to solve the problems of missing elements, high labor costs, difficulty in material integration, and insufficient automation capabilities in the existing technology when generating 3D scenes.
[0007] To achieve the above objectives, this invention provides a method for automatically generating vegetation elements in a 3D urban traffic model based on street view recognition. The method includes the following steps: Construct a vegetation element knowledge graph for the target area, define vegetation types and typical attributes, collect 3D models and multi-view images in a targeted manner, unify metadata and store it in the material library; Pixel-level semantic segmentation is performed on street view images, vegetation categories are identified based on transfer learning and attention mechanisms, and ROI and feature vectors are output. Based on ontology mapping, candidate sets are filtered, and the cosine similarity vector retrieval method is used to select the vegetation model that best matches the local feature vector of the identified ROI with the pre-stored feature vector in the material library, and the mapping is recorded. Based on reference correction, knowledge graph fusion, and automatic discrimination or optimization strategies, reliable scale correction, quantity or location inference of vegetation ROI, and batch import of ecologically constrained 3D models are placed into the target simulation platform or visualization platform.
[0008] Optionally, the method for automatically generating vegetation elements in a 3D urban traffic model based on street view recognition, wherein the step of constructing a vegetation element knowledge graph for the target area, defining vegetation types and typical attributes, selectively collecting 3D models and multi-view images, and unifying metadata and storing it in a material library, specifically includes: The target area is identified, a vegetation element knowledge graph for the target area is constructed, the hierarchical structure and attributes of common vegetation categories in the target area are determined, and the vegetation element knowledge graph is used as a semantic reference ontology for collection and retrieval. Based on the categories and attributes in the vegetation element knowledge graph, targeted crawlers and maps or data interfaces are used to collect and filter 3D materials and multi-view images related to vegetation in batches, and record the source, permission information and multi-view information. The acquired 3D materials are formatted and their metadata is standardized, a unified naming convention is established, metadata is generated, and the material library is organized according to semantic directories for retrieval and version management. Perform format verification and automatic repair on the original 3D models and textures, batch convert formats, check and automatically repair existing problems, and annotate samples that do not meet the quality standards.
[0009] Optionally, the method for automatically generating vegetation elements in a 3D urban traffic model based on street view recognition, wherein the step of performing pixel-level semantic segmentation on the street view image, identifying vegetation categories based on transfer learning and attention mechanisms, and outputting ROI and feature vectors specifically includes: Pixel-level semantic segmentation is used on street view images to locate vegetation regions (ROIs), and the region mask and local features are output to provide input samples for fine-grained recognition. Based on segmentation, a fine-grained classifier is trained to identify ROI vegetation categories; The fine-grained classifier outputs a semantic label, a local confidence score, and a corresponding local visual feature vector for each identified object.
[0010] Optionally, in the method for automatically generating vegetation elements of a three-dimensional urban traffic model based on street view recognition, the fine-grained classifier is based on a transfer learning strategy, using a deep backbone network pre-trained on a large-scale dataset as a foundation, and combining an attention mechanism to fine-tune it to a local vegetation dataset to enhance its sensitivity to texture and morphological features.
[0011] Optionally, the method for automatically generating vegetation elements in a 3D urban traffic model based on street view recognition, wherein the step of filtering the candidate set based on ontology mapping and using a cosine similarity vector retrieval method to select the most matching vegetation model between the local feature vector of the identified ROI and the pre-stored feature vector in the material library, and recording the mapping, specifically includes: Based on the predefined category and subclass correspondence in the ontology, candidate resources of the same type are selected from the material library, and preliminary filtering is performed based on size range and version constraints; The similarity between the deep feature vector extracted from the recognition region and the feature vector of the candidate material is calculated, and the similarity is sorted from high to low. The top K candidates in the sorting results are taken as the final mapping candidate list. Perform rule-based priority determination on K candidates, and output the final determined material ID and the corresponding material file paths for texture, normal, and roughness.
[0012] Optionally, the method for automatically generating vegetation elements in a 3D urban traffic model based on street view recognition, wherein the reliable scale correction, quantity or location inference of vegetation ROI, and the import of batch 3D models satisfying ecological constraints based on reference correction, knowledge graph fusion, and automatic discrimination or optimization strategies are placed into the target simulation platform or visualization platform, specifically includes: For each vegetation ROI, the size of each vegetation ROI in real space is determined by the map scale information, and the estimation results are corrected according to the typical scale of the reference object to obtain the scaling factor; For each ROI, observation features are automatically extracted based on segmentation masks, and the planting type of the ROI is determined by combining prior knowledge graphs to determine the number and size range of vegetation to be placed. Based on the vegetation model and size or quantity estimation, scale and coordinate transformations are performed to construct the affine coordinate transformation required for the simulation platform. The transformations are then imported and instantiated in batches through the simulation interface. Simultaneously, texture and material parameters are assigned in batches through the texture interface. After the matched 3D model is instantiated and imported into the target simulation platform or visualization platform, an automated acceptance check is performed. Abnormal instances are marked in the metadata and a manual or remapping process is triggered, ultimately preserving the complete mapping chain.
[0013] Optionally, in the method for automatically generating vegetation elements of a 3D urban traffic model based on street view recognition, the automated acceptance inspection includes collision body consistency inspection, texture missing inspection, and occlusion or visibility inspection.
[0014] Furthermore, to achieve the above objectives, the present invention also provides an automatic generation system for vegetation elements of a three-dimensional urban traffic model based on street view recognition, wherein the automatic generation system for vegetation elements of a three-dimensional urban traffic model based on street view recognition includes: The knowledge graph and material library construction module is used to build a vegetation element knowledge graph for the target area, define vegetation types and typical attributes, collect 3D models and multi-view images in a targeted manner, unify metadata and store it in the material library. The vegetation recognition module is used to perform pixel-level semantic segmentation on street view images, identify vegetation categories based on transfer learning and attention mechanisms, and output ROI and feature vectors. The material matching module is used to filter the candidate set based on ontology mapping. It uses the cosine similarity vector retrieval method to select the most matching vegetation model between the local feature vector of the identified ROI and the pre-stored feature vector in the material library, and records the mapping. The model import module is used to import batch 3D models of vegetation ROIs into the target simulation platform or visualization platform based on reference correction, knowledge graph fusion and automatic discrimination or optimization strategies, including reliable scale correction, quantity or location inference and ecological constraints.
[0015] Furthermore, to achieve the above objectives, the present invention also provides a terminal, wherein the terminal includes: a memory, a processor, and an automatic generation program for urban traffic 3D model vegetation elements based on street view recognition, which is stored in the memory and can run on the processor. When the automatic generation program for urban traffic 3D model vegetation elements based on street view recognition is executed by the processor, it implements the steps of the automatic generation method for urban traffic 3D model vegetation elements based on street view recognition as described above.
[0016] Furthermore, to achieve the above objectives, the present invention also provides a computer-readable storage medium, wherein the computer-readable storage medium stores an automatic generation program for vegetation elements of a three-dimensional urban traffic model based on street view recognition, and when the automatic generation program for vegetation elements of a three-dimensional urban traffic model based on street view recognition is executed by a processor, it implements the steps of the automatic generation method for vegetation elements of a three-dimensional urban traffic model based on street view recognition as described above.
[0017] In this invention, a vegetation element knowledge graph for the target area is constructed, defining vegetation types and typical attributes. 3D models and multi-view images are collected in a targeted manner, and metadata is unified and stored in a material library. Pixel-level semantic segmentation is performed on street view images, and vegetation categories are identified based on transfer learning and attention mechanisms, outputting ROIs and feature vectors. A candidate set is filtered based on ontology mapping, and a cosine similarity vector retrieval method is used to select the most matching vegetation model between the local feature vectors of the identified ROI and the pre-stored feature vectors in the material library, recording the mapping. Based on reference object correction, knowledge graph fusion, and automatic discrimination or optimization strategies, reliable scale correction, quantity or location inference of vegetation ROIs, and batch import of ecologically constrained 3D models are placed into the target simulation platform or visualization platform. This invention, through an integrated process of standardized material collection, semantic recognition, vectorized matching, and procedural generation, realizes an integrated pipeline from street view images to 3D scenes, providing efficient, intelligent, and standardized technical support for traffic simulation and urban visualization. Attached Figure Description
[0018] Figure 1 This is a flowchart of a preferred embodiment of the method for automatically generating vegetation elements in a three-dimensional urban traffic model based on street view recognition according to the present invention; Figure 2 This is a schematic diagram illustrating the process of generating vegetation elements in a three-dimensional urban traffic model based on street view recognition, according to a preferred embodiment of the present invention. Figure 3 This is a schematic diagram of vegetation material in a preferred embodiment of the method for automatically generating vegetation elements in a three-dimensional urban traffic model based on street view recognition according to the present invention. Figure 4 This is a generated effect diagram in a preferred embodiment of the method for automatically generating vegetation elements in a three-dimensional urban traffic model based on street view recognition according to the present invention. Figure 5 This is a structural diagram of a preferred embodiment of the urban traffic 3D model vegetation element automatic generation system based on street view recognition of the present invention; Figure 6 This is a structural diagram of a preferred embodiment of the terminal of the present invention. Detailed Implementation
[0019] To make the objectives, technical solutions, and advantages of this invention clearer and more explicit, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention.
[0020] With the development of urbanization and intelligent transportation, high-precision traffic simulation and 3D scenes have become crucial infrastructure in various fields such as autonomous driving verification, traffic management, urban planning, and emergency drills. However, existing high-precision maps and modeling processes generally lack key elements such as vegetation types, resulting in insufficient 3D scene fidelity and difficulty in meeting high-fidelity simulation requirements. Furthermore, traditional modeling heavily relies on manual operation, leading to low efficiency, inconsistent styles, and difficulty in large-scale replication. Inconsistent naming and parameters across different material sources also cause numerous compatibility issues for batch processing and script-based integration. Therefore, a standardized and intelligent automated pipeline from street view or map data to 3D scenes is needed to support large-scale, rapid, and reusable 3D traffic scene construction.
[0021] To address one or more of the above-mentioned problems, this invention constructs a vegetation element knowledge graph for target areas, defines vegetation categories and their typical attributes, collects 3D models and multi-view images in a targeted manner, unifies metadata and stores it in a database; performs pixel-level semantic segmentation on street view images and identifies vegetation categories and outputs ROIs and feature vectors based on transfer learning and attention mechanisms; filters candidate sets based on ontology mapping, uses a cosine similarity vector retrieval method to select the most matching vegetation model between the local feature vectors of the identified ROI and the pre-stored feature vectors in the material library, and records the mapping; through reference object correction, knowledge graph fusion and automatic discrimination or optimization strategies, it achieves reliable scale correction, quantity or location inference of vegetation ROIs, and batch import and placement of 3D models that meet ecological constraints.
[0022] The preferred embodiment of the present invention describes an automatic generation method for vegetation elements in a 3D urban traffic model based on street view recognition. Figure 1 and Figure 2 As shown, the method for automatically generating vegetation elements in a 3D urban traffic model based on street view recognition includes the following steps: Step S10: Construct a vegetation element knowledge graph for the target area, define vegetation types and typical attributes, collect 3D models and multi-view images in a targeted manner, unify metadata and store it in the material library.
[0023] Specifically, the target area is identified, and a vegetation element knowledge graph for the target area is constructed. The hierarchical structure and attributes of common vegetation categories in the target area are determined, and this vegetation element knowledge graph serves as a semantic reference ontology for subsequent data collection and retrieval. Based on the categories and attributes in the vegetation element knowledge graph, targeted crawlers and maps or data interfaces are used to batch collect and filter vegetation-related 3D materials and multi-view images, recording the source, license information, and multi-view information. The collected 3D materials are formatted and their metadata is standardized, with unified naming conventions, metadata generation, and a material library organized according to a semantic directory for retrieval and version management. Format verification and automatic repair are performed on the original 3D models (referring to the original source versions of the collected 3D model files, before format standardization) and textures (in 3D modeling, textures refer to 2D images mapped onto the surface of a 3D model, used to represent surface color, details, material properties, etc.). Format conversion is performed in batches, and existing problems are checked and automatically repaired (e.g., automatic repair of normal direction, non-manifold geometry, UV). Common issues such as overlapping and missing texture links are addressed, and samples that do not meet quality standards are labeled (e.g., labeled as "awaiting manual review").
[0024] Specifically, an ontology engineering approach is used to construct a vegetation element knowledge graph. Through expert annotation and literature or existing databases, the hierarchical relationships (class → genus → species or class → subclass, etc.) of common vegetation categories in a given region are analyzed. Key attributes are defined for each node: typical crown width range, typical height range, seasonal characteristics, typical leaf shape or color, typical growth density, and ecological constraints. The knowledge graph is implemented using an exchangeable ontology format (OWL or RDF), and versioning and annotations are maintained in a construction tool (e.g., Protecte) to serve as a semantic reference for crawler prioritization and subsequent retrieval.
[0025] Based on the aforementioned vegetation element knowledge graph, a targeted crawler framework was designed and deployed to collect materials and street view images in batches. The crawler uses vegetation categories and attribute keywords listed in the ontology as search seeds, performing semantic-priority crawling on target material websites and multi-view preview pages. For example, based on a survey of common Unreal Engine formats (FBX, OBJ, or GLB) on the Aigei material resource website, a targeted crawler framework was designed; during crawling, it automatically retrieves information such as the "name," "download link," "License," and multi-view preview images from model pages, and downloads files in mainstream formats.
[0026] Furthermore, metadata for all original models and textures is standardized. Referring to the metadata framework of the resource library, a unified JSON description file is generated, with fields covering id, name, source website, license, category, subclass, number of triangles, texture resolution, model dimensions (length × width × height), upload time, version number, and keyword tags, facilitating subsequent retrieval and traceability. All model files are renamed according to the category_subclass_source_version.fbx specification.
[0027] During the format verification and automatic repair phase, the Blender Python API is used to batch convert non-FBX formats (3DS or DAE) to FBX (Filmbox). The model is then checked for normal direction, non-manifold geometry, UV overlap (multiple faces occupying the same area in the texture coordinate system can cause issues such as texture misalignment and pattern stretching), and missing texture links. Minor geometric errors are automatically reconstructed and blank UVs are filled; serious errors are marked as "awaiting manual review."
[0028] After file verification, the original materials undergo quality screening, removing blurry, distorted, overexposed, or incompletely composed samples to ensure that the images entering the database have clear details of tree crowns and branches. This process ensures that all models are presented under the same lighting and background conditions, guaranteeing visual consistency in the subsequent scene. Ontology mapping is then performed on the keyword tags of each screened material to generate semantic fields consistent with the target layer elements.
[0029] Simultaneously, street view map data is collected, segmented information of the target area is extracted from the map, sampling points along both sides of the road are obtained at set intervals, and multi-viewpoint images or panoramic snapshots of the corresponding perspectives are obtained by calling the street view image interface based on the sampling points.
[0030] For each captured street view image, map-related metadata is recorded simultaneously, including latitude and longitude, road ID, tile_id or tile coordinates, map zoom level, map resolution, camera orientation or tilt angle, shooting time, etc., and a unified JSON description file is generated after preprocessing.
[0031] Step S20: Perform pixel-level semantic segmentation on the street view image, identify vegetation categories based on transfer learning and attention mechanisms, and output ROI and feature vector.
[0032] Specifically, pixel-level semantic segmentation is used on street view images to locate vegetation regions (ROIs), and outputs region masks and local features (the region mask is a binary image of the ROI location, where each pixel indicates whether it belongs to vegetation; the local features are feature vectors, which are descriptive information extracted from the ROI region, derived from the deep features output by the intermediate layer of the neural network, including texture, shape, color, etc.), providing input samples for fine-grained recognition; based on the segmentation, a fine-grained classifier is trained to identify the vegetation category of the ROI.
[0033] Specifically, a vegetation image dataset related to the target region is constructed, uniformly preprocessed and augmented, and proportionally divided into training, validation, and test sets to evaluate Top-1 accuracy and robustness (Top-1 accuracy is used to evaluate the model's accuracy in identifying vegetation categories; Top-1 refers to whether the category with the highest predicted probability is the true category, and Top-5 refers to whether the true category is within the top 5 predicted probabilities). The fine-grained classifier is based on a transfer learning strategy, using a deep backbone network pre-trained on a large-scale dataset as a foundation, combined with an attention mechanism to fine-tune it to the local vegetation dataset to enhance sensitivity to texture and morphological features, improve cross-domain feature alignment, and enhance recognition accuracy under small sample sizes. The fine-grained classifier outputs the semantic label, local confidence score, and corresponding local visual feature vector for each identified object.
[0034] Specifically, the street view RGB images to be processed are preprocessed according to a unified standard: first, they are centered or scaled proportionally according to their aspect ratio, and then normalized to the ImageNet mean and standard deviation; images used for pixel-level semantic segmentation are uniformly adjusted to 1024×2048 to balance segmentation accuracy and computational cost; ROI cropped images used for fine-grained classification are uniformly scaled to 224×224 pixels in subsequent steps to adapt to the input requirements of the fine-grained classifier. The georeference of the original images is also recorded during preprocessing for subsequent coordinate mapping.
[0035] For localized vegetation types, supplement the image dataset required for fine-grained recognition. Divide the preprocessed images into training, validation, and test sets in an 8:1:1 ratio. The training set is used for network training and iterative optimization; the test set is used to evaluate the final model's Top-1 Accuracy (a classification metric, the proportion of the highest predicted probability class equal to the true class) on a 20-class local vegetation recognition task; and the validation set is used to verify the network's performance and generalization ability.
[0036] Based on the DeepLabV3+ native architecture, pixel-level semantic segmentation is performed on key elements in street view map images, and feature fusion is performed during the semantic segmentation process of street view elements in complex scenes. By accurately defining the Region of Interest (ROI) of various elements, input is provided for downstream recognition and detection tasks. The backbone network (referring to the main convolutional network used to extract image features) is a ResNet-101 pre-trained on ImageNet.
[0037] DeepLabV3+'s semantic segmentation network first performs a series of downsampling and dilated convolution operations on the input RGB street view image to obtain rich high-level semantic features. Then, it generates five features in parallel (dilation rates = 1, 6, 12, 18, 24) through the Atrous Spatial Pyramid Pooling (ASPP) module, and then reduces and fuses them through 1×1 convolution. After being fused with shallow features in the Decoder, it is upsampled to the original image size through bilinear interpolation, and outputs the probability distribution of each pixel.
[0038] The loss function is pixel-level cross-entropy, and a learning rate decay strategy is employed. The optimizer uses SGD+Momentum (momentum=0.9, weight_decay=1e-4). SGD (Stochastic Gradient Descent) is a basic deep learning optimization algorithm. Momentum is the momentum mechanism, and weight_decay is the weight decay. That is, when using stochastic gradient descent, a momentum of 0.9 is introduced to accelerate convergence and reduce oscillations, while a decay of 1e-4 is added to the weight parameters to prevent overfitting.
[0039] The trained DeepLabV3+ semantic segmentation model was tested. DeepLabV3+ pre-trained weights were loaded, and forward inference was performed on a single image at full resolution to obtain a classification pixel probability map. Post-processing techniques such as DenseCRF (Conditional Random Field post-processing based on pixel and color or position similarity to refine segmentation boundaries) and connected component filtering were then applied to refine edges, remove noise, and improve the coherence of segmentation boundaries. Segmentation performance was monitored using mIoU (Mean Intersection over Union) and pixel accuracy (PA).
[0040] After completing semantic segmentation based on DeepLabV3+ and locating the ROI of vegetation elements, these regions are segmented, and semantic labels, region masks, and local features are output to provide input samples for subsequent fine-grained recognition.
[0041] Furthermore, a lightweight vegetation category detection model was trained to output the attribute labels and spatial locations corresponding to each element. The foundational network for the lightweight vegetation category detection model was determined by selecting different architectures of deep convolutional neural networks: DenseNet, Inception, ResNeXt, and MobileNet. These networks were trained on the image training dataset, and the trained models were used to perform inference on the test dataset. The performance of these models was evaluated using the Top1-ACC (Top1 accuracy). Simultaneously, the pre-trained model was fine-tuned and used to make predictions on a self-built dataset, and the recognition performance of the fine-tuned network model transferred from ImageNet was compared. The ResNet50 network, which achieved the highest Top1-ACC accuracy, was selected as the backbone network for vegetation recognition.
[0042] The training effect of local datasets on network parameters may be limited by the size of the sample. Therefore, based on large datasets such as ImageNet, fine-tuning is performed by transferring the pre-trained structure and parameters to the local dataset.
[0043] Furthermore, an attention mechanism is added to the basic network model to improve recognition accuracy. Training and testing are performed on the dataset using a "channel-first, spatial-second" CBAM (Convolutional Block Attention Module, a lightweight attention mechanism that sequentially applies channel and spatial attention to feature maps) attention mechanism. This fully considers the combination of channel and spatial domain attention, and adds average pooling and convolution operations. The optimized network is then used to select the best-matching candidate materials.
[0044] For each vegetation ROI, two parallel feature vectors are extracted: one from the intermediate layer of the semantic segmentation network; and the other from the penultimate layer of the fine-grained classifier. The extracted raw floating-point vectors are L2 normalized and PCA is performed to reduce storage and retrieval overhead; simultaneously, a set of simple color or texture statistics are generated as auxiliary features. All features are written in a uniform format to the corresponding ROI's JSON metadata field and simultaneously written to the vector database for efficient approximate nearest neighbor retrieval.
[0045] Step S30: Based on ontology mapping, filter the candidate set, and use the cosine similarity vector retrieval method to select the vegetation model that best matches the local feature vector of the identified ROI with the pre-stored feature vector in the material library, and record the mapping.
[0046] Specifically, based on the predefined category-subclass correspondence in the ontology, candidate resources of the same type are selected from the material library, and preliminary filtering is performed according to size range and version constraints; the similarity between the depth feature vector extracted from the recognition region and the feature vector of the candidate material is calculated, and the similarity is sorted from high to low. The top K candidates in the sorting result are taken as the final mapping candidate list (that is, taking the Top-K candidates as the final mapping candidate list means taking the top K with the highest similarity in the sorting result as the candidate set); rule-based priority determination is performed on the K candidates (i.e., Top-K candidates), and the final determined material ID and the corresponding texture, normal and roughness material file paths are output.
[0047] Specifically, based on the vegetation knowledge graph constructed in step S10, the species or category labels output in step S20 are mapped to semantic categories in the material library. During the mapping process, all matching material sets are first retrieved from the material library by major category or subcategory (e.g., "banyan tree" or "large tree"). Then, the sets are initially filtered based on basic constraints (license compatibility, maximum or minimum model size range, etc.) to obtain candidate sets. This step ensures that vector retrieval is performed only within a semantically relevant limited space, thereby improving retrieval efficiency and semantic consistency.
[0048] For each identified ROI, the local feature vector u provided in step S20 is taken, and L2 normalization is performed on u to obtain u'. Each candidate material in the material library has its visual or geometric feature vector v pre-calculated and stored when it is added to the library, and L2 normalization is also performed to obtain v'.
[0049] Calculate the cosine similarity between each v' and u' within the candidate set: (1) The retrieval returns a Top-K candidate list sorted by similarity from highest to lowest. The results also retain similarity scores and source material data. A priority rule is applied to the Top-K candidates to determine the final candidates. This priority rule can be a weighted average of similarity, size matching, and license scores. Size matching is calculated by comparing the relative error between the model size of the source material annotations and the estimated size of the ROI; a smaller error results in a higher score. The license score is binary / graded, with compatibility = 1 and incompatibility = 0. The candidate with the highest overall score is the preferred match. If the preferred similarity score is higher than the acceptance threshold, it is automatically accepted. If the similarity score is between the rejection and acceptance thresholds, it is labeled as a semi-automatic candidate and enters the manual review queue. If the similarity score is lower than the rejection threshold, a fallback strategy is triggered to expand the semantic filtering scope.
[0050] Step S40: Based on reference correction, knowledge graph fusion and automatic discrimination or optimization strategies, reliable scale correction, quantity or location inference of vegetation ROI and batch 3D models that meet ecological constraints are imported and placed into the target simulation platform or visualization platform.
[0051] Specifically, for each vegetation ROI, the size of each vegetation ROI in real space is determined using map scale information. The estimation results are corrected based on the typical scale of the reference object to obtain the scaling factor. For each ROI, observation features such as the aspect ratio of the minimum bounding rectangle are automatically extracted based on the segmentation mask, and the planting type (single plant, multiple plants, or low shrubs) is determined using prior knowledge from the knowledge graph, thus determining the number and size range of vegetation to be placed. Based on the vegetation model and size or quantity estimation, scale and coordinate transformations are performed, constructing the affine coordinate transformation required by the simulation platform. This transformation is then batch-imported and instantiated through the simulation interface, while texture and material parameters are batch-assigned through the texture interface. After the matched 3D model is instantiated and imported into the target simulation platform or visualization platform, automated acceptance checks are performed, including collision consistency checks, texture missing checks, occlusion or visibility checks. Abnormal instances are marked in the metadata and trigger manual or remapping processes, ultimately preserving the complete mapping chain. Figure 3 As shown, this represents vegetation material; such as Figure 4 The image shown represents the effect of automatically generating vegetation elements from a 3D urban traffic model based on street view recognition.
[0052] After determining the 3D vegetation model, the scale ratio is calculated. Based on the original model size of the material and the target size range of the ROI, the scaling factor is generated.
[0053] Specifically, the global scale s (m / px) of the street view map is read and scaled using reference objects. Within the image containing the ROI, typical road features of known width are detected, such as lane width, spacing between double yellow lines, and sidewalk slabs. These features can be obtained through semantic segmentation and rule-based detection; for example, lane width can be obtained by morphologically detecting lane line regions and then taking the average width. Assuming the detected average lane width is... pixels, while the knowledge graph gives the standard lane width for the area as Then calculate the local scale: (2) Global scale and The fusion employs a confidence-based Bayesian update: (3) in, This represents the square of the variance of the global scale uncertainty. This represents the square of the variance of the local scale uncertainty, where the uncertainty is determined based on the reliability of the reference object. This represents the posterior scale after fusion.
[0054] After obtaining the scale, perform physical scale calculations: (4) in, Indicates the pixel area of the ROI. Indicates the width of the pixel bounding box. Represents the actual physical area. Indicates the actual physical width.
[0055] Furthermore, the planting type of each ROI is determined. Based on the actual area of the vegetation, aspect ratio, number of connected components, distribution sparseness and other morphological characteristics, it is determined whether the vegetation distribution of the ROI is a single tree, a row of street trees or a shrub belt.
[0056] Specifically, if the ROI has high pixel density, a simple outline, and a crown area close to that of a typical single tree, it is judged as a single tree or a clump of vegetation; if the ROI is long and distributed along the edge of the road, it is judged as multiple roadside trees; if the ROI area is smaller than the single tree threshold and the height is insufficient, it is judged as a low shrub.
[0057] For single plants or clumps, refer to the corrected scale. Map the centroid position of the ROI to actual planar coordinates; if there are multiple plants or clumps, extract the principal direction line of the ROI, i.e., the road direction of the principal axis of the minimum bounding rectangle, according to the total length. Typical plant spacing in the knowledge graph Estimated number of plants : (5) in, Indicates rounding down.
[0058] The coordinates of each plant are generated at equal intervals or with random perturbations between the starting point and the ending point.
[0059] If it is a shrub belt, then based on the typical area density estimation in the knowledge graph, points are randomly scattered within the ROI and excessive clustering is avoided by constraining the minimum spacing.
[0060] Placement schemes must meet physical and ecological constraints, including minimum spacing limits, collision detection, and visual balance and occlusion minimization. Each placement checks the overlap with already placed instances, and attempts to fine-tune meshes that do not meet the requirements.
[0061] Further, after determining the placement plan, the model is imported into the scene. Taking CARLA scene import as an example, the CARLA Python API interface is called to create Actors in batches. A custom blueprint converted from FBX or OBJ is loaded, and for each recognition region, `world.spawn_actor(blueprint, transform)` is called to instantiate the model. The `transform` parameter includes position, size, and orientation. Through the coordinate transformation module, the inferred two-dimensional plane coordinates are transformed... Mapped to CARLA local coordinate system Affine transformation is used: (6) in, Indicates the scaling factor. , represents the rotation matrix, Offset the reference origin of the simulation platform. This represents the translation amount of the terrain grid.
[0062] Materials and textures are batch-assigned through the set_attribute('texture', asset_path) interface of the blueprint, and the matched high-resolution textures are automatically applied to the 3D model (i.e., the 3D model instantiated on the simulation platform).
[0063] Using the above methods, models can be quickly imported in batches and initially placed in a 3D scene, which greatly improves the construction efficiency and controllability of digital twin scenes for urban transportation.
[0064] This invention addresses the problems of missing map details and low modeling efficiency in existing maps by batch collecting materials, identifying vegetation based on deep semantic segmentation and fine-grained classification, matching materials through mapping and vector retrieval, and generating import schemes, thus realizing an integrated pipeline from street view images to 3D scenes.
[0065] Furthermore, such as Figure 5 As shown, based on the above-mentioned method for automatically generating vegetation elements in a 3D urban traffic model based on street view recognition, this invention also provides a system for automatically generating vegetation elements in a 3D urban traffic model based on street view recognition, wherein the system includes: The knowledge graph and material library construction module 51 is used to construct a vegetation element knowledge graph for the target area, define vegetation types and typical attributes, collect 3D models and multi-view images in a targeted manner, unify metadata and store it in the material library. The vegetation recognition module 52 is used to perform pixel-level semantic segmentation on street view images, identify vegetation categories based on transfer learning and attention mechanisms, and output ROI and feature vectors. The material matching module 53 is used to filter the candidate set based on ontology mapping. It uses the cosine similarity vector retrieval method to select the most matching vegetation model between the local feature vector of the identified ROI and the pre-stored feature vector in the material library, and records the mapping. The model import module 54 is used to import batch 3D models of vegetation ROIs into the target simulation platform or visualization platform based on reference correction, knowledge graph fusion and automatic discrimination or optimization strategies, including reliable scale correction, quantity or location inference and ecological constraints.
[0066] Furthermore, such as Figure 6 As shown, based on the above-mentioned method and system for automatically generating vegetation elements of urban traffic 3D model based on street view recognition, the present invention also provides a terminal, which includes a processor 10, a memory 20 and a display 30. Figure 6 Only some of the terminal components are shown; however, it should be understood that it is not required to implement all of the components shown, and more or fewer components may be implemented instead.
[0067] In some embodiments, the memory 20 may be an internal storage unit of the terminal, such as a hard disk or memory. In other embodiments, the memory 20 may be an external storage device of the terminal, such as a plug-in hard disk, smart media card (SMC), secure digital card (SD), flash card, etc. Further, the memory 20 may include both internal and external storage devices. The memory 20 is used to store application software and various types of data installed on the terminal, such as the program code installed on the terminal. The memory 20 may also be used to temporarily store data that has been output or will be output. In one embodiment, the memory 20 stores an automatic generation program 40 for urban traffic 3D model vegetation elements based on street view recognition. This automatic generation program 40 can be executed by the processor 10, thereby implementing the automatic generation method for urban traffic 3D model vegetation elements based on street view recognition in this application.
[0068] In some embodiments, the processor 10 may be a central processing unit (CPU), a microprocessor, or other data processing chip, used to run program code stored in the memory 20 or process data, such as executing the automatic generation method of vegetation elements in a three-dimensional urban traffic model based on street view recognition.
[0069] In some embodiments, the display 30 may be an LED display, a liquid crystal display, a touch-sensitive liquid crystal display, or an OLED (Organic Light-Emitting Diode) touchscreen. The display 30 is used to display information on the terminal and to display a visual user interface. The terminal's processor 10, memory 20, and display 30 communicate with each other via a system bus.
[0070] In one embodiment, when the processor 10 executes the automatic generation program 40 for urban traffic 3D model vegetation elements based on street view recognition in the memory 20, it implements the steps of the automatic generation method for urban traffic 3D model vegetation elements based on street view recognition as described above.
[0071] The present invention also provides a computer-readable storage medium, wherein the computer-readable storage medium stores an automatic generation program for vegetation elements of a three-dimensional urban traffic model based on street view recognition, and the automatic generation program for vegetation elements of a three-dimensional urban traffic model based on street view recognition, when executed by a processor, implements the steps of the automatic generation method for vegetation elements of a three-dimensional urban traffic model based on street view recognition as described above.
[0072] In summary, this invention provides a method, system, terminal, and computer-readable storage medium for automatically generating vegetation elements in urban traffic 3D models based on street view recognition. The method includes: constructing a vegetation element knowledge graph for the target area, defining vegetation types and typical attributes, directionally acquiring 3D models and multi-view images, unifying metadata and storing it in a material library; performing pixel-level semantic segmentation on street view images, identifying vegetation categories based on transfer learning and attention mechanisms, and outputting ROIs and feature vectors; filtering candidate sets based on ontology mapping, using a cosine similarity vector retrieval method to select the most matching vegetation model between the local feature vectors of the identified ROI and pre-stored feature vectors in the material library, and recording the mapping; and based on reference object correction, knowledge graph fusion, and automatic discrimination or optimization strategies, reliably correcting the scale of vegetation ROIs, inferring their quantity or location, and importing batches of 3D models that meet ecological constraints into a target simulation platform or visualization platform. This invention, through an integrated process of standardized material acquisition, semantic recognition, vectorized matching, and procedural generation, realizes an integrated pipeline from street view images to 3D scenes, providing efficient, intelligent, and standardized technical support for traffic simulation and urban visualization.
[0073] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or terminal that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or terminal. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or terminal that includes that element.
[0074] Of course, those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware (such as a processor, controller, etc.). The program can be stored in a computer-readable storage medium, and when executed, it can include the processes described in the above method embodiments. The computer-readable storage medium can be a memory, magnetic disk, optical disk, etc.
[0075] It should be understood that the application of the present invention is not limited to the examples above. Those skilled in the art can make improvements or modifications based on the above description, and all such improvements and modifications should fall within the protection scope of the appended claims.
Claims
1. A method for automatically generating vegetation elements in a 3D urban traffic model based on street view recognition, characterized in that, The method for automatically generating vegetation elements in a 3D urban traffic model based on street view recognition includes: Construct a vegetation element knowledge graph for the target area, define vegetation types and typical attributes, collect 3D models and multi-view images in a targeted manner, unify metadata and store it in the material library. Pixel-level semantic segmentation is performed on street view images, vegetation categories are identified based on transfer learning and attention mechanisms, and ROI and feature vectors are output. Based on ontology mapping, candidate sets are filtered, and the cosine similarity vector retrieval method is used to select the vegetation model that best matches the local feature vector of the identified ROI with the pre-stored feature vector in the material library, and the mapping is recorded. Based on reference correction, knowledge graph fusion, and automatic discrimination or optimization strategies, reliable scale correction, quantity or location inference of vegetation ROI, and batch import of ecologically constrained 3D models are placed into the target simulation platform or visualization platform.
2. The method for automatically generating vegetation elements in a three-dimensional urban traffic model based on street view recognition according to claim 1, characterized in that, The construction of a vegetation element knowledge graph for the target area, defining vegetation types and typical attributes, targeted acquisition of 3D models and multi-view images, and unified metadata storage in a material library specifically includes: The target area is identified, a vegetation element knowledge graph for the target area is constructed, the hierarchical structure and attributes of common vegetation categories in the target area are determined, and the vegetation element knowledge graph is used as a semantic reference ontology for collection and retrieval. Based on the categories and attributes in the vegetation element knowledge graph, targeted crawlers and maps or data interfaces are used to collect and filter 3D materials and multi-view images related to vegetation in batches, and record the source, permission information and multi-view information. The acquired 3D materials are formatted and their metadata is standardized, a unified naming convention is established, metadata is generated, and the material library is organized according to semantic directories for retrieval and version management. Perform format verification and automatic repair on the original 3D models and textures, batch convert formats, check and automatically repair existing problems, and annotate samples that do not meet the quality standards.
3. The method for automatically generating vegetation elements in a three-dimensional urban traffic model based on street view recognition according to claim 1, characterized in that, The process of performing pixel-level semantic segmentation on street view images, identifying vegetation categories based on transfer learning and attention mechanisms, and outputting ROIs and feature vectors specifically includes: Pixel-level semantic segmentation is used on street view images to locate vegetation regions (ROIs), and the region mask and local features are output to provide input samples for fine-grained recognition. Based on segmentation, a fine-grained classifier is trained to identify ROI vegetation categories; The fine-grained classifier outputs a semantic label, a local confidence score, and a corresponding local visual feature vector for each identified object.
4. The method for automatically generating vegetation elements in a three-dimensional urban traffic model based on street view recognition according to claim 3, characterized in that, The fine-grained classifier is based on a transfer learning strategy, using a deep backbone network pre-trained on a large-scale dataset as a foundation, and fine-tuned to a local vegetation dataset with an attention mechanism to enhance sensitivity to texture and morphological features.
5. The method for automatically generating vegetation elements in a three-dimensional urban traffic model based on street view recognition according to claim 1, characterized in that, The process of filtering the candidate set based on ontology mapping involves using a cosine similarity vector retrieval method to select the vegetation model that best matches the local feature vectors of the identified ROI with the pre-stored feature vectors in the material library, and recording the mapping. Specifically, this includes: Based on the predefined category and subclass correspondence in the ontology, candidate resources of the same type are selected from the material library, and preliminary filtering is performed based on size range and version constraints; The similarity between the deep feature vector extracted from the recognition region and the feature vector of the candidate material is calculated, and the similarity is sorted from high to low. The top K candidates in the sorting results are taken as the final mapping candidate list. Perform rule-based priority determination on K candidates, and output the final determined material ID and the corresponding material file paths for texture, normal, and roughness.
6. The method for automatically generating vegetation elements in a three-dimensional urban traffic model based on street view recognition according to claim 1, characterized in that, The aforementioned strategy, based on reference correction, knowledge graph fusion, and automatic discrimination or optimization, involves the reliable scale correction, quantity or location inference of vegetation ROIs, and the import of batch 3D models that meet ecological constraints into the target simulation platform or visualization platform. Specifically, this includes: For each vegetation ROI, the size of each vegetation ROI in real space is determined by the map scale information, and the estimation results are corrected according to the typical scale of the reference object to obtain the scaling factor; For each ROI, observation features are automatically extracted based on segmentation masks, and the planting type of the ROI is determined by combining prior knowledge graphs to determine the number and size range of vegetation to be placed. Based on the vegetation model and size or quantity estimation, scale and coordinate transformations are performed to construct the affine coordinate transformation required for the simulation platform. The transformations are then imported and instantiated in batches through the simulation interface. Simultaneously, texture and material parameters are assigned in batches through the texture interface. After the matched 3D model is instantiated and imported into the target simulation platform or visualization platform, an automated acceptance check is performed. Abnormal instances are marked in the metadata and a manual or remapping process is triggered, ultimately preserving the complete mapping chain.
7. The method for automatically generating vegetation elements in a three-dimensional urban traffic model based on street view recognition according to claim 6, characterized in that, The automated acceptance checks include checks for collider consistency, missing textures, and occlusion or visibility.
8. An automatic generation system for vegetation elements in a 3D urban traffic model based on street view recognition, characterized in that, The automatic generation system for vegetation elements in a 3D urban traffic model based on street view recognition includes: The knowledge graph and material library construction module is used to build a vegetation element knowledge graph for the target area, define vegetation types and typical attributes, collect 3D models and multi-view images in a targeted manner, unify metadata and store it in the material library. The vegetation recognition module is used to perform pixel-level semantic segmentation on street view images, identify vegetation categories based on transfer learning and attention mechanisms, and output ROI and feature vectors. The material matching module is used to filter the candidate set based on ontology mapping. It uses the cosine similarity vector retrieval method to select the most matching vegetation model between the local feature vector of the identified ROI and the pre-stored feature vector in the material library, and records the mapping. The model import module is used to import batch 3D models of vegetation ROIs into the target simulation platform or visualization platform based on reference correction, knowledge graph fusion and automatic discrimination or optimization strategies, including reliable scale correction, quantity or location inference and ecological constraints.
9. A terminal, characterized in that, The terminal includes: a memory, a processor, and an automatic generation program for urban traffic 3D model vegetation elements based on street view recognition, which is stored in the memory and can run on the processor. When the automatic generation program for urban traffic 3D model vegetation elements based on street view recognition is executed by the processor, it implements the steps of the automatic generation method for urban traffic 3D model vegetation elements based on street view recognition as described in any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores an automatic generation program for vegetation elements of a three-dimensional urban traffic model based on street view recognition. When the automatic generation program for vegetation elements of a three-dimensional urban traffic model based on street view recognition is executed by a processor, it implements the steps of the automatic generation method for vegetation elements of a three-dimensional urban traffic model based on street view recognition as described in any one of claims 1-7.
Citation Information
Patent Citations
Three-dimensional parametric modeling method for beam bridge
CN110263376A
Vegetation model auxiliary generation method and system based on aerial image and CIM
CN111798567A
Scene graph generation method and system supporting historical and cultural block scene
CN118334414A
Photovoltaic power station operation and maintenance method and system combining three-dimensional surveying and mapping and digital twinning
CN118736444A
Cited By
Multi-modal data driven three-dimensional simulation scene construction method and device
CN121527325A
Multi-category industrial detection sorting system and method based on AI large model and robot and storage medium
CN121639621A
Urban rail three-dimensional model generation method, system and device and computer medium
CN122289567A