Road modeling method and device, equipment, storage medium and program product
By acquiring image sequences from multi-view cameras to generate virtual point clouds and performing orthophoto projection, combined with a large visual model for independent modeling, the problem of local update difficulty in LiDAR modeling methods is solved, achieving low-cost and efficient dynamic road update.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- QIANXUN SPATIAL INTELLIGENCE INC
- Filing Date
- 2025-12-31
- Publication Date
- 2026-04-28
AI Technical Summary
Existing lidar modeling methods cannot achieve accurate updates when there are local changes in the road, resulting in wasted resources and low update efficiency.
Image sequences are acquired using multi-view cameras to generate virtual point clouds, and orthophotos are generated to produce digital orthophoto maps and digital elevation models. The structured road surface and the three-dimensional model above it are extracted using a large visual model, and an independent modeling process is used to achieve local updates.
It enables rapid and low-cost updates when there are local changes in the road, avoiding large-scale reconstruction, reducing resource waste and improving update efficiency.
Smart Images

Figure CN121937653A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of 3D modeling technology, and in particular relates to a road modeling method, apparatus, equipment, storage medium and program product. Background Technology
[0002] Road modeling refers to the creation of three-dimensional geometric and semantic models of roads and their ancillary facilities through digital technology, which has important application value in fields such as intelligent transportation, autonomous driving simulation, and digital twins.
[0003] In existing technologies, road 3D modeling is mainly performed using point cloud modeling methods based on lidar. Specifically, a mobile measurement platform equipped with lidar collects high-precision 3D point cloud data of the road environment, then processes the 3D point cloud data, and finally converts the processed point cloud data into a visualized 3D road model through steps such as point cloud meshing and texture mapping.
[0004] However, when local changes occur on the road, such as the addition of traffic signs, existing LiDAR modeling methods cannot achieve accurate updates to the local model. They can only re-collect point cloud data for the changed area and the surrounding large-scale roads and reconstruct the entire process. This is difficult to adapt to the actual needs of dynamic road updates, resulting in a waste of resources. Summary of the Invention
[0005] This application provides a road modeling method, apparatus, device, storage medium, and program product, which can accurately update the local model when the road changes, avoiding the problem of wasting resources by rebuilding the entire process.
[0006] In a first aspect, embodiments of this application provide a road modeling method, the method comprising: Image sequences of the road to be modeled are acquired by multi-view cameras set on a mobile platform, and virtual point clouds under a unified geographic coordinate system are generated based on the image sequences. Each view camera corresponds to an image sequence, and each frame in the image sequence corresponds to a pose information. Orthophoto projection is performed on the virtual point cloud to generate a digital orthophoto map and a digital elevation model corresponding to the road to be modeled. The digital orthophoto map and digital elevation model are input into the road surface modeling model, and the road surface modeling model is used to generate a structured three-dimensional model of the road surface. The image sequence is input into the road overhead model, and the road overhead model is used to generate a structured 3D road overhead model; By spatially associating and integrating the 3D model of the structured road surface with the 3D model above the structured road according to a unified geographic coordinate system, a structured 3D road model is generated.
[0007] In some possible implementations, digital orthophoto maps and digital elevation models are input into a road surface modeling model, and a structured 3D model of the road surface is generated using the road surface modeling model, including: The digital orthophoto is input into the visual big model, and the vector contours of road surface features are extracted using the visual big model to obtain two-dimensional feature contour data with semantic information and geographic coordinates. Two-dimensional feature outline data is back-projected onto a digital elevation model based on a unified geographic coordinate system to obtain three-dimensional feature outline data with elevation coordinates. Using 3D feature contour data as geometric constraint boundaries, constrained surface fitting is performed on the digital elevation model to generate a 3D model of the structured road surface.
[0008] In some possible implementations, an image sequence is input into a road overhead model, and the road overhead model is used to generate a structured 3D road overhead model, including: The image sequence is input into a large visual model, and the large visual model is used to extract the elements above the road to obtain the elements above the road with semantic information. The feature information of the elements above the road is extracted, and image retrieval technology is used to retrieve a structured 3D model above the road with a similarity greater than a preset threshold to the feature information in a standard model library.
[0009] In some possible implementations, after retrieving a 3D model above a structured road with a similarity greater than a preset threshold to feature information from a standard model library using image retrieval technology, the following steps are also included: Calculate the parallax information of road-above features in image sequences acquired by multi-view cameras; Based on parallax information, parameters of multi-view cameras, and pose information, the three-dimensional spatial coordinates of the features above the road in a unified geographic coordinate system are calculated; where the pose information is the pose information corresponding to the image frame where the features above the road are located. Based on three-dimensional spatial coordinates, the three-dimensional model above the structured road is instantiated into the three-dimensional model of the structured road surface.
[0010] Among some possible implementations, generating virtual point clouds in a unified geographic coordinate system based on image sequences includes: Dense stereo matching is performed on each set of images acquired synchronously by multi-view cameras to generate the corresponding disparity map; Based on the focal length and baseline corresponding to the multi-view camera, the disparity map is converted into a depth map; Based on the pose information corresponding to each group of images, the depth map is back-projected onto a unified geographic coordinate system to generate a local virtual point cloud; Under a unified geographic coordinate system, the local virtual point clouds generated from each set of images are fused and denoised to generate a virtual point cloud of the road to be modeled under the unified geographic coordinate system.
[0011] In some possible implementations, after acquiring an image sequence of the road to be modeled using a multi-view camera set up on a mobile platform, the following is also included: The image sequence is preprocessed, including at least one of the following: Based on the intrinsic distortion coefficients of multi-view cameras, distortion correction processing is performed on each frame of the image sequence. An epipolar correction algorithm is used to project each group of images onto the same epipolar plane; An adaptive histogram equalization algorithm is used to perform illumination equalization processing on each frame of the image sequence.
[0012] Secondly, embodiments of this application provide a road modeling device, the device comprising: The acquisition module is used to acquire image sequences of the road to be modeled through multi-view cameras set on the mobile platform, and generate virtual point clouds in a unified geographic coordinate system based on the image sequences. Each view camera corresponds to an image sequence, and each frame in the image sequence corresponds to a pose information. The projection module is used to perform orthophoto projection on the virtual point cloud to generate a digital orthophoto map and a digital elevation model corresponding to the road to be modeled. The first processing module is used to input digital orthophoto maps and digital elevation models into the road surface modeling model, and use the road surface modeling model to generate a structured three-dimensional road surface model. The second processing module is used to input the image sequence into the road overhead model and use the road overhead model to generate a structured road overhead 3D model. The integration module is used to spatially associate and integrate the 3D model of the structured road surface with the 3D model above the structured road according to a unified geographic coordinate system, thereby generating a structured 3D road model.
[0013] Thirdly, embodiments of this application provide an electronic device, the device comprising: A processor and a memory storing computer program instructions; a road modeling method that implements any of the above when the processor executes the computer program instructions.
[0014] Fourthly, embodiments of this application provide a computer-readable storage medium storing computer program instructions, which, when executed by a processor, implement a road modeling method as described above.
[0015] Fifthly, embodiments of this application provide a computer program product in which instructions, when executed by a processor of an electronic device, enable the electronic device to perform any of the aforementioned road modeling methods.
[0016] The road modeling method, apparatus, device, storage medium, and program product of this application acquire image sequences of the road to be modeled using a multi-view camera mounted on a mobile platform. A virtual point cloud is generated based on the image sequences. Compared to LiDAR, this acquisition method is not only lower in cost and more flexible in deployment, but also allows for rapid response to changes in the road. It only requires small-scale, targeted image re-capture of the changed area. Secondly, by orthophotoing the virtual point cloud, a digital orthophoto map and a digital elevation model (DEM) corresponding to the road to be modeled are generated. On one hand, the DEM and DEM are converted into a structured 3D road surface model using a road surface modeling model. On the other hand, a structured 3D road surface model is directly generated based on the image sequences using a road surface modeling model. These two modeling processes are independent, eliminating the need to re-acquire data and re-model a large area when local changes occur in the road. Only targeted re-capture of the changed area using a multi-view camera is needed to obtain an updated image sequence. The model corresponding to the changed area is then updated separately. Therefore, this method meets the actual needs of dynamic road updates and avoids the resource waste associated with existing reconstruction technologies. Attached Figure Description
[0017] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments of this application will be briefly introduced below. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0018] Figure 1 A schematic flowchart of a road modeling method provided in one embodiment of this application is shown; Figure 2 A flowchart illustrating a road modeling method provided in another embodiment of this application is shown; Figure 3 The process of semantic vector extraction from a large visual model based on the DOM is illustrated. Figure 4 A flowchart illustrating the generation process of a 3D model of a structured road surface is shown. Figure 5 A flowchart illustrating a road modeling method provided in another embodiment of this application is shown; Figure 6 The flowchart for generating the 3D model above the structured road is shown; Figure 7 This illustrates the integration and output process of the hierarchical structured model; Figure 8 A flowchart illustrating a road modeling method provided in another embodiment of this application is shown; Figure 9 A road modeling system is shown; Figure 10 A schematic diagram of the road modeling device provided in an embodiment of this application is shown; Figure 11 A schematic diagram of the hardware structure of the electronic device provided in an embodiment of this application is shown. Detailed Implementation
[0019] The features and exemplary embodiments of various aspects of this application will be described in detail below. To make the objectives, technical solutions, and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are only intended to explain this application and not to limit it. For those skilled in the art, this application can be implemented without some of these specific details. The following description of the embodiments is merely to provide a better understanding of this application by illustrating examples.
[0020] It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising..." does not exclude the presence of additional identical elements in the process, method, article, or apparatus that includes the element.
[0021] It should be noted that the acquisition, storage, use, and processing of data in this application embodiment all comply with the relevant provisions of national laws and regulations.
[0022] It should be noted that in the embodiments of this application, certain software, components, models and other existing solutions in the industry may be mentioned. These should be regarded as exemplary and are only intended to illustrate the feasibility of implementing the technical solution of this application. However, it does not mean that the applicant has used or necessarily used the solution.
[0023] First, let me explain the terms used in this application: DEM: Digital Elevation Model, is a digital simulation of ground terrain (i.e., a digital representation of the terrain surface morphology) achieved through limited terrain elevation data. It is a physical ground model that represents ground elevation using an ordered array of numerical values, and is a branch of Digital Terrain Model (DTM). One data format is GeoTIFF. DEMs focus on the elevation information of exposed terrain, acquiring the terrain model through 3D reconstruction.
[0024] DOM (Digital Orthophoto Map) is image data based on aerial photographs or remote sensing imagery (monochrome / color). It is scanned, processed pixel-by-pixel with radiometric correction, differential correction, and mosaicking, and then cropped to the extent of a topographic map. Topographic feature information is added to the image plane using symbols, line drawings, annotations, kilometer grids, and map borders (inner / outer), forming an image database stored in raster data format. It possesses the geometric accuracy and image characteristics of a topographic map. DOM is a two-dimensional orthophoto map generated by eliminating image distortion and combining it with topographic elevation data, emphasizing the true texture of the earth's surface.
[0025] GeoTIFF (Geographic Tagged Image File Format) is a commonly used geographic information image file format that combines image data and geographic information data. It can be used to store and transmit image data with geographic location references. The GeoTIFF format supports embedding geographic information such as geographic coordinates, projection information, and ellipsoid parameters.
[0026] Vision-Language Models (VLMs): Semantic recognition achieves a deep understanding of image content by fusing visual and linguistic information. VLMs process both image and text data simultaneously, establishing connections between visual features and linguistic semantics to achieve comprehensive analysis of objects, scenes, actions, and contextual relationships. This includes image description generation, visual question answering (VQA), object detection, and semantic segmentation.
[0027] Binocular vision principle: Depth calculation based on parallax is a classic application in computer vision. By capturing images of the same scene using two cameras (left and right), the distance between the object and the camera is calculated using the difference in the position of the object in the two images (i.e., parallax).
[0028] Dense stereo matching algorithm (SGM) is a widely used dense stereo matching algorithm that can maintain global consistency while ensuring local accuracy. It is especially suitable for applications with high real-time requirements, such as vehicle navigation and mobile mapping.
[0029] Currently, the main method for road modeling is the point cloud modeling method using LiDAR. While this method can achieve high-precision modeling, its modeling process is fixed and global. When local road features (such as road markings and signs) change, existing technologies cannot accurately update the local model for the changed areas. Instead, the entire road or relevant road segments must be rescanned to obtain new 3D point clouds, and then the entire modeling process must be run again, starting from the point cloud processing. This process not only significantly extends the modeling cycle and reduces update efficiency, but also results in a significant waste of computational and storage resources because it requires repeated modeling of unchanged parts.
[0030] To address the problems of existing technologies, this application provides a road modeling method, apparatus, device, storage medium, and program product. By acquiring image sequences of the road to be modeled using a multi-view camera mounted on a mobile platform, a virtual point cloud is generated based on the image sequences. Compared to LiDAR, this acquisition method is not only lower in cost and more flexible in deployment, but also allows for rapid response to changes in the road surface. It only requires small-scale, targeted image re-capture of the changed area. Next, by orthophotoing the virtual point cloud, a digital orthophoto map and a digital elevation model (DEM) corresponding to the road to be modeled are generated. On one hand, the DEM and DEM are converted into a structured 3D road surface model using a road surface modeling model. On the other hand, a structured 3D road surface model is directly generated based on the image sequences using a road surface modeling model. These two modeling processes are independent, eliminating the need to re-capture data and remodel a large area when local changes occur in the road. Only targeted re-capture of the changed area using a multi-view camera is needed to obtain an updated image sequence, and then the model corresponding to the changed area can be updated separately. Therefore, this method can meet the actual needs of dynamic road updates and avoid the resource waste caused by existing reconstruction technologies. The following section first introduces a road modeling method provided in the embodiments of this application.
[0031] Figure 1 A schematic flowchart of a road modeling method provided in one embodiment of this application is shown. Figure 1 As shown, the method may include the following steps: S101: Acquire image sequences of the road to be modeled by multi-view cameras set on a mobile platform, and generate virtual point clouds in a unified geographic coordinate system based on the image sequences. Each view camera corresponds to an image sequence, and each frame in the image sequence corresponds to a pose information.
[0032] In this embodiment, a multi-view camera is installed on a mobile platform (such as an inspection vehicle) to synchronously acquire image sequences of the road to be modeled. Each frame of the image is associated with pose information to generate a virtual point cloud in a unified geographic coordinate system. Compared to traditional LiDAR point cloud modeling schemes, acquiring images through a multi-view camera and converting them into virtual point clouds not only meets the point cloud requirements for road modeling but also significantly reduces hardware investment costs, solving the problem of high LiDAR equipment costs. The pose information can be provided by a Global Navigation Satellite System (GNSS) and Inertial Measurement Unit (IMU) device installed on the mobile platform. Through hardware synchronization triggering or spatiotemporal synchronization mechanisms, the pose information at the corresponding moment is recorded simultaneously while the multi-view camera captures each frame of the image.
[0033] In one example, the multi-view camera can be a binocular camera, which is an imaging device consisting of two cameras with parallel optical axes and a fixed distance between them. By simultaneously capturing two images of the same scene, depth information is calculated using the positional difference (i.e., parallax) of objects in the two images, providing a data foundation for virtual point cloud generation. To further improve the coverage of image acquisition and the accuracy of depth calculation, a tri-camera or other camera system with different numbers of lenses can also be used to reduce blind spots by increasing the number of viewing dimensions. All multi-view cameras are time-synchronized via hardware triggering to ensure that the image sequences acquired from each viewpoint are consistent in the temporal dimension.
[0034] In another example, to fully cover the area of the road to be modeled, a set of multi-view cameras can be installed on the front and right sides of the mobile platform. The optical axis of the front camera group is parallel to the direction of travel of the mobile platform and is mainly used to capture the road area in front; the optical axis of the right camera group is perpendicular to the direction of travel and points to the right, and is mainly used to capture the roadside, facilities, and other areas on the right side of the road. It should be noted that, in order to improve the overall acquisition efficiency, camera groups can also be added to the left or other positions of the mobile platform to achieve all-round acquisition of the road environment.
[0035] It should be noted that if a binocular camera is used and it is installed in the front and right directions on the mobile platform, the final output images will be: a front-left image sequence, a front-right image sequence, a right-left image sequence, and a right-right image sequence, where the resolution of each frame can be 1920×1080 pixels.
[0036] Secondly, this application relies on multi-view cameras to collect image sequences of the road to be modeled. These images naturally contain rich visual features such as texture, color, and light and shadow, which helps the large visual model to fully learn the semantic features of road elements such as lane lines, stop lines, and traffic signs. In contrast, LiDAR point cloud data only focuses on geometric location information and has a very weak capacity to carry key semantic clues such as color and texture. For road elements such as yellow lane lines and red stop lines that rely on color differentiation, the accuracy of semantic recognition is obviously insufficient.
[0037] Furthermore, current mainstream general-purpose large-scale models and semantic segmentation models are all trained on massive image data, which has a stronger learning depth and generalization ability for image visual features. When transferring them to the task of feature vector extraction in road scenes, only a small amount of fine-tuning is needed for road data to achieve high recognition and extraction accuracy. However, point cloud models are limited by the scarcity of training data and the high cost of data annotation, and the generalization ability of the models is generally weak. In complex road scenes such as multi-lane intersections and dense traffic facilities, the adaptability and stability of semantic recognition and feature extraction are difficult to guarantee.
[0038] Furthermore, the two-dimensional data structure of images significantly reduces the computational complexity of model detection and vector extraction: a single 1920×1080 pixel road image can achieve millisecond-level processing response on ordinary GPU hardware; while the three-dimensional unstructured nature of point clouds requires more computational resources for feature extraction, clustering, and other processing steps, with processing time for the same road scene being approximately 5-10 times that of images. Taking the vector extraction of features from a 1-kilometer road as an example: image-based processing can be completed in just a few minutes, while point cloud data processing may take tens of minutes or even hours, resulting in a significant efficiency difference.
[0039] In one example, to improve the accuracy and efficiency of road modeling, the image sequence can be preprocessed after acquisition. The preprocessing operations include: Based on the intrinsic distortion coefficients of multi-view cameras, distortion correction processing is performed on each frame of the image sequence.
[0040] In this embodiment, optical distortion (such as radial distortion and tangential distortion) in the camera lens can cause shape distortion of objects in the image, such as bending the edges of straight lines in the image, affecting the accuracy of subsequent dense stereo matching and semantic recognition. By using pre-calibrated camera intrinsic parameters and distortion coefficients, distortion correction is performed on each frame of the image sequence to eliminate image distortion caused by lens distortion.
[0041] In one example, camera intrinsic parameters (focal length, principal point coordinates) and distortion coefficients (radial distortion coefficient, tangential distortion coefficient) can be obtained based on Zhang's calibration method, and the coordinates of each pixel in each frame of the image can be corrected using the distortion correction formula.
[0042] An epipolar correction algorithm is used to project each group of images onto the same epipolar plane.
[0043] In the embodiments of this application, each set of images refers to a pair of images with parallax relationships, acquired by a set of multi-view cameras (such as the left and right eyes in a binocular camera) triggered at the same time. In unprocessed multi-view image pairs, due to the differences in relative position and pose between the cameras, the corresponding pixels of the same physical point in the left and right images may be located in different rows (i.e., not on the same horizontal line). This causes the algorithm to search for matching points across the entire two-dimensional image range when performing dense stereo matching, resulting in high computational complexity and low matching efficiency.
[0044] Therefore, by using the epipolar correction algorithm, each group of multi-view images is projected onto the same epipolar plane, so that the corresponding pixels in the left and right images are constrained to the same horizontal line, simplifying the matching search range from a two-dimensional image plane to a one-dimensional horizontal line, thereby reducing the amount of computation and improving the matching speed.
[0045] Specifically, a pair of projection transformation matrices can be calculated based on the pre-calibrated relative poses (rotation matrix and translation vector) between the multi-view cameras. Applying these matrices to each original set of images will output a row-aligned corrected set of images.
[0046] An adaptive histogram equalization algorithm is used to perform illumination equalization processing on each frame of the image sequence.
[0047] In this embodiment of the application, outdoor data acquisition is greatly affected by changes in lighting (shadows, backlighting, inside and outside tunnels), which can cause local areas of the image to be too bright or too dark, affecting the robustness of subsequent dense stereo matching and semantic recognition. By using an adaptive histogram equalization algorithm, the brightness and contrast of each frame of the image are adjusted to make the image lighting distribution uniform and the details in the low-light areas clearly visible.
[0048] In one example, this application does not limit the execution order or number of the above preprocessing operations. The execution order can be flexibly adjusted based on the needs of the actual modeling scenario, such as image quality, hardware computing power, and modeling accuracy requirements, or some preprocessing steps can be omitted to adapt to different application scenarios.
[0049] In this embodiment, after acquiring the image sequence of the road to be modeled, the image sequence is preprocessed to ensure both the accuracy and efficiency of road modeling. On the one hand, distortion correction is performed on each frame of the image sequence based on multi-view camera intrinsic parameters and distortion coefficients to eliminate image distortion caused by lens optical distortion, thereby providing a precise image foundation for subsequent dense stereo matching and semantic extraction, ensuring the accuracy of road modeling. On the other hand, an epipolar correction algorithm is used to project each set of synchronously acquired multi-view images onto the same epipolar plane, so that the corresponding pixels of the same object in the scene accurately fall on the same horizontal line, simplifying the two-dimensional global search in the dense stereo matching process into a one-dimensional intra-row search, significantly reducing the matching computation and improving the matching speed, thereby ensuring the efficiency of road modeling. In addition, an adaptive histogram equalization algorithm is used to perform illumination homogenization processing on each frame of the image sequence, enhancing the detailed features of road elements in low-light areas, avoiding interference from illumination differences on subsequent semantic recognition and dense stereo matching, and further ensuring the accuracy of road modeling.
[0050] S102: Perform orthophoto projection on the virtual point cloud to generate a digital orthophoto map and a digital elevation model corresponding to the road to be modeled.
[0051] In this embodiment, the original image suffers from perspective distortion (such as compression of distant road markings and deformation of road edges). Directly using it as input for subsequent road surface modeling would lead to geometric distortion. Furthermore, the virtual point cloud only contains three-dimensional coordinates and lacks texture information, making it unsuitable for direct input into the road surface modeling for semantic extraction and structured modeling. By orthophotoing the virtual point cloud, a digital orthophoto image and a digital elevation model are generated, providing an image foundation for subsequent structured three-dimensional road surface models.
[0052] It should be noted that DOM is a geographically aligned, perspective-free planar image with a resolution of 3cm×3cm; DEM is a raster model with elevation information, also with a resolution of 3cm×3cm. Furthermore, DOM and DEM share the same geographic coordinate system (such as UTM / WGS84), providing a basis for subsequent fusion and ensuring spatial consistency between the two.
[0053] In one example, virtual point clouds can be generated frame by frame, meaning each image frame can generate a local virtual point cloud. For each virtual point cloud frame and its corresponding pose information, multiple DOM and DEM frames can be generated through orthophoto projection. The specific implementation method can be flexibly chosen: for example, each virtual point cloud frame can be transformed to the same geographic coordinate system using the pose information before performing orthophoto projection; alternatively, the virtual point clouds can be orthophoto projected individually first, and then a unified geographic coordinate system transformation can be completed based on the pose information. After obtaining multiple DOM and DEM frames, they can be fused to form a fused DOM covering the entire road to be modeled. This operation avoids repeated recognition of the same road area, improving the efficiency of subsequent semantic extraction and reducing errors caused by repeated recognition, thus increasing the accuracy of the recognition results. Compared to directly modeling the road surface based on the original images, this example avoids complex stitching and registration of multiple original images and avoids the tedious and inefficient manual correction and verification of intermediate modeling results. Furthermore, DOM and DEM are compatible with Geographic Information System (GIS), allowing direct import of DOM and DEM into platforms such as ArcGIS and QGIS, facilitating manual correction and management, as well as integration and analysis with other geographic information data.
[0054] S103: Input the digital orthophoto map and digital elevation model into the road surface modeling model, and use the road surface modeling model to generate a structured three-dimensional road surface model.
[0055] In this application embodiment, traditional 3D reconstruction and semantic recognition are separated, resulting in semantic tags not being accurately mapped to 3D space. This leads to a disconnect between semantics and geometry, resulting in models that are mostly unstructured point clouds and cannot meet the needs of engineering editing and querying. This application inputs DOM and DEM into the road surface modeling model and generates a structured 3D road surface model with semantic information through semantic extraction and elevation fusion.
[0056] It should be noted that digital orthophoto maps are images generated through orthophoto projection, and therefore primarily present two-dimensional information about the road surface. Based on this, the "road surface modeling" in this application refers to the process of three-dimensional reconstruction and structured representation of the road surface and its elements reflected in the image. Specifically, the modeling objects include not only the physical bearing surface of the road (such as asphalt or concrete pavement layers), but also all key functional elements attached to its surface, such as: road markings (lane lines, stop lines, directional arrows), ground-printed markings (such as the word "stop"), curb stones, pedestrian crossings, and median strip boundaries, etc.
[0057] S104: Input the image sequence into the road overhead model and use the road overhead model to generate a structured 3D road overhead model.
[0058] In this embodiment, road top modeling refers to the process of creating 3D models of various facilities located above the road surface. These facilities include, but are not limited to, traffic signs, traffic lights, streetlights, monitoring poles, guardrails, and traffic signals. Since these facilities typically have complex 3D structures, reconstructing each facility from scratch would consume significant computational resources and be inefficient. Considering that most facilities are standardized and regulated components, the road top modeling model used in this application does not reconstruct these facilities from scratch. Instead, it identifies the type and spatial location information of the facilities based on multi-view image sequences and directly calls the corresponding standardized prefabricated models, eliminating the tedious steps of modeling from scratch and reducing the cost and time required for modeling.
[0059] Secondly, the road overground modeling process directly relies on the original image sequence and its pose information, without depending on any intermediate or final results of road surface modeling. There are no data dependencies or temporal constraints between the road overground modeling process and the road surface modeling process; they can be executed completely independently and in parallel. This effectively utilizes computational resources, shortens the overall modeling cycle, and significantly improves the overall modeling speed.
[0060] S105: Spatially associate and integrate the 3D model of the structured road surface with the 3D model above the structured road according to a unified geographic coordinate system to generate a structured 3D road model.
[0061] In this embodiment of the application, since the three-dimensional model of the structured road surface and the three-dimensional model above the structured road are both constructed based on the same unified geographic coordinate system, the two types of models can be accurately spatially associated directly based on the unified geographic coordinate system.
[0062] It should be noted that this structured 3D road model adopts a hierarchical structured storage architecture. The data of each layer is independent of each other and linked through unified geographic coordinates. For example, it includes: a pavement layer, a marking layer, and a superstructure layer. Among them, the pavement layer corresponds to the main part of the structured road surface 3D model, carrying the basic geometric shape and elevation information of the road; the marking layer contains 3D vector data of road surface semantic elements such as lane lines, stop lines, directional arrows, and ground text, which are precisely matched with the pavement layer; the superstructure layer corresponds to the structured road's upper 3D model, covering standardized 3D models of overhead facilities such as traffic lights, signs, light poles, and fences.
[0063] The hierarchical storage design enables independent updates by module. For example, lane line colors or types can be updated by modifying only the vector data of the marking layer, or a certain type of standardized model in the upper facility layer can be directly replaced (such as replacing a new type of traffic sign). This eliminates the need to rebuild the entire 3D road model, significantly reducing the update costs and time costs of model maintenance.
[0064] Furthermore, the structured 3D road model supports output in various general and professional formats to meet the needs of different downstream applications. For example, it can be output as 3D Tiles for web-based 3D visualization, as Wavefront Object File Format (OBJ) / Autodesk Filcombe Exchange Format (FBX) / Graphics Library Transfer Format (GLTF) for 3D simulation and rendering, as OpenDRIVE for autonomous driving simulation, or as Shapefile / GeoJSON for spatial analysis and asset management in GIS platforms. This multi-format output capability ensures the model's good applicability and interoperability in diverse scenarios such as intelligent transportation, infrastructure maintenance, and digital twins.
[0065] In this embodiment, compared to existing technologies that use LiDAR-based point cloud modeling for 3D road modeling, which suffer from the inability to accurately update local models when road conditions change, requiring the re-collection of point cloud data for the changed area and surrounding large-scale roads, and subsequent full-process reconstruction, this application uses a multi-view camera mounted on a mobile platform to acquire image sequences of the road to be modeled. Based on these image sequences, a virtual point cloud is generated. Compared to LiDAR, this acquisition method is not only lower in cost and more flexible in deployment, but also allows for rapid response to changes in the road, enabling targeted image re-capture only for the changed area. Furthermore, by orthographically projecting the virtual point cloud, the model to be created is generated. The digital orthophoto map and digital elevation model corresponding to the road are used to convert the digital orthophoto map and digital elevation model into a structured 3D model of the road surface using a road surface modeling model. On the other hand, a structured 3D model of the road above is generated directly based on the image sequence using a road above modeling model. The two modeling processes are independent of each other, so that when the road undergoes local changes, there is no need to re-collect data and remodel a large area. Only targeted re-photographs of the changed area are needed to obtain an updated image sequence. Then, the model corresponding to the changed area can be updated separately. Therefore, it can meet the actual needs of dynamic road updates and avoid the waste of resources caused by existing reconstruction technologies.
[0066] Figure 2 A flowchart illustrating a road modeling method according to another embodiment of this application is shown. Figure 2 As shown above, in the above Figure 1Based on the illustrated embodiment, one specific implementation of step S103 is as follows: S201: Input the digital orthophoto map into the visual big model, and use the visual big model to extract the vector contours of road surface elements to obtain two-dimensional element contour data with semantic information and geographic coordinates.
[0067] In this embodiment, the DOM is generated after eliminating distortions caused by terrain undulations and viewpoint tilt, and all geographic objects in the image are presented at true scale. This characteristic allows the road elements identified by the visual large model to have shapes and sizes closer to the actual scene, effectively avoiding geometric distortion problems in the recognition process. In addition, the visual large model itself has semantic recognition capabilities, and its output two-dimensional feature contour data carries semantic information and geographic coordinates, eliminating the need for additional coordinate transformation steps and significantly improving data processing efficiency.
[0068] It should be noted that vector data refers to data that describes the location and shape of map graphics or geographic entities in a Cartesian coordinate system using x and y coordinates. It typically reproduces the true spatial location of geographic entities to the greatest extent possible by accurately recording coordinates. In this application, the two-dimensional feature contour data output by the visual large model is vector data with two-dimensional coordinates. Since digital orthophoto maps use a unified geographic coordinate system, the extracted contours are directly represented as polygonal or polyline vectors with these unified geographic coordinates, without any subsequent projection transformation. In practical applications, its planar coordinate accuracy can reach 10 centimeters.
[0069] This process extracts engineering-usable geometric data such as road marking vectors and curb polygons. This data possesses precise geometric dimensions and location information, which can be directly applied to road engineering design, analysis, and decision-making. For example, in road planning, the extracted road marking vectors and curb polygons can be used to assess road capacity and safety.
[0070] In one example, the visual big model can be: the Segment Anything Model (SAM), or a derivative model based on it that is specifically optimized for road scenarios (such as RoadSAM).
[0071] In one example, two-dimensional feature contour data may include: lane lines, stop lines, directional arrows, ground text, curbs, pedestrian crossings, and median boundaries.
[0072] In one example, the obtained two-dimensional feature contour data can be manually corrected and attribute labeled to form structured semantic data. For example, key attribute information such as type (e.g., "left turn guide arrow"), color (e.g., "yellow solid lane line"), and length can be added to the extracted vector features to provide more accurate semantic support for subsequent three-dimensional modeling.
[0073] It should be noted that since the 2D feature outline data is generated based on DOM, and DOM not only presents all features at true scale but also comes with a unified geographic coordinate system, it can be directly imported into mainstream GIS platforms such as ArcGIS and QGIS. Therefore, this 2D feature outline data can be directly adapted to the visualization editing tools of GIS systems. With the help of the tool's built-in functions such as vector feature vertex editing, outline smoothing, and segment adjustment, manual intervention can accurately locate omissions (such as unrecognized short-distance markings) and recognition errors (such as outline distortion) that occur during the visual large model recognition process, and perform targeted completion or correction of the vector data. At the same time, the editing tool has a built-in attribute editing panel with preset standardized attribute fields that are highly matched with road features (such as "feature type," "color," "length," and "material"). Manual intervention can easily complete attribute supplementation and annotation through drop-down selection or manual input, ensuring the integrity and standardization of semantic data.
[0074] Figure 3 The process of semantic vector extraction from a large visual model based on the DOM is illustrated, such as... Figure 3 As shown, using the DOM as the input carrier, the visual large model performs automatic recognition processing on the DOM to obtain preliminary vector extraction and classification results, namely two-dimensional feature contour data. This result already carries the basic semantic labels and unified geographic coordinates of the features, realizing the preliminary binding of semantic information and spatial location. In order to further improve the accuracy of the two-dimensional feature contour data, manual correction and attribute labeling operations can be performed on the preliminary results. The manual correction and attribute labeling are completed using the built-in editable tools of GIS, and finally the structured vector data is output. This vector data is a polygon or polyline vector with unified geographic coordinates and complete semantic attributes, which can be directly used as the geometric constraints and semantic basis for the subsequent construction of the three-dimensional model of the road surface.
[0075] S202: The two-dimensional feature outline data is back-projected onto the digital elevation model based on a unified geographic coordinate system to obtain three-dimensional feature outline data with elevation coordinates.
[0076] In this embodiment of the application, the two-dimensional feature contour data obtained above only contains planar coordinates (X, Y) and lacks elevation information (Z). Since the DOM and DEM share a unified geographic coordinate system, the X and Y coordinates of the two-dimensional feature contour data can be directly back-projected onto the spatial area corresponding to the DEM to obtain three-dimensional feature contour data with supplemented elevation coordinates (Z value).
[0077] In one example, since some data points of the 2D feature contour data may not fall exactly on the grid vertices of the DEM, but may fall inside a DEM grid cell, the accurate elevation Z value of such data points can be calculated using a bilinear interpolation algorithm. Specifically: first, the corresponding DEM grid cell is determined based on the X and Y coordinates of the data point; then, the known elevation values of the four adjacent vertices of that grid cell are extracted; finally, the elevation values of the four vertices are weighted according to the relative position of the data point within the grid cell to obtain the accurate elevation Z value of the data point at its current position.
[0078] In another example, to reduce computational complexity, the nearest neighbor interpolation method can also be used. By calculating the distance between the data point and the four vertices of the grid cell, the vertex with the smallest distance is selected, and the elevation value of the vertex is directly used as the elevation value of the data point. This can significantly reduce the amount of computation and speed up data processing.
[0079] S203: Using the three-dimensional feature contour data as geometric constraint boundaries, constrain the surface fitting of the digital elevation model to generate a three-dimensional model of the structured road surface.
[0080] In this embodiment, three-dimensional feature contour data is used as geometric constraint boundaries to perform surface fitting on the DEM, thereby generating a structured road surface three-dimensional model. It should be noted that since the two-dimensional feature contour data itself contains semantic information, and the structured road surface three-dimensional model is obtained by supplementing the elevation of the two-dimensional feature contour data with three-dimensional feature contour data, combined with DEM constraint fitting, it also contains semantic information. It should also be noted that the structured road surface three-dimensional model supports output in standard formats such as FBX format, 3D tiles format, and OpenDRIVE (Open Dynamic Road Information Format for Vehicle Environment).
[0081] In one example, constrained surface fitting can be performed on a digital elevation model based on either a Triangulated Irregular Network (TIN) model or Non-Uniform Rational B-Splines (NURBS modeling). If constrained surface fitting is performed based on TIN, the resulting structured road surface 3D model is an editable triangular mesh, allowing direct adjustment of the vertex coordinates of individual triangles to achieve fine-grained correction of local morphology. If constrained surface fitting is performed based on NURBS, the resulting structured road surface 3D model is an editable parametric surface, allowing smooth modification of the road surface morphology by adjusting the position and weight of surface control points.
[0082] In another example, since both DOM and DEM can be directly imported into GIS systems such as ArcGIS and QGIS for editing, the structured road surface 3D model generated based on them (linked with semantic information and unified geographic coordinates) can also be directly adapted to the editing functions of GIS systems. Furthermore, because the model carries semantic information, semantic units can be queried, edited, and rendered using dedicated tools on the GIS platform. For example: using the attribute query tool of the GIS system, the target structured road surface 3D model can be filtered by semantic tags, and the system will automatically retrieve the attribute data associated with that structured road surface 3D model; through the GIS feature editing panel, selecting the target structured road surface 3D model (such as a lane line) allows direct modification of its attribute fields (such as changing "white dashed line" to "yellow solid line"), or adjusting the vertex coordinates of its 3D vector contour to correct the geometric shape; using the GIS symbolic rendering tool, rendering rules can be set based on semantic attribute fields (such as "feature type" and "color"), and the system will automatically perform differentiated rendering of different target structured road surface 3D models according to the rules (such as lane lines displayed according to the attribute label color, and pedestrian crossings filled with a specific texture), while also supporting the export of the rendered 3D scene view.
[0083] Figure 4 The flowchart illustrating the generation process of a 3D model of a structured road surface is shown, such as... Figure 4 As shown, the DOM is first input into the visual large model. The visual large model is then used to extract the vector outlines of road surface features from the DOM. Then, the geometric deviations of the vectors are manually corrected using GIS tools. Finally, a DOM semantic vector carrying semantic labels and unified geographic coordinates is obtained, which is a two-dimensional feature outline data with semantic information and geographic coordinates. Then, based on the elevation information provided by the DEM, it is fused with the DOM semantic vector to create a model. Finally, a structured three-dimensional road surface model containing the overall road model and subdivided semantic units such as lanes, markings, and materials is generated. This model not only restores the undulation of the road surface based on the DEM, but also inherits the complete semantic attributes of the DOM semantic vector, which can be directly used for subsequent editing, querying and other engineering operations.
[0084] In this embodiment, the digital orthophoto map is generated after eliminating distortions caused by terrain undulations and viewing angle tilt. Inputting it into a large visual model allows for the extraction of highly accurate two-dimensional feature contour data based on its geometrically distortion-free and realistically proportioned features. Furthermore, the large visual model possesses semantic segmentation capabilities, thus the obtained two-dimensional feature contour data naturally carries semantic information. Since the digital orthophoto map and the digital elevation model share the same unified geographic coordinate system, the two-dimensional feature contour data can be back-projected onto the digital elevation model to obtain three-dimensional feature contour data. This three-dimensional feature contour data relies on the semantic information of the two-dimensional feature contour data. Therefore, the structured road surface 3D model generated based on the three-dimensional feature contour data directly carries semantic information, avoiding the cumbersome steps of first constructing a geometric model and then separately performing semantic annotation in the traditional modeling process. This reduces the risk of misalignment between semantic and geometric information, significantly shortens the overall modeling cycle, and effectively improves the efficiency of road 3D modeling.
[0085] Figure 5 A flowchart illustrating a road modeling method according to another embodiment of this application is shown. Figure 5 As shown above, in the above Figure 1 Based on the illustrated embodiment, one specific implementation of step S104 is as follows: S501: Input the image sequence into the visual big model, use the visual big model to extract the road-above elements, and obtain the road-above elements with semantic information.
[0086] In this embodiment, by utilizing the semantic segmentation and object detection capabilities of a large visual model, road-above elements with semantic information can be extracted from image sequences. Examples include traffic signs, light poles, traffic signals, and fences.
[0087] In one example, to further improve the accuracy of identifying elements above the road and reduce the interference of irrelevant information about the road surface (such as lane lines and road stains) on the model recognition, the image sequence can be preprocessed first, and only image segments of the area above the road surface can be input into the large visual model. Specific processing methods can include: determining the boundary threshold between the road surface and the overhead area in the image using a road surface edge detection algorithm, or delineating the imaging area above the road surface based on the pose information of the mobile platform; precisely cropping the original image to remove redundant pixels in the road surface area; and then inputting the cropped image above the road surface into the large visual model for recognition.
[0088] In another example, because the time intervals between consecutive frames in an image sequence are extremely short, the same road overhead feature (such as a lamppost or a traffic sign) may appear repeatedly in multiple adjacent images. Directly retaining all recognition results would lead to feature redundancy, spatial location conflicts, and consequently, modeling chaos. In this case, based on the pose information corresponding to each frame, the spatial coordinates of the same type of feature identified in different frames can be calculated. By setting a distance threshold, repeatedly identified features can be matched and deduplicated, retaining only the unique road overhead feature.
[0089] S502: Extract feature information of elements above the road and use image retrieval technology to retrieve a structured 3D model above the road with a similarity greater than a preset threshold to the feature information from a standard model library.
[0090] In this embodiment, a large number of standardized 3D models above roads and their associated data are pre-stored in the standard model library. Using image retrieval technology, the corresponding structured 3D models above roads can be directly matched and retrieved from the model library based on the identified road features, thus avoiding the need for 3D reconstruction from scratch for each facility and significantly improving modeling efficiency.
[0091] In one example, image retrieval technology can combine feature extraction based on convolutional neural networks (CNNs) with fast nearest neighbor retrieval based on FAISS (Facebook AI Similarity Search). Specifically, firstly, feature vectors of identified road-above elements are extracted using a CNN model. Then, the FAISS tool is used to quickly retrieve the structured 3D road-above-the-road model with the highest similarity to the feature vectors of the road-above-road elements from a standard model library, achieving efficient matching. The standard model library can store structured road-above-the-road 3D model files supporting common formats such as FBX and GLTF.
[0092] In one example, to ensure the accuracy of the matching results, the validity of the matched 3D model above the structured road can be verified based on the semantic information of the previously identified road-above elements. For instance, the standard model library not only stores 3D model files but also associates and stores the semantic labels corresponding to each model. During matching, the semantic labels of the identified results can be compared with the semantic labels of the candidate models, prioritizing the model with the most consistent semantics and visual features, thereby improving the accuracy of the matching. It should be noted that the semantic information of the 3D model above the structured road can directly use the semantic information obtained in the previous identification or can use the semantic information pre-stored in the standard model library.
[0093] In another example, to further improve matching speed, the standard model library can be preprocessed and optimized. Feature information of all 3D models above standardized roads in the library can be extracted and stored in advance using a CNN model. When matching is performed later, there is no need to perform feature extraction operations on the model library resources in real time. The pre-stored model feature information can be directly called and compared with the feature vector of the element to be matched, thereby shortening the overall time of feature extraction and matching, and thus improving the efficiency of 3D modeling.
[0094] In another example, when no structured road overhead 3D model with a similarity greater than a preset threshold is matched in the standard model library (such as when encountering niche customized facilities, new road ancillary components, etc.), a standardized 3D model can be customized and newly created through manual intervention. After the model is created, it is synchronously fed back into the standard model library and added to the training set of the visual big model, thereby continuously enriching the coverage of the standard model library and continuously improving the visual big model's ability to recognize niche and new road overhead facilities.
[0095] In this embodiment, a large visual model is used to identify road-above elements in an image sequence. Image retrieval technology is then used to search a standard model library for structured 3D road-above models with a similarity greater than a preset threshold to the feature information. This avoids the repetitive 3D reconstruction of numerous standardized facilities, significantly improving the efficiency of overall road modeling. Furthermore, when a road-above facility changes, the update can be completed simply by replacing the corresponding structured road-above 3D model in the scene, without needing to re-collect data or perform a complete reconstruction process. This effectively avoids the waste of computational and storage resources caused by complete reconstruction and significantly shortens the model update response time.
[0096] Figure 6 The flowchart illustrates the generation of a 3D model above a structured road. (For example...) Figure 6 As shown, firstly, the acquired image sequences are automatically identified and classified using a large visual model to accurately locate and extract road overhead equipment and facilities (such as traffic signs, surveillance cameras, traffic lights, etc.) from the images, while assigning corresponding semantic classification labels to these elements. Then, combining the pose information corresponding to the image sequence, the 3D spatial coordinates of the road overhead equipment and facilities in a unified geographic coordinate system are calculated. This can be supplemented with manual correction to ensure that the obtained 3D spatial coordinates are consistent with the real scene. Next, based on image retrieval technology, the visual features of the identified road overhead equipment and facilities are matched with resources in a pre-built standard model library. Through retrieval, standardized models matching the road overhead equipment and facilities can be quickly obtained, avoiding reconstruction from scratch. Finally, the standardized models obtained through matching are instantiated by combining the previously corrected pose and classification information (i.e., 3D spatial coordinates and semantic information in a unified geographic coordinate system) to obtain a structured 3D model of the road overhead.
[0097] In one example, after obtaining the 3D model above the structured road, the process also includes: Calculate the parallax information of road-above features in image sequences acquired by multi-view cameras; In this embodiment, only the category and geometric information of the 3D model above the structured road are obtained, which is insufficient to determine its precise location in the real world. To achieve accurate positioning, the disparity information of each identified feature above the road in the image sequence acquired by the multi-view camera needs to be calculated. This information is the basis for subsequently estimating its 3D position in a unified geographic coordinate system.
[0098] In one example, the object detection or semantic segmentation capabilities of a large visual model can be used to identify and determine the pixel region corresponding to the road-above element in each frame of the image (such as the semantic mask or bounding box of the output element, which clarifies the specific pixel range of the element in the image). This limits the effective area for subsequent disparity calculation. Then, for image pairs (such as left and right frame images) acquired by multi-view cameras (such as binocular cameras), pixel-level matching calculation is performed only within the pixel region corresponding to the road-above element, using a dense stereo matching algorithm to obtain the disparity value corresponding to each pixel in the element region, forming a disparity map covering the entire element region.
[0099] In another example, considering that road features will appear continuously in multiple frames of images, the accuracy can be improved by multi-frame disparity fusion optimization: first, calculate the single-frame disparity information of the feature in each frame of the image, and then use optimization algorithms such as average value, median filtering or weighted average of the disparity values of multiple frames to finally determine the optimal disparity information of the road features. This can effectively reduce the impact of single-frame image noise, illumination interference or feature point matching errors, and significantly improve the accuracy of disparity calculation.
[0100] Based on parallax information, parameters of multi-view cameras, and pose information, the three-dimensional spatial coordinates of the features above the road in a unified geographic coordinate system are calculated; where the pose information is the pose information corresponding to the image frame where the features above the road are located.
[0101] In this embodiment, based on the parameters of the multi-view camera (baseline length, camera focal length), the depth value corresponding to each pixel of the road above feature can be calculated using the relationship formula between parallax and depth (depth = baseline × focal length / parallax). Then, combining the camera extrinsic parameters (including the rotation matrix and translation vector describing the camera's own spatial attitude) and the pose information corresponding to the image sequence, the depth value of each pixel is associated with the two-dimensional coordinates of the image, and converted into the three-dimensional coordinates of the feature in the camera coordinate system. Finally, based on the pose parameters of the camera in a unified geographic coordinate system, a coordinate transformation matrix is constructed to map the three-dimensional coordinates in the camera coordinate system to the unified geographic coordinate system, thereby obtaining the three-dimensional spatial coordinates of the road above feature in the unified geographic coordinate system.
[0102] In one example, to fully reconstruct the spatial orientation of road surface elements in a 3D scene (such as the installation tilt angle of traffic signs and the shooting direction of surveillance cameras), their pose information can be calculated simultaneously. The specific pose calculation process is as follows: First, the 3D feature points of the corresponding road surface element prefabricated models (such as the corner points of traffic lights and the vertices of the signboard borders) are retrieved from the standard model library. Then, the 2D pixel coordinates of these 3D feature points in multi-view images are matched. Using the Perspective-n-Point (PnP) algorithm combined with camera intrinsics, the rotation matrix and translation vector of the elements in the camera coordinate system are solved. Finally, a coordinate transformation matrix is used to map them to a unified geographic coordinate system to determine their pose parameters in the target scene (such as horizontal yaw angle and vertical pitch angle). Ultimately, the positional accuracy of the structured road surface 3D model in the unified geographic coordinate system can reach 10cm, and the pose accuracy can reach 1 degree.
[0103] In one example, to further ensure the geographic accuracy of the 3D spatial coordinates of features above the road, DOM can be used for auxiliary verification and revision. Since DOM is generated through orthographic projection and uses a unified geographic coordinate system, it not only clearly presents the road surface features, but also includes the orthographic projection images (i.e., the projection outlines and coordinates of the features on the horizontal plane) of features above the road (such as traffic signs and light poles). Theoretically, the planar projection (X,Y) of the 3D spatial coordinates (X,Z) of the features above the road in the unified geographic coordinate system should completely match the coordinates of the orthographic projection of the feature in the DOM. Therefore, this spatial relationship can be used to verify the accuracy of the coordinates.
[0104] The specific verification method is as follows: First, extract the two-dimensional coordinates of the feature points of the orthophoto projection of the road above the DOM (the vertex or center coordinates of the projection outline can be obtained through GIS tools). Then, project the calculated three-dimensional spatial coordinates of the road above the DOM vertically onto the horizontal plane (i.e., the two-dimensional plane where the DOM is located) to obtain the two-dimensional coordinates of the projection points. Finally, compare the Euclidean distance between the two-dimensional coordinates of the projection points and the two-dimensional coordinates of the feature points. If the Euclidean distance is less than a preset threshold, it is determined that the three-dimensional spatial coordinates of the road above the DOM are without deviation. If a deviation is determined, the coordinates of the orthophoto projection of the features in the DOM can be used as a reference to reverse-calibrate the planar components of the three-dimensional spatial coordinates.
[0105] Based on three-dimensional spatial coordinates, the three-dimensional model above the structured road is instantiated into the three-dimensional model of the structured road surface.
[0106] In this embodiment, based on the three-dimensional spatial coordinates of the road above features in a unified geographic coordinate system, the structured road above three-dimensional model matched from the standard model library can be instantiated into a structured road surface three-dimensional model, thereby obtaining a structured three-dimensional road model.
[0107] In another example, based on the three-dimensional spatial coordinates and attitude parameters of the road above elements calculated above, the three-dimensional model of the structured road above can be instantiated into the three-dimensional scene according to the real scene information to obtain a complete three-dimensional scene model of the road above facilities. Finally, the three-dimensional scene model of the road above facilities and the three-dimensional model of the structured road surface are spatially aligned and integrated according to a unified geographic coordinate system.
[0108] Figure 7 The process of integrating and outputting a hierarchical structured model is shown, such as... Figure 7 As shown, the road overhead equipment and facilities model is the previously generated 3D model of the road overhead facilities, which includes the structured 3D form of overhead facilities such as traffic lights, light poles, and surveillance cameras. It comes with unified geographic coordinates and complete semantic attributes. The road surface model (pavement layer + marking layer) is the structured 3D model of the road surface, which includes the pavement layer that supports the basic form of the road and the marking layer that includes elements such as lane lines and guide arrows. It is also built based on a unified geographic coordinate system and has accurate spatial location information. Therefore, based on the unified geographic coordinate system, the structured 3D model of the road overhead facilities and the structured 3D model of the road surface can be merged to obtain a layered structured 3D road model.
[0109] In this embodiment, after obtaining the three-dimensional model above the structured road, the precise three-dimensional spatial coordinates of the road's overhead elements in a unified geographic coordinate system are calculated. This allows the three-dimensional model above the structured road to be accurately placed at the corresponding spatial position on the three-dimensional model of the structured road surface. By using the unified geographic coordinate system as a reference constraint, the spatial relationship between the road's overhead facilities and the pavement and marking layers is ensured to perfectly match the real road scene, while also fully preserving the independent structural attributes of each layer model. If a certain type of road element needs to be updated subsequently, only the corresponding layer's model module needs to be modified individually, without needing to reconstruct the entire three-dimensional road model, which greatly improves the model's maintenance efficiency and engineering practicality.
[0110] Figure 8 A flowchart illustrating a road modeling method according to another embodiment of this application is shown. Figure 8 As shown above, in the above Figure 1 Based on the illustrated embodiment, a specific implementation method for generating virtual point clouds in a unified geographic coordinate system based on image sequences is as follows: S801: Performs dense stereo matching on each set of images simultaneously acquired by multi-view cameras to generate the corresponding disparity map.
[0111] In this embodiment, dense stereo matching is performed on each set of synchronized multi-view images to generate a pixel-by-pixel disparity map, reflecting the disparity value of each pixel in the image. SGM balances local matching accuracy and global consistency by calculating cost aggregation in different directions (such as horizontal and vertical), and optimizes the matched disparity map to remove noise points and mismatched points, ensuring the integrity and accuracy of the disparity map.
[0112] S802: Converts disparity maps into depth maps based on the focal length and baseline of multi-view cameras.
[0113] In this embodiment, parallax is the difference in pixel units. By using the focal length and baseline corresponding to the multi-view camera, the difference in each pixel unit in the parallax map can be converted into distance in the physical world, i.e., depth map.
[0114] S803: Based on the pose information corresponding to each group of images, the depth map is back-projected onto a unified geographic coordinate system to generate a local virtual point cloud.
[0115] In this embodiment, for the pose information corresponding to each group of images, the pixel points of each frame depth map can be combined with their corresponding depth values and camera intrinsic parameter matrices to calculate their three-dimensional coordinates in the camera coordinate system. Then, based on the transformation matrix of the camera coordinate system in the unified geographic coordinate system, each pixel point is converted into three-dimensional coordinates in the unified geographic coordinate system, thereby generating multiple frames of local virtual point clouds.
[0116] S804: Under a unified geographic coordinate system, the local virtual point clouds generated from each set of images are fused and denoised to generate a virtual point cloud of the road to be modeled under a unified geographic coordinate system.
[0117] In this embodiment of the application, the mobile platform collects tens of thousands of frames of images, which will generate a large number of overlapping and noisy local point clouds. All local point clouds are loaded into the same global geographic coordinate system. Since they have absolute coordinates, point cloud fusion and filtering and noise reduction can be performed on the overlapping areas to generate a virtual point cloud of the road to be modeled in a unified geographic coordinate system.
[0118] In this embodiment, by performing dense stereo matching on each group of images synchronously acquired by multi-view cameras, a disparity map is generated. Combined with the parameters and pose information of the multi-view cameras, the disparity map can be converted into a virtual point cloud in a unified geographic coordinate system. Compared to traditional LiDAR solutions, multi-view camera equipment is far less expensive than LiDAR, and subsequent processing does not rely on expensive professional point cloud software and highly specialized personnel; the geometric data required for 3D modeling can be generated solely through image algorithms. Furthermore, multi-view cameras offer flexible deployment and high acquisition efficiency, allowing for rapid re-acquisition based on changes in road infrastructure, thereby achieving dynamic updates to the road's 3D model. In contrast, LiDAR equipment is expensive, and the acquisition process is subject to numerous environmental and site constraints, resulting in relatively low efficiency. In scenarios requiring frequent updates, flexibility is insufficient. The image data relied upon in this application has a uniform format and a relatively moderate amount of data. After preprocessing, it can be directly used in subsequent modeling processes without complex additional preprocessing steps. In contrast, point cloud data acquired by LiDAR is usually an unstructured set of 3D coordinates, which is prone to problems such as data sparsity, uneven density, and the presence of a large number of invalid noise points. Its preprocessing process (such as denoising, registration, and hole filling) is cumbersome, requiring not only specialized optimization algorithms but also a large amount of computing resources to ensure the accuracy of subsequent modeling. Therefore, compared with the traditional method of modeling based on point clouds acquired by LiDAR, the virtual point cloud constructed by this application based on multi-view images can reduce the threshold of 3D modeling, improve the flexibility of dynamic updates, and increase data processing efficiency.
[0119] Figure 9 A road modeling system is shown, such as Figure 9As shown, the system includes: a binocular camera system 901, an image preprocessing module 902, a stereo matching module 903, a DOM / DEM generation module 904, a GNSS / IMU device 905, a visual large model semantic extraction module 906, a prefabricated model library 907, an overpass recognition module 908, a hierarchical model integration module 909, and a 3D road model output module 910. This road modeling system adopts a modular architecture design, and each functional module can be independently developed, tested, and deployed. The system supports elastic expansion and can flexibly add or delete processing modules or adapt to different hardware sensors according to actual scenario requirements. At the same time, the system supports edge computing deployment and can complete part or all of the data processing flow on a mobile acquisition platform or roadside computing unit to achieve real-time or near real-time 3D modeling and updating.
[0120] Binocular camera system 901: Responsible for acquiring stereoscopic images of the road, providing raw data for subsequent processing.
[0121] Image preprocessing module 902: Performs distortion correction, stereo correction and illumination normalization on the acquired images to improve image quality.
[0122] Stereo matching module 903: Employs the SGM algorithm for dense stereo matching to generate a disparity map, providing a foundation for subsequent depth calculations and point cloud generation.
[0123] DOM / DEM generation module 904: Based on virtual point cloud and GNSS / IMU pose information, it generates DOM and DEM, providing geographic reference for semantic extraction and model building.
[0124] GNSS / IMU device 905: Provides position and attitude information for mobile platforms for DOM and DEM generation and georegistration.
[0125] Visual large model semantic extraction module 906: Extracts the vector contours of road semantic elements from the DOM using the visual large model, providing support for the semanticization of the model.
[0126] Prefabricated Model Library 907: Stores standard 3D models of various road overpasses for model matching and instantiation.
[0127] Upper structure identification module 908: Identifies structures above the road surface, calculates their three-dimensional position and orientation, and matches them with the prefabricated model library.
[0128] Layered Model Integration Module 909: This module integrates the 3D model of the structured road surface and the 3D model above the structured road to form a layered 3D road model.
[0129] 3D Road Model Output Module 910: Outputs the generated 3D road model in multiple formats to meet the needs of different application scenarios.
[0130] Based on the road modeling method provided in the above embodiments, this application also provides a specific implementation of a road modeling device. Please refer to the following embodiments.
[0131] First see Figure 10 , Figure 10 The diagram shows a road modeling device structure provided in an embodiment of this application. The road modeling device 1000 provided in this embodiment includes: an acquisition module 1001, a projection module 1002, a first processing module 1003, a second processing module 1004, and an integration module 1005.
[0132] The acquisition module 1001 is used to acquire image sequences of the road to be modeled through a multi-view camera set on a mobile platform, and generate a virtual point cloud in a unified geographic coordinate system based on the image sequences. Each view camera corresponds to an image sequence, and each frame in the image sequence corresponds to a pose information. Projection module 1002 is used to perform orthophoto projection on virtual point cloud to generate digital orthophoto map and digital elevation model corresponding to the road to be modeled. The first processing module 1003 is used to input the digital orthophoto map and digital elevation model into the road surface modeling model, and use the road surface modeling model to generate a structured three-dimensional road surface model. The second processing module 1004 is used to input the image sequence into the road overhead model and use the road overhead model to generate a structured road overhead 3D model. The integration module 1005 is used to spatially associate and integrate the three-dimensional model of the structured road surface with the three-dimensional model above the structured road according to a unified geographic coordinate system, so as to generate a structured three-dimensional road model.
[0133] In one example, the first processing module 1003 includes: The input submodule is used to input digital orthophotos into the visual big model, and use the visual big model to extract the vector contours of road surface features to obtain two-dimensional feature contour data with semantic information and geographic coordinates. The back projection submodule is used to backproject two-dimensional feature outline data onto a digital elevation model based on a unified geographic coordinate system to obtain three-dimensional feature outline data with elevation coordinates. The fitting submodule is used to perform constrained surface fitting on the digital elevation model using the 3D feature contour data as geometric constraint boundaries, and generate a 3D model of the structured road surface.
[0134] In one example, the second processing module 1004 includes: The input submodule is also used to input image sequences into the visual big model, and use the visual big model to extract road-above elements to obtain road-above elements with semantic information. The extraction submodule is used to extract feature information of elements above the road and use image retrieval technology to search for structured 3D models above the road with a similarity greater than a preset threshold in a standard model library.
[0135] In one example, the road modeling device 1000 also includes: The calculation module is used to calculate the parallax information of road features in the image sequence acquired by the multi-view camera; The calculation module is also used to calculate the three-dimensional spatial coordinates of the road above features in a unified geographic coordinate system based on parallax information, parameters of multi-view cameras, and pose information; where pose information is the pose information corresponding to the image frame where the road above features are located. The instantiation module is used to instantiate the 3D model above the structured road onto the 3D model of the structured road surface based on 3D spatial coordinates.
[0136] In one example, obtaining module 1001 includes: The dense stereo matching submodule is used to perform dense stereo matching on each group of images acquired synchronously by multi-view cameras to generate the corresponding disparity map. The transformation submodule is used to convert the disparity map into a depth map based on the focal length and baseline corresponding to the multi-view camera. The backprojection submodule is used to backproject the depth map to a unified geographic coordinate system based on the pose information corresponding to each group of images, thereby generating a local virtual point cloud. The generation submodule is used to fuse and denoise the local virtual point clouds generated from each set of images under a unified geographic coordinate system, thereby generating a virtual point cloud of the road to be modeled under a unified geographic coordinate system.
[0137] In one example, the road modeling device 1000 also includes: The preprocessing module is used to perform distortion correction on each frame of the image sequence based on the intrinsic distortion coefficients of the multi-view camera. The preprocessing module is used to project each group of images onto the same epipolar plane using an epipolar correction algorithm; The preprocessing module is used to perform illumination homogenization processing on each frame of the image sequence using an adaptive histogram equalization algorithm.
[0138] Figure 11 A schematic diagram of the hardware structure of the electronic device provided in an embodiment of this application is shown.
[0139] An electronic device may include a processor 1101 and a memory 1102 storing computer program instructions.
[0140] Specifically, the processor 1101 may include a central processing unit (CPU), an application specific integrated circuit (ASIC), or one or more integrated circuits that can be configured to implement the embodiments of this application.
[0141] Memory 1102 may include mass storage for data or instructions. For example, and not limitingly, memory 1102 may include a hard disk drive (HDD), a floppy disk drive, flash memory, optical disk, magneto-optical disk, magnetic tape, or a Universal Serial Bus (USB) drive, or a combination of two or more of these. In one instance, memory 1102 may include removable or non-removable (or fixed) media, or memory 1102 may be a non-volatile solid-state memory.
[0142] In one instance, memory 1102 may be read-only memory (ROM). In one instance, the ROM may be a mask-programmed ROM, a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM), an electrically rewritable ROM (EAROM), or flash memory, or a combination of two or more of these.
[0143] Memory 1102 may include read-only memory (ROM), random access memory (RAM), disk storage media device, optical storage media device, flash memory device, electrical, optical, or other physical / tangible memory storage device. Therefore, generally, memory includes one or more tangible (non-transitory) computer-readable storage media (e.g., memory devices) encoded with software including computer-executable instructions, and when the software is executed (e.g., by one or more processors), it is operable to perform the operations described with reference to the method according to one aspect of this disclosure.
[0144] The processor 1101 implements a road modeling method in the above-described embodiment by reading and executing computer program instructions stored in the memory 1102.
[0145] In one example, the electronic device may also include a communication interface 11011 and a bus 1104. Wherein, as... Figure 11As shown, the processor 1101, memory 1102, and communication interface 11011 are connected through bus 1104 and complete communication with each other.
[0146] The communication interface 11011 is mainly used to realize communication between various modules, devices, units and / or equipment in the embodiments of this application.
[0147] Bus 1104 includes hardware, software, or both, that couples components of an online data traffic metering device together. For example, and not as a limitation, the bus may include an Accelerated Graphics Port (AGP) or other graphics bus, an Extended Industry Standard Architecture (EISA) bus, a Front Side Bus (FSB), a Hyper Transport (HT) interconnect, an Industry Standard Architecture (ISA) bus, an Infinite Bandwidth Interconnect, a Low Pin Count (LPC) bus, a memory bus, a Microchannel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI-Express (PCI-X) bus, a Serial Advanced Technology Attachment (SATA) bus, a Video Electronics Standards Association Local (VLB) bus, or other suitable buses, or combinations of two or more of these. Where appropriate, bus 1104 may include one or more buses. Although specific buses are described and illustrated in embodiments of this application, this application contemplates any suitable bus or interconnect.
[0148] The road modeling method described in the above embodiments can be implemented using a computer-readable storage medium provided in this application. The computer storage medium stores computer program instructions; when these computer program instructions are executed by a processor, they implement any of the road modeling methods described in the above embodiments.
[0149] This application also provides a computer program product, including a computer program that, when executed by a processor, implements any of the road modeling methods described in the above embodiments.
[0150] It should also be noted that the exemplary embodiments mentioned in this application describe methods or systems based on a series of steps or apparatus. However, this application is not limited to the order of the above steps; that is, the steps can be performed in the order mentioned in the embodiments, or in a different order, or several steps can be performed simultaneously.
[0151] The aspects of this disclosure have been described above with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this disclosure. It should be understood that each block in the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that these instructions, executable via the processor of the computer or other programmable data processing apparatus, enable the implementation of the functions / actions specified in one or more blocks of the flowchart illustrations and / or block diagrams. Such a processor can be, but is not limited to, a general-purpose processor, a special-purpose processor, a special application processor, or a field-programmable logic circuit. It is also understood that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can also be implemented by special-purpose hardware performing the specified functions or actions, or can be implemented by a combination of special-purpose hardware and computer instructions.
[0152] The above are merely specific embodiments of this application. Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, modules, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here. It should be understood that the protection scope of this application is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in this application, and these modifications or substitutions should all be covered within the protection scope of this application.
Claims
1. A road modeling method, characterized in that, include: Image sequences of the road to be modeled are acquired by multi-view cameras set on a mobile platform, and virtual point clouds under a unified geographic coordinate system are generated based on the image sequences. Each view camera corresponds to an image sequence, and each frame in the image sequence corresponds to a pose information. Orthophoto projection is performed on the virtual point cloud to generate a digital orthophoto map and a digital elevation model corresponding to the road to be modeled. The digital orthophoto image and the digital elevation model are input into the road surface modeling model, and the road surface modeling model is used to generate a structured three-dimensional road surface model. The image sequence is input into the road overhead model, and the road overhead model is used to generate a structured three-dimensional road overhead model; The structured road surface 3D model and the structured road above 3D model are spatially associated and integrated according to the unified geographic coordinate system to generate a structured 3D road model.
2. The method according to claim 1, characterized in that, The step of inputting the digital orthophoto image and the digital elevation model into the road surface modeling model, and using the road surface modeling model to generate a structured three-dimensional road surface model, includes: The digital orthophoto image is input into a large visual model, and the vector contours of road surface elements are extracted using the large visual model to obtain two-dimensional element contour data with semantic information and geographic coordinates. The two-dimensional feature outline data is back-projected onto the digital elevation model based on the unified geographic coordinate system to obtain three-dimensional feature outline data with elevation coordinates. Using the three-dimensional feature contour data as geometric constraint boundaries, the digital elevation model is fitted with a constraint surface to generate the three-dimensional model of the structured road surface.
3. The method according to claim 1, characterized in that, The step of inputting the image sequence into the road overhead model and generating a structured 3D road overhead model using the road overhead model includes: The image sequence is input into a large visual model, and the large visual model is used to extract the road overhead features to obtain road overhead features with semantic information. The feature information of the elements above the road is extracted, and image retrieval technology is used to retrieve the three-dimensional model above the structured road that has a similarity greater than a preset threshold with the feature information from a standard model library.
4. The method according to claim 3, characterized in that, After retrieving a 3D model above a structured road with a similarity greater than a preset threshold to the aforementioned feature information from a standard model library using image retrieval technology, the process further includes: Calculate the parallax information of the road above features in the image sequence acquired by the multi-view camera; Based on the parallax information, the parameters of the multi-view camera, and the pose information, the three-dimensional spatial coordinates of the road above features in the unified geographic coordinate system are calculated; wherein, the pose information is the pose information corresponding to the image frame where the road above features are located; Based on the three-dimensional spatial coordinates, the three-dimensional model above the structured road is instantiated onto the three-dimensional model of the structured road surface.
5. The method according to claim 1, characterized in that, The generation of a virtual point cloud in a unified geographic coordinate system based on the image sequence includes: Dense stereo matching is performed on each group of images synchronously acquired by the multi-view camera to generate a corresponding disparity map; Based on the focal length and baseline corresponding to the multi-view camera, the disparity map is converted into a depth map; Based on the pose information corresponding to each group of images, the depth map is back-projected onto the unified geographic coordinate system to generate a local virtual point cloud; Under the unified geographic coordinate system, the local virtual point clouds generated from each set of images are fused and denoised to generate the virtual point cloud of the road to be modeled under the unified geographic coordinate system.
6. The method according to claim 1, characterized in that, After acquiring image sequences of the road to be modeled using a multi-view camera set up on a mobile platform, the process also includes: The image sequence is preprocessed, and the preprocessing includes at least one of the following: Based on the intrinsic distortion coefficients of the multi-view camera, distortion correction processing is performed on each frame of the image sequence. An epipolar correction algorithm is used to project each group of images onto the same epipolar plane; An adaptive histogram equalization algorithm is used to perform illumination equalization processing on each frame of the image sequence.
7. A road modeling device, characterized in that, The device includes: The acquisition module is used to acquire image sequences of the road to be modeled through a multi-view camera set on a mobile platform, and generate a virtual point cloud in a unified geographic coordinate system based on the image sequences. Each view camera corresponds to an image sequence, and each frame in the image sequence corresponds to a pose information. The projection module is used to perform orthophoto projection on the virtual point cloud to generate a digital orthophoto image and a digital elevation model corresponding to the road to be modeled. The first processing module is used to input the digital orthophoto image and the digital elevation model into the road surface modeling model, and use the road surface modeling model to generate a structured three-dimensional road surface model. The second processing module is used to input the image sequence into the road overhead model and use the road overhead model to generate a structured road overhead 3D model. An integration module is used to spatially associate and integrate the three-dimensional model of the structured road surface with the three-dimensional model above the structured road according to the unified geographic coordinate system to generate a structured three-dimensional road model.
8. An electronic device, characterized in that, The device includes: a processor and a memory storing computer program instructions; the processor, when executing the computer program instructions, implements a road modeling method as described in any one of claims 1-6.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer program instructions, which, when executed by a processor, implement a road modeling method as described in any one of claims 1-6.
10. A computer program product, characterized in that, When the instructions in the computer program product are executed by the processor of the electronic device, the electronic device is able to perform a road modeling method as described in any one of claims 1-6.