A method for automatically generating measurable real-scene images of traffic elements

By collecting high-precision point cloud data and panoramic images, combined with deep learning and orthogonal correction technology, measurable projection maps and three-dimensional coordinates are generated, which solves the problems of low data collection efficiency and high modeling costs in traffic element management, and realizes high-precision and efficient real-scene image generation.

CN120431272BActive Publication Date: 2025-09-16BEIJING VISION ZHIXING TECH CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510932881.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-08
Publication Date
2025-09-16
Estimated Expiration
2045-07-08

AI Technical Summary

Technical Problem

Existing technologies in transportation element asset management have problems such as high manual collection costs, low efficiency, slow data updates, complex post-processing of mobile scanning technology data, high three-dimensional modeling costs and large differences between results and the current situation, and cannot meet the requirements of high precision and timeliness.

Method used

Using high-precision point cloud data and panoramic image acquisition, the system extracts contour lines through orthogonal correction, depth map generation, and deep learning to generate a three-dimensional model. It then performs automatic classification and parameter setting to ultimately generate measurable projection maps and three-dimensional coordinates.

Benefits of technology

It has achieved high-precision and rapid generation of real-life images of traffic elements, reduced system deviation, improved data accuracy and user experience, achieved a success rate of 99.8%, and solved the problems of high cost and difficulty in updating of 3D modeling.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120431272B_ABST
    Figure CN120431272B_ABST
Patent Text Reader

Abstract

The present invention provides a method for automatically generating measurable real-life images of traffic elements, which relates to the field of intelligent transportation technology. It includes: collecting data and pre-processing to obtain panoramic images after orthogonal correction, cleaned high-precision point cloud data and depth maps; extracting the contour lines of traffic elements, and generating a three-dimensional model of the target traffic element based on the contour lines of the traffic elements; automatically classifying the extracted traffic elements according to the three-dimensional model, and setting reasonable parameters to obtain the classification results of the traffic elements; generating an optimal panoramic image based on the classification results of the traffic elements; generating a measurable projection map based on the optimal panoramic image; and calculating the three-dimensional coordinates of the target traffic element based on the measurable projection map and the recorded parameters. The present invention solves the multiple problems faced by traditional traffic element asset management, such as low manual collection efficiency, complex post-processing of mobile scanning, high cost of three-dimensional modeling, and a large gap with reality.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of intelligent transportation technology, and in particular to a method for automatically generating measurable real-scene images of traffic elements. Background Art

[0002] In the development of transportation asset management, traditional technologies face many problems:

[0003] (1) Limitations of manual collection: In the past, manual collection (photography) was used to obtain transportation asset data, which was costly and inefficient. Furthermore, the timeliness of collection was difficult to guarantee, and data updates were slow. When faced with large amounts of data, maintenance work was almost impossible. Due to the subjectivity and limitations of manual collection, it cannot meet the high-precision and timeliness requirements of modern transportation asset management.

[0004] (2) Existing problems of mobile scanning technology: Although the technical means based on mobile scanning (including vehicle-mounted scanning and backpack scanning) have replaced manual collection to a certain extent and transferred the field data collection work to the office, it has greatly improved the efficiency of data acquisition and reduced some costs. However, the data post-processing (in-house work) link is still complex. Although point cloud data can achieve three-dimensional measurement, due to its poor visibility, it is difficult for industry users to operate and understand it, and the user experience is poor. Although the panoramic image of the scanning system can provide real-life information of traffic elements, it has a large deformation problem and does not meet the requirements of asset archiving. In addition, the same traffic element often corresponds to multiple panoramic images, making it difficult to select the optimal image and unable to achieve object-oriented management.

[0005] (3) Dilemma of 3D modeling technology: Although 3D modeling of traffic elements based on mobile scanning data can construct 3D models, the modeling cost is extremely high and model updating is extremely difficult. More importantly, the modeling results are significantly different from the actual status of traffic elements, which cannot meet the requirements of traffic asset management for the current status of elements and is difficult to play an effective role in actual asset management work. Summary of the Invention

[0006] In order to overcome the deficiencies of the prior art, the present invention aims to provide a method for automatically generating measurable real-scene images of traffic elements.

[0007] To achieve the above object, the present invention provides the following solutions:

[0008] A method for automatically generating measurable real-scene images of traffic elements, comprising:

[0009] Collect high-precision point cloud data and panoramic images and perform preprocessing to obtain orthogonally corrected panoramic images, cleaned high-precision point cloud data and depth maps;

[0010] Extracting the contours of traffic elements based on the cleaned high-precision point cloud data, the orthogonally corrected panoramic image, and the depth map, and generating a three-dimensional model of the target traffic element based on the contours of the traffic element;

[0011] Automatically classify the extracted traffic elements according to the three-dimensional model, and set reasonable parameters to obtain classification results of the traffic elements;

[0012] Generate the optimal panoramic image based on the classification results of traffic elements;

[0013] generating a measurable projection map based on the optimal panoramic image;

[0014] The three-dimensional coordinates of the target traffic element are calculated based on the measurable projection map and the recorded parameters.

[0015] Preferably, the high-precision point cloud data and panoramic images are collected using a vehicle-mounted mobile measurement system and a single-soldier mobile measurement system.

[0016] Preferably, the method for acquiring the panoramic image after orthogonal correction is:

[0017] Construct an initial spherical coordinate system;

[0018] Determine the three-dimensional coordinates corresponding to the panoramic image, wherein the three-dimensional coordinates are: ;

[0019] x original ,y original , z original When the panoramic pixel is converted into a panoramic sphere, the values ​​of the x-axis, y-axis, and z-axis in the three-dimensional coordinates corresponding to the pixel are taken as the radius;

[0020] Determine a corresponding rotation matrix according to the initial spherical coordinate system and the panoramic image;

[0021] Determine the corresponding inverse matrix according to the rotation matrix;

[0022] A new three-dimensional coordinate model corresponding to the panoramic image is determined according to the inverse matrix and the three-dimensional coordinates; wherein the expression of the new three-dimensional coordinate model is:

[0023] ;

[0024] in, is the inverse matrix;

[0025] An orthogonally corrected panoramic image is determined according to the new three-dimensional coordinate model.

[0026] Preferably, the depth map uses RGB three bands to store depth information.

[0027] Preferably, the extraction of the contour lines of traffic elements based on the cleaned high-precision point cloud data, the orthogonally corrected panoramic image and the depth map, and the generation of a three-dimensional model of the target traffic element according to the contour lines of the traffic elements include:

[0028] Construct a traffic element image dataset;

[0029] Use convolutional neural networks to train and validate the collected panoramic images to extract the contours of traffic elements;

[0030] Determine the corresponding representation point according to the extracted contour line;

[0031] A three-dimensional model of the target traffic element is determined according to the representation points.

[0032] Preferably, the convolutional neural network comprises:

[0033] U-Net and Mask R-CNN.

[0034] Preferably, the extracted traffic elements are automatically classified according to the three-dimensional model, and reasonable parameters are set to obtain the classification results of the traffic elements, including:

[0035] determining a classification standard for traffic elements based on the three-dimensional model;

[0036] Training a preset deep learning model according to the classification criteria to obtain a classification model;

[0037] A classification result is determined according to the classification model and parameters of various traffic elements are set according to the classification result, wherein the parameters of various traffic elements include: a classification confidence threshold, a shape similarity score and a geometric feature threshold.

[0038] Preferably, generating an optimal panoramic image according to the classification results of traffic elements includes:

[0039] Based on the classification results of traffic elements, various types of traffic elements are screened within a preset distance range to obtain a filtered panoramic image;

[0040] evaluating the filtered panoramic image to obtain an evaluation result;

[0041] According to the evaluation result, the panoramic images with occlusions are eliminated from the filtered panoramic images to generate an optimal panoramic image.

[0042] Preferably, generating a measurable projection map based on the optimal panoramic image includes:

[0043] Using the selected optimal panoramic image, the outline of the target traffic element is projected onto the preset projection surface through the camera pose information to obtain the initial projection map;

[0044] Calculating the actual physical position of each pixel in the initial projection image according to preset projection parameters;

[0045] A measurable projection map is generated based on the actual physical location.

[0046] Preferably, calculating the three-dimensional coordinates of the target traffic element based on the measurable projection map and the recorded parameters includes:

[0047] Extract the position information of each pixel in the projection image, and perform normalized coordinate transformation according to the projection parameters to obtain normalized coordinates;

[0048] Determine the actual coordinates of each pixel in the world coordinate system using the preset projection surface parameters and normalized coordinates;

[0049] The three-dimensional coordinates of the target traffic element are determined according to the actual coordinates of each pixel point in the world coordinate system.

[0050] The present invention discloses the following technical effects:

[0051] The present invention provides a method for automatically generating measurable real-world images of traffic elements, comprising: collecting high-precision point cloud data and panoramic images and preprocessing them to obtain orthogonally corrected panoramic images, cleaned high-precision point cloud data, and depth maps; extracting the contours of traffic elements based on the cleaned high-precision point cloud data, orthogonally corrected panoramic images, and depth maps, and generating a three-dimensional model of the target traffic element based on the contours of the traffic elements; automatically classifying the extracted traffic elements based on the three-dimensional model and setting reasonable parameters to obtain a classification result of the traffic elements; generating an optimal panoramic image based on the classification result of the traffic elements; generating a measurable projection map based on the optimal panoramic image; and calculating the three-dimensional coordinates of the target traffic element based on the measurable projection map and recorded parameters. The present invention integrates a precise positioning strategy: "Image detection" and "point cloud positioning" are integrated to solve the problems of simple images lacking three-dimensional coordinates and the low success rate of point cloud extraction, achieving precise spatial positioning and attitude determination. An innovative solution ensures data accuracy: combining methods such as obtaining "element representation points" reduces the error rate of element boundary projections, mitigates the impact of system deviations, and provides more accurate data. Scientific Classification Improves Processing Quality: Scientifically classifying traffic elements facilitates subsequent processes and enhances the quality of projection map generation. Solving business challenges: "Measurable" + "Realistic" technologies overcome the high cost, difficulty in updating, and low fidelity of 3D modeling, providing a direct view of element status. Highlighting application value: Testing across thousands of kilometers of projects in Shenzhen demonstrated a full-element extraction success rate exceeding 99.8%, demonstrating strong practicality, high stability, and broad application prospects. BRIEF DESCRIPTION OF THE DRAWINGS

[0052] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0053] Figure 1 A flow chart of a method for automatically generating measurable real-scene images of traffic elements provided by an embodiment of the present invention;

[0054] Figure 2 A schematic diagram of a framework of a method for automatically generating measurable real-scene images of traffic elements provided by an embodiment of the present invention;

[0055] Figure 3 A schematic diagram of generating vertical projection elements from spatial elements provided by an embodiment of the present invention;

[0056] Figure 4 A schematic diagram of key parameters related to generating vertical projection elements from spatial elements provided by an embodiment of the present invention;

[0057] Figure 5 A schematic diagram of a process for generating a projection image from a panoramic sphere provided by an embodiment of the present invention;

[0058] Figure 6 A schematic diagram of the actual length corresponding to each pixel in the projection image provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0059] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0060] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the present invention is further described in detail below with reference to the accompanying drawings and specific embodiments.

[0061] like Figure 1-2 As shown, the present invention provides a method for automatically generating measurable real-scene images of traffic elements, comprising:

[0062] Step 100: Collect high-precision point cloud data and panoramic images and perform preprocessing to obtain orthogonally corrected panoramic images, cleaned high-precision point cloud data, and depth maps;

[0063] Step 200: extracting the contours of traffic elements based on the cleaned high-precision point cloud data, the orthogonally corrected panoramic image, and the depth map, and generating a three-dimensional model of the target traffic element based on the contours of the traffic elements;

[0064] Step 300: Automatically classify the extracted traffic elements according to the three-dimensional model, and set reasonable parameters to obtain classification results of the traffic elements;

[0065] Step 400: Generate an optimal panoramic image based on the classification results of traffic elements;

[0066] Step 500: generating a measurable projection map based on the optimal panoramic image;

[0067] Step 600: Calculate the three-dimensional coordinates of the target traffic element based on the measurable projection map and the recorded parameters.

[0068] Specifically, the above methods can be summarized into the following categories:

[0069] Data Collection and Preprocessing: Vehicle-mounted scanning technology and backpack-mounted scanning technology are used to collect high-precision point cloud data and synchronized panoramic images, which are then flexibly integrated. Preprocessing then proceeds, including point cloud cleaning, panoramic image orthogonal correction, and high-precision depth map generation. Panoramic image orthogonal correction constructs an initial spherical coordinate system into which the original image pixels are projected, addressing non-standard poses while preserving heading values. Depth map generation utilizes a matrix transformation method and uses three RGB bands to store depth information, reducing computational complexity and improving storage accuracy.

[0070] Element Contour Extraction and 3D Conversion: Based on the characteristics of traffic elements, deep learning algorithms are used in conjunction with a specially constructed traffic element image dataset to train a deep learning model to accurately extract traffic element contours from orthogonally transformed panoramic images. A combination of morphological erosion, distance transformation, and safe zone sampling is used to obtain "element characterization points" away from the boundary, minimizing the impact of sensor bias. Based on the depth map, the depth values ​​of these "characterization points" are then calculated and converted to an absolute coordinate system to fit the 3D plane where the element resides.

[0071] Traffic Element Classification: Automatically classify all road elements into categories such as surface on-line, strip on-line, parallel strip, road surface unit, and general unit. This classification helps set appropriate parameters for the projection plane, resulting in a more natural and aesthetically pleasing projection angle. It also enables end-to-end output of strip-type elements during deep learning training.

[0072] Optimal panoramic image search and projection map generation: Based on the feature's projection plane, normal, and bounding rectangle, the optimal panoramic image is searched through distance screening, angle screening, and occlusion analysis. Then, according to a specific generation principle, the optimal panoramic image is calculated pixel by pixel on the projection surface to generate a measurable projection map.

[0073] Three-dimensional measurement based on projection maps: Although projection maps do not have a fixed sampling distance, through the recorded parameters, such as camera height h, the minimum horizontal distance d from the element surface's enclosing rectangle to the camera, the distance f from the projection surface to the camera, the angle α between the element surface and the horizontal plane, the rolling angle θ between the projection surface and the element surface, etc., through coordinate conversion, rotation and translation operations, the corresponding three-dimensional coordinates on the spatial plane can be calculated from the pixel points on the projection surface, thereby realizing the three-dimensional measurability of traffic elements.

[0074] Furthermore, the high-precision point cloud data and panoramic images are collected using a vehicle-mounted mobile measurement system and a single-soldier mobile measurement system.

[0075] Specifically, a mobile measurement system (including a vehicle-mounted mobile measurement system and a single-soldier mobile measurement system) is used to scan and photograph traffic roads, and a (general) combined navigation algorithm is used to calculate point clouds with absolute coordinates and synchronously registered panoramic images with poses. Three data processing processes are then performed: "point cloud cleaning", "panoramic image orthogonal correction", and "high-precision depth map generation".

[0076] Furthermore, the method for obtaining the panoramic image after orthogonal correction is:

[0077] Construct an initial spherical coordinate system;

[0078] Determine the three-dimensional coordinates corresponding to the panoramic image, wherein the three-dimensional coordinates are: ;

[0079] x original ,y original , z original When the panoramic pixels are converted into a panoramic sphere, the values ​​of the x-axis, y-axis, and z-axis in the three-dimensional coordinates corresponding to the pixels are taken as the radius.

[0080] Determine a corresponding rotation matrix according to the initial spherical coordinate system and the panoramic image;

[0081] Determine the corresponding inverse matrix according to the rotation matrix;

[0082] A new three-dimensional coordinate model corresponding to the panoramic image is determined according to the inverse matrix and the three-dimensional coordinates; wherein the expression of the new three-dimensional coordinate model is:

[0083] ;

[0084] in, is the inverse matrix;

[0085] An orthogonally corrected panoramic image is determined according to the new three-dimensional coordinate model.

[0086] Specifically, the present invention aims to solve the problem of non-standard posture in panoramic images captured by mobile scanning systems. This non-standard posture can be decomposed into three axial rotation angles, namely heading, pitch, and roll, all of which are non-zero values. In particular, when the pitch and roll angles are not zero, the presentation of ground objects in the panoramic image becomes irregular, which brings great trouble to the subsequent extraction of traffic elements. In order to overcome this problem, the present invention constructs an initial spherical coordinate system (that is, in this coordinate system, heading, pitch, and roll are all set to zero), and projects each pixel in the original image into this initial spherical coordinate system one by one based on the posture information of the acquired panoramic image, thereby generating a new, orthogonally corrected panoramic image.

[0087] The main calculation process includes coordinate system conversion and pixel reprojection:

[0088] Let R be the rotation matrix corresponding to the panoramic image pose acquired by the mobile scanning system. This matrix describes the rotation of the panoramic image relative to the initial spherical coordinate system. In the initial spherical coordinate system, the rotation matrix corresponding to the pose is the identity matrix (the identity matrix indicates no rotation, that is, the heading, pitch, and roll angles are all zero).

[0089] For any pixel in the panoramic image, its three-dimensional coordinates in the original image coordinate system are expressed as column vectors , to project it into the initial spherical coordinate system, we get the new coordinates .

[0090] According to the properties of the rotation matrix, the relationship between coordinate transformations is:

[0091] ;

[0092] in, Is the inverse matrix of the rotation matrix R. Since the rotation matrix is ​​an orthogonal matrix, it satisfies ( for The transposed matrix of , that is, the inverse matrix of the rotation matrix R can be obtained by transposing it.

[0093] The above matrix multiplication operation is performed on every pixel in the panoramic image, transforming each pixel from the original coordinate system with attitude deviation to the initial spherical coordinate system (orthogonal coordinate system). This completes the pixel-by-pixel projection of the panoramic image, ultimately generating a new panoramic image that conforms to the requirements of the initial spherical coordinate system. The three rotation angles of the new panoramic sphere (the three-dimensional panoramic sphere model formed by the panoramic image) are all zero, and the coordinates of the sphere center remain unchanged. The present invention retains the heading value (i.e., the heading value) of the three rotation angles and only resets the pitch and roll values ​​to zero.

[0094] Furthermore, the depth map uses RGB three bands to store depth information.

[0095] Specifically, the surrounding point cloud is transformed into a panoramic spherical coordinate system. Then, the spherical mesh corresponding to each laser point is calculated. This mesh only records the distance to the sphere's center, retaining only the minimum value. After calculating the entire point cloud, the value stored in the mesh is the depth. The computational complexity of this solution is (panoramic matrix × number of laser points), and the rest is just simple comparison. The specific calculation process is as follows:

[0096] Get the inverse matrix of the panoramic spherical pose ;

[0097] Calculate the coordinates of the laser point under the panoramic sphere: ;

[0098] Calculate the latitude and longitude of the laser point in the panoramic spherical coordinate system and Length , which is the depth map, can be expressed as: ; That is, the vector from the center of the panoramic sphere to the laser point cloud.

[0099] When there are multiple laser points corresponding to the same When taking the depth the smaller one;

[0100] In addition, the storage of depth maps generally adopts its own format or grayscale image format. The former has poor visual effects and the latter has insufficient storage accuracy. To this end, the present invention proposes a method for storing depth information in RGB three bands. Since the laser range (depth) in the field of mobile scanning is less than 1000 meters, the present invention designs a specific storage method: using the red (R (u, v)) channel to store meter-level depth information, the green (G (u, v)) channel to store decimeter to centimeter-level depth information, and the blue (B (u, v)) channel to store millimeter-level depth information. This storage method can record depth data efficiently and accurately within a limited RGB color space. For a point with coordinates (u, v) projected onto the depth map, its depth value D i Assign to RGB channels:

[0101] .

[0102] Furthermore, the extraction of the contour lines of traffic elements based on the cleaned high-precision point cloud data, the orthogonally corrected panoramic image and the depth map, and the generation of a three-dimensional model of the target traffic element according to the contour lines of the traffic elements include:

[0103] Construct a traffic element image dataset;

[0104] Use convolutional neural networks to train and validate the collected panoramic images to extract the contours of traffic elements;

[0105] Determine the corresponding representation point according to the extracted contour line;

[0106] A three-dimensional model of the target traffic element is determined according to the representation points.

[0107] Specifically, based on the characteristics of traffic elements, advanced deep learning algorithms are used to accurately detect traffic elements from panoramic images after orthogonal transformation, and further extract their contour lines and feature lines.

[0108] In order to ensure the training effect of the deep learning model, the present invention constructs a special traffic element image dataset, covering a rich variety of traffic element categories, including road markings, traffic signs, retaining walls and guardrails, vehicle barriers, isolation belts, anti-glare panels, sidewalks, safety islands, manhole covers, warning posts, gantries, lighting street lamps, sewer grates, etc., and collects panoramic images of these traffic elements in different scenes to fully reflect various situations in actual applications.

[0109] Classic convolutional neural network models, such as U-Net and Mask R-CNN, are trained on the constructed dataset. During training, model parameters are continuously adjusted to enable the model to accurately identify and extract traffic elements and their contours in panoramic images. These models, with their powerful feature extraction capabilities, effectively capture detailed information about traffic elements, resulting in highly accurate contour extraction. This paper employs a holistic approach that integrates traditional image processing and deep learning methods to extract traffic element contours. First, preprocessing enhances target features through denoising, color space conversion, and adaptive threshold segmentation. Differentiated strategies are then employed for different elements. For poles (such as streetlights), vertical boundaries are extracted using a multi-scale feature fusion network (such as U-Net) combined with morphological operations. For faces (such as road signs), geometric contours are fitted after segmentation using a semantic segmentation model (such as Mask R-CNN). For strips (such as guardrails), parallel boundaries are extracted by fusing point cloud data with image geometric constraints (such as Hough transform and RANSAC). Key techniques include deep learning object detection / segmentation models, Canny / Sobel edge detection, and polygon simplification algorithms. Post-processing optimizes broken connections and noise filtering.

[0110] After extracting the contours of traffic elements, it is necessary to obtain "element representation points" within the contours. These "representation points" must belong to the element and be far away from the element's boundaries. This is because in actual mobile scanning systems, the image of the element and the laser system will have a certain deviation (e.g., about 0.1 to 0.5 degrees). Directly using the element contour (boundary line) will result in inaccurate corresponding depth. After extracting the traffic element contours, the present invention obtains "element representation points" far from the boundary through the following methods to reduce the impact of sensor bias: ① Morphological erosion (such as adaptive kernel erosion to shrink the boundary area, but it is necessary to avoid excessive shrinkage that may cause loss of key information); ② Distance transformation + safe area sampling (calculating the farthest distance from the pixel to the boundary, screening high-distance areas and uniformly sampling, balancing accuracy and efficiency); ③ Geometric feature sampling (extracting the core points of regularly shaped elements based on the minimum enclosing rectangle / ellipse or grid method); ④ Deep learning key point detection (end-to-end regression of the coordinates of representation points far from the boundary, adapting to complex scenarios but relying on data training); ⑤ Inverse compensation of image deviation by combining LiDAR point cloud. This method uses a combination of approaches: distance transforms are used to generate safe zones and uniform sampling for conventional scenarios; corrosion preprocessing and deep learning fine-tuning are used for complex dynamic scenarios, with LiDAR registration compensation employed. Furthermore, sensor deviation angles are quantified and the safe distance threshold is dynamically adjusted. The optimization results are validated using non-maximum suppression (NMS) and projection error. If the resulting "representation points" are densely packed, appropriate downsampling can be employed.

[0111] Furthermore, after obtaining the pixel "representation point", the three-dimensional plane where the feature is located is obtained based on the above depth map. The specific algorithm is as follows:

[0112] ① Depth value acquisition: For the pixel "representation point", the depth value d of the point is directly obtained according to its corresponding position on the depth map. The depth value is obtained by decoding the RGB three bands of the depth map according to specific rules.

[0113] ② Calculate the three-dimensional coordinates of the "representation point". Let its pixel coordinates be u and v, and the panorama width and height be W / H. Then, let the distance from the three-dimensional coordinates corresponding to the pixel (u, v) to the projection to the horizontal plane be xy. Then, xy = d × cos ((vH / 2) / H*PI-0.5*PI);

[0114] The three-dimensional coordinates X, Y, and Z corresponding to the pixel (u, v) can be derived as follows:

[0115] X=xy×cos(u / W*PI);

[0116] Y=xy×sin(u / W*PI);

[0117] Z=d×sin((vH / 2) / H*PI-0.5*PI);

[0118] Where PI is the circumference of a circle (in radians).

[0119] ③ According to the absolute position (T) of the panoramic image, the acquired 3D points need to be converted into the absolute coordinate system, namely:

[0120]

[0121] is the world 3D coordinate corresponding to the pixel, is the local (panoramic spherical coordinate system) three-dimensional coordinate corresponding to the pixel, and T is the position of the panoramic sphere in the world coordinate system.

[0122] ④ Fitting the spatial plane where the elements are located. First, determine the geometric center of all points (the average value) and use this as the plane's reference origin. Analyze the spatial distribution of the points and, by calculating the degree of data dispersion in each direction, identify the primary direction of the data extension, thereby determining the plane's normal direction (the direction perpendicular to the plane). Finally, based on the normal direction and the plane's reference origin, calculate the plane's offset and determine the plane's specific position in space. This process is achieved by minimizing the sum of the squared distances from the points to the plane, ultimately yielding a plane equation that best represents the data distribution.

[0123] Furthermore, the present invention aims to generate a "measurable + real scene" map. The "real scene" not only has photo-level clarity, but also needs "element background" to show the environmental conditions of the element. This is different from the traditional "orthophoto". Therefore, it is necessary to take into account the spatial plane formed by the element's own contour line and the effect of the element in the "human eye perspective", such as Figure 3 shown.

[0124] Figure 3 Here, "u" represents the spatial plane where different types of road elements are located (referred to as the "element plane"), p is the vertical plane, which is the "projection plane" (referred to as the "projection plane") to be generated, and the camera position is the camera position of the mobile scanning system, that is, the position of the human eye.

[0125] Regardless of the above type of elements, because the "coplanar operation" has been performed, it can be considered that the elements are in the same spatial surface u. The calculation process from the "element surface" u to the "projection surface" p requires the following information, such as Figure 4 As shown. Among them:

[0126] h is the height of the camera from the ground;

[0127] d is the minimum horizontal distance from the feature surface's bounding rectangle to the camera;

[0128] α is the angle between the element surface and the horizontal plane, and its value range is 0°-90°, that is, from horizontal (such as road surface elements) to vertical (such as signs and other elements);

[0129] f is the distance from the projection surface to the camera, that is, the focal length;

[0130] θ is not shown in the figure. It refers to the roll angle between the projection surface and the feature surface. The feature surface actually exists. The projection surface needs to be defined according to the feature type to achieve a natural visual effect (from the perspective of the camera). The angle θ is then calculated.

[0131] The element face enclosing rectangle refers to the enclosing rectangle of the spatial plane where the element is located when performing coplanar operations. The enclosing rectangle can be "extended with blank space" according to the element type to achieve an effect close to that of a hand-taken photo. The calculation process of generating the projection map p is as follows: Figure 5 As shown, where q is the panoramic sphere.

[0132] Rasterize the feature's bounding rectangle u at a given resolution. Find the intersection of a line from the camera center to the subdivided grid points with the image panorama sphere. Extract the RGB value at that intersection and place it at the intersection of the line and the projection surface p. By traversing each grid point in the feature's bounding rectangle, you can generate a real-world projection of the feature.

[0133] Furthermore, the extracted traffic elements are automatically classified according to the three-dimensional model, and reasonable parameters are set to obtain the classification results of the traffic elements, including:

[0134] determining a classification standard for traffic elements based on the three-dimensional model;

[0135] Training a preset deep learning model according to the classification criteria to obtain a classification model;

[0136] A classification result is determined according to the classification model and parameters of various traffic elements are set according to the classification result, wherein the parameters of various traffic elements include: a classification confidence threshold, a shape similarity score and a geometric feature threshold.

[0137] Specifically, before generating a projection map, the present invention automatically classifies all road elements. The purpose of element classification is to achieve a more natural and aesthetically pleasing angle for the generated projection map. This "natural and aesthetically pleasing" aspect is reflected in the selection of a more reasonable "projection plane." This projection plane must be perpendicular to the horizontal plane, but its rotation angle around the z-axis must be confirmed. For example, the projection plane corresponding to an "oncoming traffic sign" should preferably have an azimuth angle of 0° to 30° with the direction of travel (or road direction), and the azimuth angle of a "side road sign" with the direction of travel should preferably be 60° to 120°.

[0138] Based on the above principles and taking into account the panoramic shooting position during data acquisition, the present invention divides traffic elements into the following types and defines parameters such as angle value and blank space for each type.

[0139] ①On-line surface: surface targets on the road surface that intersect with the trajectory line (closed polygons contain multiple panoramas), such as intersections;

[0140] ②Online strip: Road surface strip targets (closed polygons containing multiple panoramas) that intersect with the trajectory line, such as roadways;

[0141] ③ Parallel strips: Strip-shaped targets along the road that do not intersect the trajectory line, such as metal isolation strips and retaining walls;

[0142] ④ Road surface monomers: monomers on the road surface that need to be marked on the panoramic view at a certain distance (not the closest), such as road surface text, arrows, and manhole covers;

[0143] ⑤ General monomer: monomer that can be observed from the nearest panoramic view, such as trash cans, sewer grates, and safety islands;

[0144] Another advantage of feature classification is that more parameters can be set for the projection plane according to the type, such as angle and distance during panoramic optimization.

[0145] After summarizing, the present invention classifies traffic elements one by one as shown in the following table:

[0146] Table 1 Classification of traffic elements

[0147] id type 100102-Carriageway CXDDP Online ribbon 100103-Non-motorized FJDDP Parallel bands 100104-Ramp ZDDP Online ribbon 100106-Slope BPDP Parallel bands 100107-Open space along the highway YXKDDP Parallel bands 100108-Level Intersection PMJCKDP Line shape 100110-LJDP shoulder Parallel bands 100111-Retaining Wall DQDP_TD Parallel bands 100112-MDDP Parallel bands 100115-Dirt shoulder Parallel bands 100116-Hard shoulder Parallel bands 200101-Retaining Wall DQDP_JC Parallel bands 200102-Guardrail Railing HLDP_LG Parallel bands 200103-Soundproof Screen GYPDP Parallel bands 200104-Anti-glare net FXWDP Parallel bands 200105-Isolation Barrier GLSDP Parallel bands 200106-Greening LHDP_JC Parallel bands 200108-Sidewalk RXDDP Parallel bands 200110-Anti-falling net Parallel bands 600101-Bridge QLDP Online ribbon 600103-Culvert HDDP Parallel bands 600104-Channel TDDP Parallel bands 600105-Tunnel SDDP Parallel bands 600106-Pedestrian Crossing RXHTDDP Online ribbon 600107-Cross-Channel CXHTDDP Online ribbon 600108-Cement Isolation Belt HLDP_SNGLD Parallel bands 600109-Metal Isolation Belt HLDP_JSGLD Parallel bands 600110-Curbstone LYSDP Parallel bands 600111-Drainage facilities_1PSSSDP_1 Parallel bands 600112-Drainage facilities_2PSSSDP_2 Parallel bands 700101-Anti-glare board FXBDP Parallel bands 700102-Anti-collision facilities_Anti-collision pier FZSSDP_FZD Parallel bands 100101-BXDP_LINE Road surface monomer 100109-Marking Arrow BXDP_JT Road surface monomer 300103-Manhole cover JCJDP_JG Road surface monomer 400105-Line marking text BXDP_WENZI Road surface monomer 600113-Speeding pad General monomer 200107-Safety Island AQDDP General monomer 100105-Entrance and Exit CRKDP General monomer 300101-Anti-collision barrel FZSSDP_FZT General monomer 300102-Warning pile JSZDP General monomer 300104-LHDP_TREE General monomer 300111-Car stop stone General monomer 300112-Street Light ZMLDDP General monomer 300113-Signal light General monomer 400102-Corridor LLDP General monomer 400103-Gantry LMJDP Road surface monomer 400104-Lighting Street Light ZMLDDP General monomer 400106-Elevator DTDP General monomer 400107-Drain grate JCJDP_XSBZ General monomer 400108-Pedestrian Bridge QLDP_RXTQ Road surface monomer 500101-km pile GLZDP General monomer 500102-Bus Station GJZDP General monomer 600102-Bridge Expansion Joint QLSSFDP General monomer 800101-Camera SXTDP General monomer 800102-Signal Light XHDDP General monomer 900101-Large sign BZDP_L General monomer 900102-Small sign BZDP_S General monomer 900103-Variable Information Board JBBDP General monomer 900104-Taxi waiting point sign HCDDP General monomer

[0148] After defining the feature type, you need to add the corresponding type during deep learning training to achieve end-to-end direct output of typed features.

[0149] Furthermore, generating an optimal panoramic image according to the classification results of traffic elements includes:

[0150] Based on the classification results of traffic elements, various types of traffic elements are screened within a preset distance range to obtain a filtered panoramic image;

[0151] Evaluate the selected panoramic images, determine their angles and calculation angles with the target traffic element, and ensure that the selected panoramic images can reflect the front view of the traffic element;

[0152] Panoramic images with significant occlusion are eliminated to generate the optimal panoramic image.

[0153] Specifically, based on the feature's projection plane (or normal) and the projection plane's bounding rectangle, the optimal panoramic image of the feature is searched to prepare for generating the final real-life image. The main steps are as follows:

[0154] ① Distance screening. For each traffic feature, calculate the distance from each panoramic image to the feature's midpoint (the midpoint of the enclosing rectangle), and retain those within the threshold range (e.g., 1 to 40 meters).

[0155] ② Angle screening. Calculate the vector from the center of each panorama to the midpoint of the feature, and then calculate the angle between this vector and the normal of the projection surface. Ensure that the two vectors are in opposite directions (i.e., panoramas on the back of the feature cannot be selected). Select panoramas with angles less than a threshold (e.g., 20°) are retained.

[0156] ③ Occlusion determination. Calculate whether there is occlusion in the corresponding depth map within the feature's contour line range, and remove those with severe occlusion. The determination method is: the contour line area ratio exceeds a threshold (such as 30%) and is smaller than the projection surface depth.

[0157] After the above screening, the panorama retained is the optimal panorama. There can be multiple optimal panoramas, but generally two are retained in actual engineering applications (the two with the longest time interval, not two consecutive ones).

[0158] Furthermore, generating a measurable projection map based on the optimal panoramic image includes:

[0159] Using the selected optimal panoramic image, the outline of the target traffic element is projected onto the defined projection surface through the camera pose information;

[0160] Calculate the actual physical position of each pixel in the projection image according to the preset projection parameters;

[0161] A measurable projection map is generated according to the determined actual physical position.

[0162] Specifically, after achieving "element classification" + "optimal panoramic selection", a projection map is generated for each element.

[0163] Furthermore, calculating the three-dimensional coordinates of the target traffic element based on the measurable projection map and the recorded parameters includes:

[0164] Extract the position information of each pixel in the projection image and perform normalized coordinate transformation according to the projection parameters;

[0165] Using the recorded projection surface parameters, determine the actual coordinates of each pixel point in the world coordinate system;

[0166] The three-dimensional coordinates of the target traffic element are determined according to the actual coordinates of each pixel point in the world coordinate system.

[0167] Specifically, the "projection map" generated above is different from the "orthophoto" in photogrammetry. The latter has a fixed GSD (Ground Sampling Distance), that is, a fixed sampling distance. However, the projection map does not have this property. The actual length corresponding to different pixels in the projection map is different. Take the "ground feature" as an example, Figure 6 shown.

[0168] That is, the features represented by the same window have different sizes; therefore, the true 3D coordinates of the feature cannot be directly obtained from the image pixel coordinates (u, v) and the reference GSD, making 3D measurement impossible. However, when the projection image is generated, important parameter information is recorded. Based on these parameters, 3D information can be "recovered." Specifically, given h, d, f, α, and θ, the 3D coordinates (x, y, z) of the spatial plane u corresponding to any pixel point (w, e) on the projection surface p can be calculated. The process can be described as follows: ① Convert the pixel point (w, e) on the projection surface p to normalized coordinates in the camera coordinate system. With the center of the projection surface as the origin, the normalized coordinate values ​​are calculated based on the relationship between the distance from the pixel to the center and the camera focal length. ② Considering the camera height h, the z coordinate of the camera optical center in the world coordinate system is determined. Simultaneously, the perpendicular distance from the camera optical center to the feature surface is calculated based on the angle α between the feature surface u and the horizontal plane, thereby obtaining the z coordinate of the spatial point. ③ By rotating the camera coordinate system by an angle α around the y-axis and an angle θ around the x-axis, the coordinates in the camera coordinate system are converted to the world coordinate system. Finally, adding the translation amount determined by the camera height, the three-dimensional coordinates (x, y, z) corresponding to the pixel point (w, e) on the projection surface p in the spatial plane u can be obtained.

[0169] The following is the detailed derivation process:

[0170] (1) Establishing a camera projection model

[0171] First, based on the camera's pinhole model, in the camera coordinate system, let the camera's optical center be the origin O and the optical axis be the z-axis. The camera height h and focal length f are known.

[0172] For the pixel point (w, e) on the projection surface p, we can first convert it to the normalized coordinates in the camera coordinate system. Assuming that the center of the projection surface p is the coordinate origin (0, 0), the horizontal and vertical distances from the pixel point (w, e) to the center of the projection surface are w and e respectively, then the normalized coordinates (x n ,y n ,z n )for:

[0173] ;

[0174] (2) Consider camera height and angle

[0175] If the camera height h is known, then the z coordinate of the camera's optical center in the world coordinate system is h.

[0176] It is known that the angle between the element surface u and the horizontal plane is α, and the roll angle between the projection surface p and the element surface u is θ.

[0177] First, convert the coordinates in the camera coordinate system to the world coordinate system. According to the principle of similar triangles, let the coordinates of the corresponding point on the spatial plane u be (x, y, z).

[0178] Since (z=hd) (here d is the vertical distance from the camera optical center to the element surface u), according to the trigonometric function relationship, d can be calculated by h and α, (d=h·cos(α)), so (z=hh·cos(α)).

[0179] For x and y coordinates, we need to consider rotation and translation. First, the coordinates in the camera coordinate system (x n ,y n ,z n ) Rotate the camera by an angle α around the y-axis, then rotate it by an angle θ around the x-axis, and finally translate it (the amount of translation is determined by the camera height h).

[0180] The rotation matrix for rotating around the y-axis by an angle α: ;

[0181] The rotation matrix for the angle θ around the x-axis: ;

[0182] Coordinate vector in the camera coordinate system , the rotated coordinate vector Then translate, assuming the translation vector , then the coordinates (x, y, z) on the spatial plane u satisfy: ;

[0183] In summary, through the above steps, the corresponding three-dimensional coordinates (x, y, z) on the spatial plane u can be obtained from the pixel point (w, e) on the projection surface p.

[0184] The above three-dimensional coordinates are obtained from the pixel coordinates (w, e) It is performed in the camera coordinate system, and the obtained coordinates are relative to the camera. If you want to obtain absolute coordinates, you need to multiply the above coordinates by the absolute position T of the panoramic sphere.

[0185] The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the various embodiments can be referenced to each other.

[0186] This document uses specific examples to illustrate the principles and implementation methods of the present invention. The above examples are only intended to help understand the method and core concept of the present invention. At the same time, those skilled in the art will find that the specific implementation methods and application scopes may vary based on the concept of the present invention. In summary, the contents of this specification should not be construed as limiting the present invention.

Claims

1. A method for automatically generating measurable real-scene images of traffic elements, characterized in that: include: Collect high-precision point cloud data and panoramic images and perform preprocessing to obtain orthogonally corrected panoramic images, cleaned high-precision point cloud data and depth maps; Extracting the contours of traffic elements based on the cleaned high-precision point cloud data, the orthogonally corrected panoramic image, and the depth map, and generating a three-dimensional model of the target traffic element based on the contours of the traffic element; Automatically classify the extracted traffic elements according to the three-dimensional model, and set reasonable parameters to obtain classification results of the traffic elements; Generate the optimal panoramic image based on the classification results of traffic elements; generating a measurable projection map based on the optimal panoramic image; Calculating the three-dimensional coordinates of the target traffic element based on the measurable projection map and the recorded parameters; The method of extracting the contour lines of traffic elements based on the cleaned high-precision point cloud data, the orthogonally corrected panoramic image, and the depth map, and generating a three-dimensional model of the target traffic element according to the contour lines of the traffic elements, includes: Construct a traffic element image dataset; Use convolutional neural networks to train and validate the collected panoramic images to extract the contours of traffic elements; Determine the corresponding representation point according to the extracted contour line; Determining a three-dimensional model of a target traffic element according to the representation points; The extracted traffic elements are automatically classified according to the three-dimensional model, and reasonable parameters are set to obtain the classification results of the traffic elements, including: determining a classification standard for traffic elements based on the three-dimensional model; Training a preset deep learning model according to the classification criteria to obtain a classification model; A classification result is determined according to the classification model and parameters of various traffic elements are set according to the classification result, wherein the parameters of various traffic elements include: a classification confidence threshold, a shape similarity score and a geometric feature threshold.

2. The method for automatically generating measurable real-scene images of traffic elements according to claim 1, characterized in that: The high-precision point cloud data and panoramic images are collected using a vehicle-mounted mobile measurement system and a single-soldier mobile measurement system.

3. The method for automatically generating measurable real-scene images of traffic elements according to claim 1, characterized in that: The method for obtaining the panoramic image after orthogonal correction is: Construct an initial spherical coordinate system; Determine the three-dimensional coordinates corresponding to the panoramic image, wherein the three-dimensional coordinates are: ; x original ,y original , z original When the panoramic pixel is converted into a panoramic sphere, the values ​​of the x-axis, y-axis, and z-axis in the three-dimensional coordinates corresponding to the pixel are taken as the radius; Determine a corresponding rotation matrix according to the initial spherical coordinate system and the panoramic image; Determine the corresponding inverse matrix according to the rotation matrix; A new three-dimensional coordinate model corresponding to the panoramic image is determined according to the inverse matrix and the three-dimensional coordinates; wherein the expression of the new three-dimensional coordinate model is: ; in, is the inverse matrix; An orthogonally corrected panoramic image is determined according to the new three-dimensional coordinate model.

4. The method for automatically generating measurable real-scene images of traffic elements according to claim 1, characterized in that: The depth map uses RGB three bands to store depth information.

5. The method for automatically generating measurable real-scene images of traffic elements according to claim 1, characterized in that: The convolutional neural network includes: U-Net and Mask R-CNN.

6. The method for automatically generating measurable real-scene images of traffic elements according to claim 1, characterized in that: Generating an optimal panoramic image based on the classification results of traffic elements includes: Based on the classification results of traffic elements, various types of traffic elements are screened within a preset distance range to obtain a filtered panoramic image; evaluating the filtered panoramic image to obtain an evaluation result; According to the evaluation result, the panoramic images with occlusions are eliminated from the filtered panoramic images to generate an optimal panoramic image.

7. The method for automatically generating measurable real-scene images of traffic elements according to claim 1, characterized in that: Generating a measurable projection map according to the optimal panoramic image includes: Using the selected optimal panoramic image, the outline of the target traffic element is projected onto the preset projection surface through the camera pose information to obtain the initial projection map; Calculating the actual physical position of each pixel in the initial projection image according to preset projection parameters; A measurable projection map is generated based on the actual physical location.

8. The method for automatically generating measurable real-scene images of traffic elements according to claim 1, characterized in that: Calculating the three-dimensional coordinates of the target traffic element based on the measurable projection map and the recorded parameters, including: Extract the position information of each pixel in the projection image, and perform normalized coordinate transformation according to the projection parameters to obtain normalized coordinates; Determine the actual coordinates of each pixel in the world coordinate system using the preset projection surface parameters and normalized coordinates; The three-dimensional coordinates of the target traffic element are determined according to the actual coordinates of each pixel point in the world coordinate system.

Citation Information

Patent Citations

  • Method and system for establishing road asset management system

    CN108388995A

  • Road boundary extraction and vectorization method fusing vehicle-mounted image and point cloud

    CN115690138A