Image-based road facility three-dimensional model automatic matching method
Through image processing and deep learning models, the three-dimensional model of road facilities is automatically matched, and the problem of insufficient description of facilities in high-precision maps is solved, efficient and accurate facility matching is achieved, and the safety and accuracy of intelligent driving and navigation are improved.
Patent Information
- Application Number
- CN202510689106.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-27
- Publication Date
- 2025-07-29
AI Technical Summary
The morphological description of road facilities in the existing high-precision map is insufficient, and manual collection of transportation infrastructure is expensive, making it difficult to achieve efficient and accurate facility matching.
Through image processing and deep learning models, the three-dimensional models of road facilities are automatically matched, including distortion correction, semantic segmentation, feature extraction and similarity calculation, and a high-precision map is generated.
It achieves efficient and accurate matching of road facilities, improving the safety and accuracy of intelligent driving and navigation.
Smart Images

Figure BDA0005421423680000021 
Figure BDA0005421423680000023 
Figure BDA0005421423680000031
Abstract
Description
Technical Field
[0001] The present invention relates to a method for automatically matching three-dimensional models of road facilities, and particularly to a method for automatically matching three-dimensional models of road facilities based on images. Background Art
[0002] The traffic facility information in high-precision maps is closely related to traffic safety and driving order. Existing high-precision maps mainly target the field of autonomous driving services, resulting in insufficient expression of the true form of road facilities and facing two technical dilemmas: First, when OpenDRIVE processes the morphological information of roadside facilities such as street lights, it often adopts a general way and fails to provide an accurate description of the detailed morphological features of these entities; Second, currently, the collection of traffic infrastructure using vehicle-mounted mobile measurement systems still relies on manual methods, and the type of current traffic infrastructure is determined by comparing panoramic images and model library models of facilities, resulting in high labor costs. Summary of the Invention
[0003] Object of the Invention: The object of the present invention is to propose a method for automatically matching three-dimensional models of road facilities based on images, which can perform retrieval and matching through images and a basic three-dimensional model library to expand the semantic information of road facilities in high-precision maps.
[0004] Technical Solution: The present invention includes:
[0005] S1. If the input image is a panoramic image, perform distortion correction on the panoramic image and reduce the image resolution at the same time; otherwise, do not perform any processing and proceed to the next module.
[0006] S2. Use semantic segmentation technology to segment the road facility images in the image, and the segmented facility images are used as the input of the retrieval image.
[0007] S3. Generate a feature view of the road infrastructure model library according to the angle of the road facility, extract the key features of the facility image and the feature map of the model library through a deep learning model, and convert them into numerical information that can be compared by a computer.
[0008] S4. Calculate the similarity according to the retrieval image and the feature view feature vector, sort the retrieval results, and query the corresponding facility model.
[0009] In the above S1, through the corresponding conversion relationship between the panoramic image and the cube view coordinates, generate the front, back, left, right, up, and down six views, and take the front, left, and right three views as the input data for automatic collection of traffic signs.
[0010] The corresponding conversion relationship between the panoramic image and the cube view coordinates is obtained by inputting the corresponding face index k and the cube view coordinates (i, j) and calculating the coordinates (u, v) in the panoramic image. The formula is as follows:
[0011]
[0012] Among them, S C is the side length of the cube face, W and H are the panoramic image resolutions, the cube face index k ∈ {1, …, 6}, representing the front, back, left, right, top, and bottom six faces respectively, and the local coordinates (u, v) ∈ [0, W) × [0, H). is the floor function operator.
[0013] In the above S2, through image segmentation technology, candidate masks are generated based on the image and the predicted bounding boxes, and the outer rectangle of the signboard is generated according to the masks.
[0014] The input image I ∈ R H*W*3 , the rectangular prompt box B = (x min , y min , x max , y max ), after SAM segmentation, the outline of the road facilities is output.
[0015] The above S3 specifically includes:
[0016] S31. Feature view generation: Generate the front, back, left, and right four-direction feature views of the basic model library according to the angle of the road facilities;
[0017] S32. Feature vector extraction: Use the three-dimensional feature matrix view set generated by the road infrastructure model library to train the neural network and extract the feature vectors of the two images.
[0018] The above road facilities generate feature views based on the highway center line as a reference.
[0019] The above S4 specifically includes:
[0020] S41. Similarity calculation: Calculate the similarity of the feature vectors for the extracted feature vectors;
[0021] S42. Recommended facility model output: Return the recommended three-dimensional model according to the similarity or distance.
[0022] The above similarity calculation uses the Euclidean distance to calculate the similarity between the target facility and the feature vectors of the model feature views:
[0023]
[0024] The above S42 is specifically: Return the three three-dimensional models with the smallest distances as the three-dimensional models that may match:
[0025]
[0026] Among them, S* represents the set of the three models closest to M, where M is the retrieved image and S is the three-dimensional model dataset corresponding to the feature view.
[0027] Beneficial effects: The present invention has high efficiency, accuracy, and practicality, and can be used in fields such as intelligent driving, high-precision map updating, and road facility management, helping vehicles or systems better understand the road environment and improving driving safety and navigation accuracy. Brief description of the drawings
[0028] Figure 1 It is a flowchart of the present invention. Detailed implementation manners
[0029] The present invention will be further described below with reference to the accompanying drawings.
[0030] As Figure 1 shown, a method for automatically matching three-dimensional models of road facilities based on images in this embodiment is used to automatically identify and extract semantic information of road facilities (such as traffic signs, street lights, guardrails, etc.) from images and supplement it to the high-precision map. This method first applies the cube projection technology to the panoramic image to reduce the distortion of the panoramic image and at the same time reduce the resolution to reduce the data volume, while directly entering the next module for other types of images; then uses image segmentation technology to accurately segment the road facilities in the planar image and extract the contour and position information of the facilities. At the same time, a basic three-dimensional model library containing various road facilities is established, and each model generates a feature view from a specific angle for subsequent matching. Then, a deep learning model is used to extract feature vectors from the segmented facility image and the feature view in the basic model library, and the matching degree between the two is judged through similarity calculation. Finally, the successfully matched three-dimensional models of road facilities are supplemented to the high-precision map to enrich the details of the map. This method has high efficiency, accuracy, and practicality, and can be used in fields such as intelligent driving, high-precision map updating, and road facility management, helping vehicles or systems better understand the road environment and improving driving safety and navigation accuracy. Specifically, it includes:
[0031] S1. If the input image is a panoramic image, perform distortion correction on the panoramic image through the cube projection technology to generate a cube view, and at the same time reduce the image resolution to reduce the data volume, otherwise do not process and proceed to the next module.
[0032] Input the collected panoramic image RGB with an aspect ratio of 2:1. Through the coordinate correspondence conversion relationship between the panoramic image and the cube view, generate the front, back, left, right, top, and bottom six views, and take the front, left, and right three views as the input data for automatic collection of traffic signs. The parameters are as follows:
[0033] Resolution of the panoramic image: 8192*4096
[0034] Cube view resolution: 2048*2048
[0035] The corresponding conversion relationship between the panoramic image and the cube view coordinates. By inputting the corresponding face index k and the cube view coordinates (i, j), the coordinates (u, v) in the panoramic image are calculated. The formula is as follows:
[0036]
[0037] Where S C is the side length of the cube face, W and H are the resolutions of the panoramic image. The cube face index k ∈ {1, …, 6}, representing the front, back, left, right, top, and bottom six faces respectively. The local coordinates (u, v) ∈ [0, W) × [0, H), is the floor function operator.
[0038] S2. Use semantic segmentation technology to accurately segment the road facility images in the image. The segmented facility images are used as the input of the retrieval image; use image segmentation technology to generate candidate masks according to the image and the predicted bounding box, and generate the outer rectangle of the signboard according to the mask.
[0039] Based on the position relationship between the facilities and the image in the high-precision map, through the frame selection of the data collected by the vehicle-mounted mobile measurement system, and using image segmentation technology (such as SAM) to segment the retrieval image, accurately identify and separate the specific positions and forms of the road facilities.
[0040] The input image I ∈ R H*W*3 , and the rectangular prompt box B = (x min , y min , x max , y max ). After being segmented by SAM, the outline of the road facility is output
[0041] S3. Generate the feature views of the road infrastructure model library according to the specific angles of the road facilities. Extract the key features of the facility image and the feature map of the model library through the deep learning model, and convert them into numerically comparable information by the computer. Specifically include:
[0042] S31. Feature view generation: Generate the front, back, left, and right four-direction feature views of the basic model library according to the specific angles of the road facilities;
[0043] S32. Feature vector extraction: Use the three-dimensional feature matrix view set generated by the road infrastructure model library to train the neural network and extract the feature vectors of the two images.
[0044] The position angles of the road facilities based on the highway center line are summarized in Table 1. According to the distribution characteristics of the road facilities, generate the front, back, left, and right four-direction feature views of the basic model library.
[0045] Table 1 Angle of the Location of Road Facilities
[0046]
[0047] A convolutional neural network (such as the VGG-NET16 network) trained with a three-dimensional feature matrix view set generated from a road infrastructure model library is used to extract the feature vectors of two images. The extracted feature vectors are represented by f(x) and f(y) respectively.
[0048] S4. Calculate the similarity based on the retrieved image and the feature view feature vectors, and sort the retrieval results to query the corresponding facility model, specifically including:
[0049] S41. Similarity calculation: Calculate the similarity of the feature vectors for the extracted feature vectors;
[0050] The Euclidean distance is used to calculate the similarity between the target facility and the model feature view feature vectors. The formula (5) is as follows:
[0051]
[0052] The above calculation result is the similarity of two vectors, while the retrieval input is the retrieved image y after SAM segmentation, and the returned result is a three-dimensional model. The distance between the retrieved image and the three-dimensional models in the dataset can be calculated through the distance of the feature vectors.
[0053] S42. Recommended facility model output: Return the recommended three-dimensional model for people to select according to the similarity or distance.
[0054] By returning the three three-dimensional models with the smallest distances as the three-dimensional models that may match, the formula (6) is as follows.
[0055]
[0056] Where S* represents the set of the three models closest to M, M is the retrieved image, and S is the three-dimensional model dataset corresponding to the feature view.
Claims
1. An automatic matching method for 3D models of road facilities based on images, characterized in that, It includes the following steps: S1. If the input image is a panoramic image, perform distortion correction on the panoramic image and reduce the image resolution at the same time; otherwise, do not process it and proceed to the next module; S2. Use semantic segmentation technology to segment the road facility images in the image, and the segmented facility images are used as the input of the retrieval image; S3. Generate the feature views of the road infrastructure model library according to the angles of the road facilities, extract the key features of the facility images and the feature maps of the model library through a deep learning model, and convert them into numerical information that can be compared by a computer; S4. Calculate the similarity according to the retrieval image and the feature view feature vector, sort the retrieval results, and query the corresponding facility model.
2. The method for automatically matching a 3D model of road facilities based on images according to claim 1, wherein In S1, through the corresponding conversion relationship between the panoramic image and the cube view coordinates, generate the front, back, left, right, top, and bottom six views, and take the front, left, and right three views as the input data for automatic collection of traffic signs.
3. The automatic matching method for three-dimensional models of road facilities based on images according to claim 2, characterized in that For the corresponding conversion relationship between the panoramic image and the cube view coordinates, by inputting the corresponding face index k and the cube view coordinates (i, j), the coordinates (u, v) in the panoramic image are calculated, and the formula is as follows: Among them, S C is the side length of the cube face, W and H are the panoramic image resolutions, the cube face index k ∈ {1, …, 6}, representing the front, back, left, right, top, and bottom six faces respectively, and the local coordinates (u, v) ∈ [0, W) × [0, H). is the floor operator.
4. The method for automatically matching a 3D model of road facilities based on images according to claim 1, wherein In S2, through image segmentation technology, generate candidate masks according to the image and the predicted bounding box, and generate the outer rectangle of the signboard according to the masks.
5. The method for automatically matching a 3D model of road facilities based on images according to claim 4, wherein Input image \(I\in\mathbb{R}\) H*W*3 , rectangular prompt box \(B=(x\) min ,y min ,x max ,y max ), after SAM segmentation, output the outline of road facilities 6. The method for automatically matching a 3D model of road facilities based on an image according to claim 1, wherein S3 specifically includes: S31. Feature view generation: Generate the front, back, left, and right four-direction feature views of the basic model library according to the angles of the road facilities; S32. Feature vector extraction: Use the three-dimensional feature matrix view set generated by the road infrastructure model library to train the neural network and extract the feature vectors of the two images.
7. The method for automatically matching a 3D model of road facilities based on an image according to claim 6, characterized in that, The road facilities generate feature views based on the highway center line as the reference.
8. The method for automatically matching a 3D model of road facilities based on images according to claim 1, wherein S4 specifically includes: S41. Similarity calculation: Calculate the similarity of the feature vectors for the extracted feature vectors; S42. Recommended facility model output: Return the recommended three-dimensional model according to the similarity or distance.
9. The method for automatically matching a 3D model of road facilities based on images according to claim 8, wherein The similarity calculation uses the Euclidean distance to calculate the similarity between the target facility and the feature vector of the model feature view:
10. The method for automatically matching a 3D model of road facilities based on an image according to claim 8, characterized in that S42 is specifically: Return the three three-dimensional models with the smallest distance as the three-dimensional models that may match: Among them, S* represents the set of the three models closest to M, M is the retrieval image, and S is the three-dimensional model data set corresponding to the feature view.