An intersection reconstruction method based on high-altitude video semantic element extraction and medium

By setting up targets at intersections and using high-altitude video and RTK equipment to obtain UTM coordinates, combined with perspective transformation and artificial intelligence algorithms to identify intersection elements, the problem of difficulty in obtaining intersection plan maps in existing technologies has been solved, and efficient and accurate measurement and model reconstruction of intersection elements have been achieved.

CN115965710BActive Publication Date: 2026-04-28WUHAN UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
WUHAN UNIV OF TECH
Filing Date
2023-01-16
Publication Date
2026-04-28

AI Technical Summary

Technical Problem

Existing technologies make it difficult to quickly and accurately obtain intersection plans, resulting in a lack of effective data support for intersection evaluation and improvement.

Method used

By setting up targets at intersections, using high-altitude video and RTK equipment to obtain UTM coordinates, and combining perspective transformation matrices and artificial intelligence algorithms to identify intersection elements, extract their feature points, and establish logical relationships, intersection reconstruction is achieved.

Benefits of technology

It achieves efficient automatic identification and accurate measurement of intersection elements, generating more accurate intersection models to support subsequent evaluation and optimization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115965710B_ABST
    Figure CN115965710B_ABST
Patent Text Reader

Abstract

The application discloses a kind of intersection reconstruction method and medium based on high-altitude video semantic element extraction, which comprises the following steps: at least four targets are arranged in the intersection to be reconstructed, and the UTM coordinates of the target are obtained;High-altitude video or image of the intersection is obtained, and the perspective transformation matrix is calculated according to the image pixel coordinates of the target and the corresponding UTM coordinates;The perspective transformation matrix is used to perform perspective transformation on the image of the intersection, and the orthographic projection image of the intersection is obtained;The intersection elements are extracted by identifying the orthographic projection image of the intersection through artificial intelligence algorithm, and the characteristic points of the intersection elements are obtained;According to the semantic of intersection element and the logical relationship of intersection element, the intersection element is identified according to the basic element rule of intersection and the traffic flow rule;The intersection is reconstructed in combination with the logical relationship of intersection element.The application can extract the key elements of intersection through artificial intelligence algorithm by the picture or video of intersection shot from high altitude, and realize the reconstruction of intersection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of intersection re-representation and visualization, specifically involving an intersection reconstruction method and medium based on the extraction of semantic elements from high-altitude video. Background Technology

[0002] As key nodes in road traffic, the safety and efficiency of intersections during their operation are a major concern for traffic management departments. Creating plan views of operational intersections allows for rapid intersection modeling, facilitating the analysis of safety or efficiency issues and enabling the development of targeted improvement measures. However, many existing intersections lack accurate plan views due to their long design time and multiple reconstructions and expansions. Therefore, effective and usable data for intersection evaluation and improvement lacks suitable modeling. Consequently, how to quickly and accurately obtain complete plan views of intersections is a pressing issue that needs to be addressed. Summary of the Invention

[0003] The purpose of this invention is to provide a method and medium for intersection reconstruction based on semantic element extraction from high-altitude video, thereby reconstructing the intersection by extracting intersection elements from high-altitude video.

[0004] The technical solution of the present invention is as follows:

[0005] A method for reconstructing intersections based on semantic feature extraction from high-altitude video includes the following steps:

[0006] At least four targets are deployed at the intersection to be reconstructed, and the UTM coordinates of the targets are obtained; high-altitude video or images of the intersection are obtained, and the perspective transformation matrix is ​​calculated based on the image pixel coordinates and corresponding UTM coordinates of the targets;

[0007] The perspective transformation matrix is ​​used to perform perspective transformation on the image of the intersection to obtain the orthographic projection image of the intersection;

[0008] The intersection features are extracted by recognizing the orthographic projection image of the intersection using artificial intelligence algorithms, and the feature points of the intersection features are obtained; the UTM coordinates of the feature points of the intersection features are obtained based on their pixel positions.

[0009] Based on the basic intersection element rules and traffic flow rules, the semantics of intersection elements are identified and the logical relationships between intersection elements are established. Identifying the semantics of intersection elements means determining the semantic elements of intersections to which each feature point belongs, and constructing intersection elements through the feature points of the semantic elements. Establishing the logical relationships between intersection elements means determining the parameters of each intersection element and the matching relationships between each intersection element.

[0010] Reconstruct the intersection by combining the logical relationships between the intersection elements.

[0011] Furthermore, the targets are rectangular or circular grids with alternating black and white lines; each target is dispersed and its UTM coordinates are obtained by an RTK device located at the center of the target.

[0012] Furthermore, the perspective transformation matrix is ​​as follows:

[0013]

[0014] In the formula, a 11 a 12 a 13 a 21 a 22 a 23 a 31 a 32 For coefficients;

[0015] Perspective transformation is as follows:

[0016]

[0017] In the formula, x and y are the horizontal and vertical coordinates of the original image, respectively. ′ ,y ′ The x and y coordinates are the transformed image.

[0018] Furthermore, intersection elements include lane lines, stop lines, pedestrian crossings, traffic islands, guardrails, and lane arrows.

[0019] Furthermore, the endpoints and key transformation points of the line segments of the intersection elements are extracted, and straight lines or circular curves are fitted to them to obtain the feature points of the intersection elements; the feature points of the intersection elements include the endpoints of the line segments, the radius of the circular curve, the center of the circle, and the starting angle.

[0020] Furthermore, the semantic elements of an intersection include lane lines, stop lines, pedestrian crossings, traffic islands, the edges of the median strip and other channelization facilities, lane function arrows, and curbs;

[0021] The semantic elements of the intersection to which each feature point belongs are modified and reviewed through human interaction.

[0022] Furthermore, determining the parameters of each intersection element involves extracting the location, length, width, and other quantitative values ​​of the intersection semantic elements, including: the angle of each lane, lane length and width, stop line position, length and width of the intersection median strip, start and end positions and width of the central median strip, length and width of the pedestrian crossing area, and turning radius of the curb strip.

[0023] Furthermore, artificial intelligence algorithms include image recognition algorithms based on deep learning.

[0024] Furthermore, intersection reconstruction includes:

[0025] Create a new intersection object with the following attributes: function of lane x; angle, length, and width of lane x; stop line location; length and width of the intersection's motor vehicle / non-motor vehicle separation zone; length and width of the central divider zone; and length and width of the pedestrian crossing zone.

[0026] Perform measurement procedures to determine the geometric dimensions of intersection elements, including: the angle of each lane; lane length and width; stop line location; length and width of the intersection median strip area; length and width of the central median strip area; length and width of the pedestrian crossing area; and improve the intersection element database.

[0027] Based on a comprehensive database of intersection elements, intersections are generated on a two-dimensional coordinate system, or lanes, stop lines, intersection median strips, central median strips, and pedestrian crossings are manually set on a two-dimensional coordinate system to obtain a reconstructed intersection.

[0028] A computer-readable storage medium, characterized in that the computer-readable storage medium stores program code, wherein, when the program code is executed, the intersection reconstruction method based on high-altitude video semantic element extraction described above is executed.

[0029] Compared with the prior art, the present invention has the following advantages and beneficial effects:

[0030] This invention can extract key elements of intersections through aerial photography or video and artificial intelligence algorithms, achieving efficient and automatic recognition of intersection elements and their semantics. Based on the automatically extracted element measurement locations, a more accurate intersection model can be established, and the characteristics of the intersection can be accurately analyzed through the model, which helps in subsequent evaluation and optimization of the intersection. Attached Figure Description

[0031] Figure 1 This is a flowchart of the intersection reconstruction method based on high-altitude video semantic element extraction according to the present invention;

[0032] Figure 2 This is a schematic diagram illustrating the semantic recognition of intersection elements according to the present invention. Detailed Implementation

[0033] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention. Furthermore, the technical features involved in the various embodiments of this invention described below can be combined with each other as long as they do not conflict with each other.

[0034] Existing methods for drawing intersection plans involve two main approaches: one is manual measurement using a theodolite on the road, which is inefficient; the other is using aerial video, but the images are uncorrected and distorted, lacking measurability. Addressing the issues of low efficiency in traditional data acquisition and the inability of road simulation software to automate and accurately generate intersections, this invention provides an intersection reconstruction method and medium based on semantic element extraction from aerial video. This method can automatically, efficiently, and accurately acquire key parameters of a planar intersection and reconstruct the intersection accordingly, achieving automated acquisition of road parameter data and providing a solution for rapid intersection modeling and digital representation. Furthermore, existing intersection simulation platforms primarily use two-dimensional planar maps supplemented with basic road parameters to display intersection models. However, for operating intersections, there is no quick and simple method to measure their accurate data. This invention automatically extracts intersection elements from aerial video, achieves accurate measurement of intersection elements through perspective transformation, and ultimately extracts important intersection parameters for reconstruction.

[0035] The intersection reconstruction method based on high-altitude video semantic element extraction of the present invention includes the following steps:

[0036] Step 1: Set up targets and take aerial photos. Place at least four targets at the intersection to be reconstructed, ensuring the targets are as dispersed as possible and within the shooting range. Use RTK equipment to obtain the latitude, longitude, and UTM coordinates of the targets. Simultaneously, use aerial camera equipment to obtain panoramic photos or videos of the intersection to confirm that all targets can be photographed.

[0037] While acquiring high-altitude video, at least four targets are deployed on the driving plane at the target intersection. The targets are rectangular or circular grids with alternating black and white lines. The latitude and longitude (l, b) of the targets are determined using RTK equipment, and the UTM coordinates (x, y) are obtained. The actual planar physical coordinates of the targets are recorded.

[0038] Step 2, Image Transformation. Acquire an image. Based on the relationship between the UTM plane coordinates (x, y) of each target in the image and the image pixel coordinates (u, v), calculate a unique perspective transformation matrix between the target pixel coordinates and the UTM plane coordinates. Using this perspective transformation matrix, perform a perspective transformation on the image to generate an orthographic projection image.

[0039] Based on the relationship between the pixel coordinates of the target and the actual UTM coordinates, a perspective transformation is performed on the image to obtain the orthographic projection image of the target intersection. At this time, the pixel position in the orthographic projection image represents a fixed physical position, which can be used for subsequent measurements.

[0040] Step 3: Using artificial intelligence algorithms, identify intersection elements such as lane lines, stop lines, pedestrian crossings, and lane arrows from the transformed image to obtain the intersection elements, their feature point pixel positions, and actual UTM coordinate values.

[0041] Artificial intelligence algorithms are used to identify typical elements required for intersection reconstruction from images, including but not limited to lane lines, stop lines, pedestrian crossings, traffic islands, guardrails, lane arrows, etc.; and straight lines and circular curves are fitted to the endpoints and key transformation points of the key elements of the intersection to obtain the pixel positions of the feature points of the intersection elements and then convert them into UTM coordinates.

[0042] Step four involves identifying the semantics of intersection elements and establishing logical relationships based on the basic intersection element rules and traffic flow rules, and further obtaining parameters of intersection elements, such as lane width and stop line position. This step achieves the modification and confirmation of semantic elements through manual intervention.

[0043] By analyzing intersection rules and typical intersection elements and their locations, algorithms are used to identify the semantic elements of the intersection represented by these elements. These elements include, but are not limited to: lane lines, stop lines, pedestrian crossings, traffic islands, edges of median strips, edges of channelization facilities such as central dividers, lane function arrows, and curbs. After automatic semantic identification, the semantic elements can be modified and reviewed through human interaction. This step transforms the aforementioned scattered elements into elements that reconstruct the intersection.

[0044] The step of extracting the spatial location of the semantic elements of the intersection from the transformed video above extracts the location, length, width and other quantitative values ​​required to represent the intersection, including but not limited to: the angle of each lane; the length and width of the lanes; the location of the stop line; the length and width of the intersection's motor vehicle and non-motor vehicle separation zone; the start and end positions and width of the central divider; the length and width of the pedestrian crossing area; the turning radius of the curb, etc.

[0045] Step 5: Combining the logical relationships of intersection elements, the intersection is numerically represented in a certain structure and stored in a database to complete the intersection reconstruction.

[0046] The present invention also provides a computer-readable storage medium storing program code, wherein, when the program code is executed, the intersection reconstruction method based on high-altitude video semantic element extraction described above is performed.

[0047] In summary, the intersection reconstruction method and medium provided by this invention utilize video data from UAV aerial photography or high-altitude cameras, while simultaneously deploying ground RTK targets, to automatically generate accurate road parameters without requiring manual on-site road surveys. Furthermore, it employs various artificial intelligence algorithms to correct video data, eliminating video data distortion and ensuring that accurate and reliable intersection planar parameters can be generated from video data.

[0048] Example:

[0049] The embodiments of the present invention provide a method and medium for intersection reconstruction based on high-altitude video semantic element extraction, thereby achieving at least a certain degree of automation of intersection reconstruction and performing intersection reconstruction efficiently and accurately.

[0050] like Figure 1 As shown, the specific process of the intersection reconstruction method based on high-altitude video semantic element extraction is as follows:

[0051] Step S101: Acquire aerial images or videos of the intersection, obtain the pixel values ​​(u, v) of at least four target positions in the image, and simultaneously obtain the sub-meter UTM plane coordinates (x, y) corresponding to the four targets in the image to obtain the perspective transformation matrix:

[0052]

[0053] Specifically, the high-altitude video data mainly comes from high-altitude cameras or drone aerial photography; the target mainly refers to a checkerboard pattern on the ground, with an RTK device placed at the center. It should be noted that the RTK device here has sub-meter level positioning accuracy.

[0054] Step S102: Based on the obtained perspective transformation matrix, perform image perspective transformation to obtain the target orthographic projection image, and perform image masking to obtain a cross-shaped image (e.g., extracting the region of interest). The calculation method for the two-dimensional perspective transformation matrix is ​​as follows:

[0055]

[0056] Where x and y are the x and y coordinates of the original image, x ′ ,y ′ Let a be the x and y coordinates of the transformed image. 11 a 12 a 13 a 21 a 22 a 23 a 31 a 32 The eight coefficients can be obtained by solving eight equations.

[0057] Step S103: Use artificial intelligence algorithms to identify typical elements required for intersection reconstruction from the image, including but not limited to lane lines, stop lines, pedestrian crossings, traffic islands, guardrails, lane arrows, etc.; obtain key points and fit straight lines and circular curves; obtain the pixel positions of the feature points of the intersection elements and then convert them into UTM coordinates.

[0058] Specifically, artificial intelligence algorithms can be based on image recognition algorithms including but not limited to deep learning, and the UTM coordinates of feature points are obtained based on the coordinate transformation formula of the RTK device used.

[0059] For example, this step can be divided into two steps. First, the endpoint UTM values ​​(x, y) of line segments of key intersection elements (intersection median strip edges, central divider and other channelization facility edges, lane lines, stop lines, lane function arrows, curb strips) and their key transformation points are marked on the image using artificial intelligence or manual labeling methods. Then, a straight line or circular curve fitting program is performed on the image to obtain the feature points of the aforementioned key intersection elements, such as: line segment endpoints, circular curve radius / center / starting angle.

[0060] Step S104: Identify the semantics of intersection elements based on the basic intersection element rules and traffic flow rules; this step can be achieved through manual intervention to modify and confirm the semantic elements.

[0061] The following is a detailed explanation of the method for identifying the semantics of intersection elements using a specific example. Figure 2 This is a schematic diagram of the method for recognizing the semantics of intersection elements provided in an embodiment of the present invention. Please refer to... Figure 2 The example includes typical elements such as approach lane lines, stop lines, sidewalks, curbs, medians, and targets. The feature points of approach lane lines are their endpoints (two or more); the feature points of stop lines are their two endpoints; the feature points of sidewalks are their four endpoints or more (which can be used to create a polygon); the feature points of curbs are the points on the identified circular edge of the curb (multiple points can be used to create a precise circle); and the feature points of the median are... Figure 2 The shadowed polygon has multiple corner points; the target's feature point is the center of the checkerboard rectangle. Solid lines in the diagram represent the lane lines at the entrance, and dashed lines represent the lane lines at the exit.

[0062] For example, such as Figure 2 As shown, after processing the image to obtain the key element feature points of the intersection, keywords are labeled for the feature points based on the basic element rules of the intersection (or the intersection design specifications) and traffic flow rules. Labeling the endpoints of the approach lane lines can represent the relevant symbols for "approach lane lines"; labeling the two endpoints of the stop lines can represent the relevant symbols for "stop lines"; labeling the four endpoints or multiple points of the pedestrian crossing can represent the relevant symbols for "pedestrian crossings"; labeling the points on the circular edge of the curb can represent the relevant symbols for "curb strips"; labeling the multiple corner points of the median polygon can represent the relevant symbols for "median strips"; and labeling the center of the checkerboard rectangle can represent the relevant symbols for "targets".

[0063] Step S105: Establish the logical relationship between the semantic elements of the intersection based on the basic element rules and traffic flow rules of the intersection.

[0064] Specifically, the process begins by manually specifying which lane, median, and central divider the key elements of the intersection belong to. Then, the approximate function of each lane is marked based on the lane markings. The matching of elements on the lanes is then refined to ultimately obtain the logical relationships between the intersection elements.

[0065] Step S106: Combine the logical relationships between intersection elements and store the intersection in a database with a certain structure to complete the intersection reconstruction.

[0066] Specifically, the first step is to create a new intersection object in the database language. The new object attributes (i.e., the key elements of the intersection) include: the function of lane x; the angle, length, and width of lane x; the location of the stop line; the length and width of the intersection's median strip area; the length and width of the central divider area; and the length and width of the pedestrian crossing area. The second step is to execute a measurement program to determine the geometric dimensions of the key elements of the intersection composed of the above elements, including: the angle of each lane; the length and width of each lane; the location of the stop line; the length and width of the intersection's median strip area; the length and width of the central divider area; and the length and width of the pedestrian crossing area. This completes the database of key intersection elements, updating the measurement parameter values. The third step is to generate the intersection on a two-dimensional coordinate system based on the completed database of key intersection elements obtained in the second step, or to manually set the lanes, stop lines, intersection median strip area, central divider area, and pedestrian crossing area on a two-dimensional coordinate system to obtain the reconstructed intersection.

[0067] Based on the same inventive concept, embodiments of the present invention also provide a readable computer database storage medium storing a computer program thereon, which is executed by a processor as described in any of the above embodiments of the intersection reconstruction method, and establishes a reconstructed intersection database.

[0068] In summary, this invention provides a method and medium for intersection reconstruction based on semantic element extraction from high-altitude video. The method includes acquiring panoramic photos or videos taken from a high altitude at a selected intersection, simultaneously deploying at least four targets on the driving plane, and using RTK to determine the latitude and longitude of the targets to obtain UTM coordinates. Perspective transformation is then performed on the image based on the relationship between the target feature point pixels and the UTM coordinates to generate an orthographic projection image. Artificial intelligence algorithms are used to extract intersection elements such as lane lines, stop lines, pedestrian crossings, and lane arrows, obtaining their actual coordinates through the positional feature point pixels. The semantics of the intersection elements are identified and logical relationships are established based on basic intersection element rules and traffic flow rules. Combining these logical relationships, the intersection is numerically represented in a specific structure and stored in a database to complete the intersection reconstruction. This invention can automatically extract measurement data representing intersection elements from high-altitude video of the intersection for semantic recognition, and digitally represent, store, and reconstruct the intersection through video surveys.

[0069] It should be noted that, depending on the implementation needs, the various steps / components described in this application can be broken down into more steps / components, or two or more steps / components or parts of the operation of steps / components can be combined into new steps / components to achieve the purpose of this invention.

[0070] Those skilled in the art will readily understand that the above are merely preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A method for reconstructing intersections based on semantic element extraction from high-altitude video, characterized in that, Includes the following steps: At least four targets are deployed at the intersection to be reconstructed, and the UTM coordinates of the targets are obtained; high-altitude video or images of the intersection are obtained, and the perspective transformation matrix is ​​calculated based on the image pixel coordinates and corresponding UTM coordinates of the targets; The perspective transformation matrix is ​​used to perform perspective transformation on the image of the intersection to obtain the orthographic projection image of the intersection; The intersection features are extracted by identifying the orthographic projection image of the intersection using artificial intelligence algorithms. The endpoints of the line segments of the intersection features and their key transformation points are extracted, and straight lines or circular curves are fitted to obtain the feature points of the intersection features. The UTM coordinates of the feature points of the intersection features are obtained based on their pixel positions. The feature points of the intersection features include the endpoints of the line segments, the radius of the circular curve, the center of the circle, and the starting angle. Based on the basic rules of intersection elements and traffic flow rules, the semantics of intersection elements are identified and the logical relationships between them are established. Identifying the semantics of intersection elements involves determining the semantic elements to which each feature point belongs, and constructing intersection elements from the feature points of these semantic elements. Establishing the logical relationships between intersection elements involves determining the parameters of each intersection element and the matching relationships between them. Determining the parameters of each intersection element involves extracting the position, length, width, and other quantitative values ​​of the semantic elements, including: the angle of each lane, lane length and width, stop line position, length and width of the intersection median strip, start and end positions and width of the central median strip, length and width of the pedestrian crossing area, and the turning radius of the curb. Reconstructing a planar intersection by combining the logical relationships between intersection elements includes: Create a new intersection object with the following attributes: function of lane x; angle, length, and width of lane x; stop line location; length and width of the intersection's motor vehicle / non-motor vehicle separation zone; length and width of the central divider zone; and length and width of the pedestrian crossing zone. Perform measurement procedures to determine the geometric dimensions of intersection elements, including: the angle of each lane; lane length and width; stop line location; length and width of the intersection median strip area; length and width of the central median strip area; length and width of the pedestrian crossing area; and improve the intersection element database. Based on a comprehensive database of intersection elements, intersections are generated on a two-dimensional coordinate system, or lanes, stop lines, intersection median strips, central median strips, and pedestrian crossings are manually set on a two-dimensional coordinate system to obtain a reconstructed planar intersection.

2. The intersection reconstruction method based on high-altitude video semantic element extraction according to claim 1, characterized in that, The targets are rectangular or circular grids with alternating black and white lines; each target is dispersed and its UTM coordinates are obtained by an RTK device located at the center of the target.

3. The intersection reconstruction method based on high-altitude video semantic element extraction according to claim 1, characterized in that, The perspective transformation matrix is ​​as follows: R= In the formula, , , , , , , , For coefficients; Perspective transformation is as follows: In the formula, The x and y coordinates of the original image. The x and y coordinates are the transformed image.

4. The intersection reconstruction method based on high-altitude video semantic element extraction according to claim 1, characterized in that, Intersection elements include lane lines, stop lines, pedestrian crossings, traffic islands, guardrails, and lane arrows.

5. The intersection reconstruction method based on high-altitude video semantic element extraction according to claim 1, characterized in that, The semantic elements of an intersection include lane lines, stop lines, pedestrian crossings, traffic islands, the edges of the separation strip between motorized and non-motorized lanes, the central divider, lane function arrows, and curb strips; The semantic elements of the intersection to which each feature point belongs are modified and reviewed through human interaction.

6. The intersection reconstruction method based on high-altitude video semantic element extraction according to claim 1, characterized in that, Artificial intelligence algorithms include image recognition algorithms based on deep learning.

7. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores program code, wherein, when the program code is executed, the intersection reconstruction method based on high-altitude video semantic element extraction as described in any one of claims 1 to 6 is performed.

Citation Information

Patent Citations

  • High-precision traffic element target extraction method based on image point cloud fusion

    CN112434706A

  • High-precision map construction method and device, electronic equipment and storage medium

    CN113034566A