RTK and image reconstruction combined road scene space rapid investigation system
By combining the system of RTK and image reconstruction, the road topology skeleton is built, the mapping unit is divided, the structural semantic areas are extracted and the tile boundary pose registration is carried out, and the spatial drift, misalignment and structural incompleteness of road scene mapping in the existing technology is solved, and efficient road space scene mapping and stitching are achieved.
Patent Information
- Application Number
- CN202510577610.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-07
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2045-05-07
AI Technical Summary
The existing technology has problems such as spatial drift, misalignment of splicing, lack of structural element identification and modeling optimization in the spatial mapping and survey of road scenarios, resulting in fuzzy road geometric expression and incomplete semantic structure.
The road scene space rapid survey system combined with RTK and image reconstruction is realized through the road trajectory acquisition module, the road map analysis module, the structure enhancement module and the road survey module to realize the construction of road topological skeleton, the division of map construction units, the extraction of structural semantic regions and the local attitude registration of block boundaries.
The mapping efficiency in road space scenarios is improved, the continuous splicing and structural consistency of road scenarios is achieved, and the clarity of road geometric expression and the integrity of semantic structure are enhanced.
Smart Images

Figure CN120107504A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of road scene space survey, and more specifically, to a road scene space rapid survey system combining RTK and image reconstruction. Background Art
[0002] With the rapid development of smart transportation, autonomous driving, high-precision map construction, and urban digital management, the acquisition and expression of road scene spatial information has become a basic and key link. As the most important structural network in urban space, the spatial topological relationship, surface morphology, element distribution and other information of roads directly affect the overall performance of path planning, traffic organization and urban perception systems.
[0003] Deficiencies of existing technologies: Spatial mapping and surveying of road scenes mainly rely on a single sensor (such as GNSS or visual SLAM) for path positioning and 3D reconstruction. Positioning methods based on images or SLAM are severely affected by changes in lighting, texture and perspective, and lack global consistency, resulting in spatial drift and splicing misalignment in the reconstruction results. Existing reconstruction methods mostly use uniform sampling and blind point cloud generation strategies, and fail to identify and optimize structural elements such as lane lines and road edges, resulting in fuzzy road geometry expression and incomplete semantic structure. The overall mapping process often lacks a clear spatial organizational structure, lacks control point drive and main axis constraint mechanisms, and is difficult to achieve modular, efficient tile division and parallel mapping, resulting in low spatial splicing efficiency in road scenes. Summary of the invention
[0004] In order to overcome the above defects of the prior art, there is a solution as follows to solve the problem of poor road mapping splicing effect in the above background technology.
[0005] To achieve the above object, the present invention provides the following technical solutions: A road scene space rapid survey system combining RTK and image reconstruction, including a road trajectory acquisition module, a road mapping analysis module, a structure enhancement module and a road survey module, and each module is connected by a signal; The road trajectory acquisition module is used to collect road inspection trajectories and synchronize image frames with trajectory points, remove abnormal jump points, extract continuous control points, and establish a road topology skeleton based on control points and direction vectors; The road mapping analysis module is used to divide the road space into continuous mapping units based on control points and direction information, set the spatial range and main axis direction and regional attribution mapping for the mapping units, and construct independent mapping task packages to map overlapping areas; The structure enhancement module is used to extract the structural semantic area and edge information in the image, construct a structure mask and map it to the three-dimensional space, and mark it as a structure enhancement point; The road survey module is used to identify the overlapping areas of tile boundary structures and establish anchor points, complete local rigid body registration, and then build a global error minimization model to uniformly optimize the posture and position of all tiles and conduct overall road space survey.
[0006] In a preferred embodiment, for collecting multi-source data in the network and performing graph structured modeling, the specific steps are as follows: for collecting road inspection trajectories and synchronizing image frames and trajectory points, the specific steps are as follows: The three-dimensional spatial positioning data collected in the road scene is used as RTK trajectory data. The RTK trajectory data corresponds to the real-time spatial position of each frame of image collection in the road scene. The collection object is the installation position of the image acquisition terminal, including vehicle-mounted cameras, mobile work platforms, handheld devices, and trajectory scanning devices; the collection path is all the spatial points passed along the road topology path during the road inspection process; The coverage area includes motor vehicle lanes, sidewalks, bicycle lanes, curbs, lane line intersections; turning points, intersections, ramps, intersections, under viaducts, culverts, and overpasses; Perform RTK trajectory point collection as a trajectory point sequence. The RTK trajectory point includes the three-dimensional positioning coordinates corresponding to the collection position and the corresponding timestamp; The collected image frame data includes image frames and image frame collection time; For each image frame, find the RTK track point closest to the timestamp and perform time binding between the image and the track point; Set the time matching tolerance threshold. If the absolute value of the difference between the image frame acquisition time and the timestamp of the closest RTK track point is less than or equal to the matching tolerance threshold, the frame image is bound to an RTK track point and the pairing data is retained.
[0007] In a preferred embodiment, abnormal jump points are removed, and continuous control points are extracted, and a road topology skeleton is established based on the control points and direction vectors. The specific steps include: Perform RTK trajectory smoothing and pseudo-jump point removal, and convert three consecutive trajectory points into , the direction vector is calculated as ,in, It is from the point Pointing Point The direction vector of It is from the point Pointing Point The direction vector of Structural direction change angle: ,in, Represents the Euclidean norm of a vector; represents the turning angle of the trajectory at the i-th point; Calculate the spatial distance between adjacent trajectory points : ,like , there is a drastic space jump, is the distance threshold; The pseudo jump point determination rules are set as follows: If the trajectory point satisfies both: and , it is determined to be an abnormal jump point and recorded as a pseudo jump point; is the set turning angle threshold; If the pseudo jump point is only isolated, it is directly removed. If the trajectory points before and after the jump point are continuous, linear interpolation or spline reconstruction is used to fill the gap; Summarize the interpolation points to form the final set of control points: ,in, is a three-dimensional Euclidean real number space, is the kth control point; n is the total number of extracted control points; Each control point from the trajectory point set or interpolation point is used as a space partition anchor point and skeleton graph node.
[0008] In a preferred embodiment, each control point from the trajectory point set or interpolation point is used as a space partition anchor point and a skeleton graph node, and the specific steps are as follows: Construct a direction vector for any pair of consecutive k-th and k+1-th control points , C is the control point set, and the unit direction vector is defined as: ,in, , represents the main axis direction of the k-th space segment; Construct a set of attitude vectors and a set of direction vectors: , each direction vector With corresponding control points Bind to form a posture reference pair: , the pose information is used as the orientation reference for the partition, and as the main vector for initializing the camera pose estimation for image reconstruction, and as the pose constraint input during the tile stitching process; Construct the road skeleton graph structure and define the road skeleton graph as: : ; A set of directed edges: ; Each edge records the following attribute tuple: ,in, is the side length, which indicates the spatial distance between control point pairs; is the direction angle, which indicates the turning angle between road segments.
[0009] In a preferred embodiment, the method for dividing the road space into continuous mapping units based on control points and direction information, setting the spatial range and main axis direction and area attribution mapping for the mapping units, and constructing independent mapping task packages for overlapping area mapping includes the following steps: Construct a spatial principal axis reference set based on the output control point sequence and the corresponding direction vector set. Each pair of adjacent control points is regarded as a segment interval, and the direction vector represents the basic segment principal axis direction. After obtaining the main axis reference of each segment, for the kth segment, the spatial area center of the corresponding local mapping unit is recorded as: ; The boundary region of the mapping unit is constructed as follows: , where Box represents the spatial enclosing area defined by the center and direction parameters; is the main direction vector; its length is L; its vertical direction is , the width W is set by the environmental constraints; the height direction is set to H; After the structure is constructed, each mapping unit area is used as a spatial carrier for independent image reconstruction, and the image data, point cloud data and texture mapping processing are all limited to the mapping unit. According to the overlapping length of the mapping unit, , make an equidistant extension on the original mapping unit length L, and determine the effective area of the mapping unit as: ; In the process of image frame attribution, if the center point of the image frame falls into the valid area of the adjacent mapping unit, the image frame will be assigned to two mapping units for redundant mapping; After completing the division of spatial mapping units, the spatial parameters, image data, direction constraints and reconstruction configuration of each mapping unit are structured and packaged to form an independent mapping task package; Construct a task unit for each mapping unit: ,in, is a set of image frames belonging to the mapping unit; is the set of RTK trajectory points corresponding to the image frame; is the main direction vector; is the center position of the mapping unit.
[0010] In a preferred embodiment, the steps for extracting the structural semantic area and edge information in the image, constructing a structural mask and mapping it to the three-dimensional space, and marking it as a structural enhancement point are as follows: Extract regional information with road semantics and geometric structure significance from image frames assigned to local mapping units, and use a joint semantic segmentation and structural edge extraction mechanism to identify typical road structure areas in each image frame; Perform pixel-level structural classification on the image to obtain a semantic label map; At the same time, the Canny operator is used for gradient edge detection to obtain the image edge structure map; Combine the semantic map with the edge map to construct a multi-structure joint mask; After image feature extraction, the camera pose and spatial orientation information obtained by RTK registration is used to perform three-dimensional projection of the high-confidence structure area in each image to obtain a set of structural guide points, which are used as structural enhancement points for structural enhancement.
[0011] In a preferred embodiment, after image feature extraction is completed, the camera pose and spatial orientation information obtained by RTK registration are used to perform three-dimensional projection on the high-confidence structure area in each image to obtain a set of structure guide points, which are used as structure enhancement points for structure enhancement. The specific steps are as follows: For each image frame, the pixel-level mask is projected into a set of points in three-dimensional space using the camera's internal and external parameters; Convert the two-dimensional image coordinates into normalized camera coordinate system vectors, and map them into three-dimensional space projection directions through camera extrinsics; Combine the RTK trajectory points corresponding to the image frame to generate the line of sight space points, and record the three-dimensional points after all significant structures are projected as a three-dimensional point set; In the point cloud generation stage, an adaptive sampling mechanism is used for the significant structure area to enhance the spatial position of the structure point area. Use the reconstruction pipeline as a sampling reference constraint to construct a spatial density adjustment factor : ,in, Represents the candidate position for reconstruction of the current point cloud, Enhanced impact range thresholds for structures; In response to the slope control factor, q represents the spatial position coordinates of the structural reinforcement point area; After the initial generation of the point cloud, the structure area is subjected to geometric optimization and texture refinement. For each point cloud in the mapping unit, a local neighborhood subset close to the structure guide point is extracted, and geometric boundary repair and texture mapping optimization operations are performed. Based on the local optimization strategy of the reprojection error, the optimization goal is to minimize the multi-view reprojection error of the structure area.
[0012] In a preferred embodiment, the specific steps for identifying the overlapping area of the tile boundary structure and establishing the anchor point to complete the local rigid body registration are as follows: Between each pair of adjacent tiles, the structural salient points of the boundary area are extracted, including lane corners, curb edges, sudden corners or traffic sign outlines, to form splicing anchor points. The boundary structural feature point set of the kth and k+1th tiles is Respectively expressed as: ,in, , Respectively represent the three-dimensional coordinate vectors of the i-th and j-th structural feature points in the k-th block; Determine whether there are significant structural coincidence point pairs through nearest neighbor matching: ; And form a candidate matching set: ,in, is the set of candidate matching structure point pairs of blocks k and k+1; If satisfied , then the two blocks have good spatial connectivity at the structural boundary, and further registration is performed, where is the minimum matching point pair number threshold; Initial rigid body registration is performed using structural coincident point pairs to unify adjacent tiles into a continuous road principal axis coordinate system; In obtaining the set of structural coincidence points Finally, construct the rigid body transformation and determine the rigid body space transformation matrix of tile k+1 relative to tile k: ,in, represents the rigid body transformation group in three-dimensional space, is the rotation matrix in three-dimensional space, indicating the change of attitude direction; t is the translation vector, indicating the displacement of the coordinate origin; Minimize the coincidence error function: , get the transformation relationship from block k+1 to block k coordinate system; Transform all point clouds in the k+1th tile into: ,in, is the original 3D coordinate of the jth point cloud point in tile k+1; It means that the points in the k+1 block are mapped to the new coordinates in the block k coordinate system after rigid body transformation.
[0013] In a preferred embodiment, a global error minimization model is reconstructed to uniformly optimize the postures and positions of all image blocks and perform an overall road space survey. The specific steps are as follows: After completing the local rough alignment, use the global alignment optimization mechanism to make all tiles have a consistent coordinate system in the global road space; Transform the coordinates of each of the N tiles contained into: ,in, is the global pose of the kth tile; Describe the orientation / posture change of tile k in space; Represents the position offset of tile k relative to the global coordinate system; Minimize the structure point coincidence error between all tiles: ,in, , Represent the coordinates of the structural points of tiles i and j respectively; , Respectively represent the global pose transformation of tiles i and j; SY is the index set of all tile pairs with overlapping structural points; The coincidence error of structural points between blocks is minimized, the optimal transformation of each block is solved, and the optimal transformation result is used as the input of road scene modeling to conduct overall road space survey.
[0014] The technical effects and advantages of the road scene space rapid survey system combining RTK and image reconstruction of the present invention are as follows: The present invention collects high-precision RTK trajectory data during road inspections, synchronizes it with image frames, builds a one-to-one correspondence between images and spatial positions, designs continuity detection and anomaly elimination mechanisms for common multipath effects and pseudo-jump point problems in urban roads, extracts physically reasonable and spatially stable control point sequences, builds a road topology skeleton structure based on control points and direction vectors, and divides the entire road space into multiple local mapping units (Tiles) on this basis. Each unit has a main axis direction and a boundary range, and redundant areas are built by setting spatial overlap zones to ensure structural connectivity between tiles. The system further extracts semantic structural information such as lane lines, curbs, and signs from the image, fuses edge detection to build a structural mask, and maps it to three-dimensional space for guiding density distribution adjustment in point cloud reconstruction. Finally, local pose alignment of tile boundaries is achieved through structural point matching, a global error minimization model is constructed, and the poses and positions of all tiles are uniformly optimized to achieve continuous splicing and structural consistency of the road scene spatial model, thereby improving the mapping efficiency in road space scenes. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] Figure 1 The present invention is a schematic diagram of the structure of a road scene space rapid survey system combining RTK and image reconstruction. DETAILED DESCRIPTION
[0016] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0017] In order to achieve the above objectives, Figure 1A schematic diagram of the structure of a road scene space rapid survey system combining RTK and image reconstruction is given in the present invention, which specifically includes a road trajectory acquisition module, a road mapping analysis module, a structure enhancement module and a road survey module, and each module is connected by a signal; The road trajectory acquisition module is used to collect road inspection trajectories and synchronize image frames with trajectory points, remove abnormal jump points, extract continuous control points, and establish a road topology skeleton based on control points and direction vectors; The road mapping analysis module is used to divide the road space into continuous mapping units based on control points and direction information, set the spatial range and main axis direction and regional attribution mapping for the mapping units, and construct independent mapping task packages to map overlapping areas; The structure enhancement module is used to extract the structural semantic area and edge information in the image, construct a structure mask and map it to the three-dimensional space, and mark it as a structure enhancement point; The road survey module is used to identify the overlapping areas of tile boundary structures and establish anchor points, complete local rigid body registration, and then build a global error minimization model to uniformly optimize the posture and position of all tiles and conduct overall road space survey.
[0018] In the multi-source perception data collection of road space, RTK-GNSS (real-time dynamic high-precision satellite positioning system) usually collects position information independently at a high frequency, while image acquisition devices (such as cameras or multi-view cameras) collect image sequences based on frame rates. For example, RTK provides high-frequency (such as above 10Hz) spatial positioning data, while the image acquisition frame rate is relatively low (such as 5Hz to 15Hz). Since the two types of equipment have differences in time reference and data update cycle, if there is no unified time reference or synchronization strategy, it is very easy to cause time drift or spatial alignment error between image frames and trajectory points, which in turn affects the accuracy of subsequent image reconstruction and spatial modeling; Step 1: construct a priori road topology for RTK trajectories. Based on the high-precision positioning characteristics of RTK, a road topology prior model with spatial continuity, attitude reference and structural expressiveness is constructed. The specific steps are as follows: Perform RTK trajectory data acquisition and time synchronization. During the acquisition phase, a time synchronization mechanism is established between image frames and RTK trajectory points. Each image frame is bound to a spatial trajectory point through the principle of minimum time difference. The three-dimensional spatial positioning data collected in the road scene is used as RTK trajectory data. The RTK trajectory data corresponds to the real-time spatial position of each frame of image collection in the road scene. The collection object is the installation position of the image acquisition terminal (such as vehicle-mounted cameras, mobile work platforms, handheld devices, trajectory scanning devices, etc.), and the collection path is all the spatial points passed along the road topology path during the road inspection process; The coverage areas include but are not limited to: motor vehicle lanes, sidewalks, bicycle lanes; the intersection area of curbs and lane lines; turning corners, intersections, ramps, and intersections; under viaducts, culverts, overpasses, and other areas with strong spatial constraints.
[0019] Perform RTK track point collection, collect as a track point sequence, set as: ,in, , Represents the three-dimensional positioning coordinates at the i-th moment, located in the geocentric coordinate system or the local projection coordinate system, Indicates the acquisition timestamp of the trajectory point, recorded in a unified system reference time (such as UTC time or system relative time); The image frame data is represented as: ,in, is the image frame; is the image frame acquisition time; For each image frame, find the RTK trajectory point closest to the timestamp, recorded as: The matching result is , used to represent the temporal binding of images to spatial points; Set the time match tolerance threshold , if the match is satisfied: , the data pair is retained; otherwise the image frame is discarded, usually The value range is 0.02 to 0.05 seconds, which can be determined according to the RTK frequency and camera synchronization capability; Representation and Image Frame The timestamp of the closest RTK track point.
[0020] Although the RTK system has high-precision positioning capabilities, it is often affected by high-rise buildings, electromagnetic interference or multipath effects in non-ideal environments such as urban roads, resulting in sudden pseudo jump points. Such jump points are usually accompanied by sudden changes in direction or a sharp increase in distance, which seriously damages the continuity and structural expression of the trajectory. If not processed, it will directly lead to subsequent module division offset, direction estimation failure and even image reconstruction collapse. Therefore, the collected trajectory sequence is subjected to continuity detection and outlier removal processing to restore the physical rationality and spatial consistency of the trajectory; To smooth the RTK trajectory and remove false jump points, the specific steps are as follows: Suppose any three consecutive trajectory points are , calculate the direction vector: ,in, It is from the point Pointing Point The direction vector of It is from the point Pointing Point The direction vector of Structural direction change angle: ,in, represents the vector dot product, Represents the Euclidean norm of a vector; represents the turning angle of the trajectory at the i-th point, in radians;
[0021] Calculate the spatial distance between adjacent trajectory points: ,like , it means there is a violent space jump, is the distance threshold; The pseudo jump point determination rules are set as follows: If a point satisfies both: and , it is determined to be an abnormal jump point and recorded as a pseudo jump point; is the set turning angle threshold; If the point only jumps in isolation, it is directly removed. If the trajectory points before and after the jump point are continuous, linear interpolation or spline reconstruction is used to fill the gap. Summarize the interpolation points to form the final set of control points: ,in, is a three-dimensional Euclidean real space, that is, a three-dimensional vector space composed of three real coordinates, such as a set of triplets of (x, y, z); n is the total number of extracted control points, which changes dynamically depending on the complexity of the path. From the original trajectory point set or interpolated points, the order is kept consistent with the trajectory path as space partition anchor points and skeleton graph nodes.
[0022] In the mapping based on image reconstruction, the camera posture (i.e., viewing direction) directly affects the registration accuracy between images, the stability of depth estimation, and the projection direction of the point cloud. If the initialization of the camera direction is ignored, the direction of the local mapping unit will diverge and the reconstruction result will be discontinuous. Therefore, it is necessary to estimate the direction vector based on the spatial relationship between the control point pairs to initialize the image posture and the spatial module orientation. The specific steps for control point posture estimation and viewpoint constraint establishment are as follows; Construct a direction vector for any pair of consecutive k-th and k+1-th control points , define the unit direction vector as: ,in, , represents the main axis direction of the kth spatial segment, which is used to indicate the main observation direction of the camera in this area or the extension direction of the road; Construct a set of attitude vectors and a set of directions: , each direction vector With corresponding control points Bind to form a posture reference pair: ,This attitude information will be used as the orientation reference of the partition, as the main vector of the camera attitude estimation for image reconstruction, and as the attitude constraint input in the tile stitching process.
[0023] In a multi-tile, multi-path road scene, a global data structure that can express the road structure topology is constructed, that is, a road skeleton graph structure is constructed to define the graph. : ; A set of directed edges: ; Each edge records the following attribute tuple: ,in, is the side length, which indicates the spatial distance between control point pairs; is the direction angle, indicating the turning angle between road segments; A road skeleton graph is constructed for connection planning, splicing path reconstruction and direction navigation between spatial segments, and is used for error constraint propagation in the global alignment stage as a topological index structure of the road scene.
[0024] Step 2: Based on the mapping module division of the local reconstruction unit, the road space is divided into controllable and splicable modular units according to the control point sequence and road skeleton graph structure constructed in step 1. By combining the spatial distribution law, trajectory direction vector and perspective consistency principle, the overall road space reconstruction task is decomposed into several structurally independent but spatially continuous mapping units. The specific steps are as follows: Establish the space division benchmark and reconstruction axial reference as follows: According to the output control point sequence and the corresponding direction vector set, the division center and main direction of each local mapping unit are determined as the basis for the subsequent construction of spatial mapping units; Let the kth control point be , whose direction vector is , then the spatial principal axis reference set can be constructed: , where n is the total number of control points, and each pair of adjacent control points It is considered as a basic segment interval, representing a road logical unit. The direction vector indicates the main axis direction of the segment, which is used to guide the orientation layout of subsequent building blocks in space. After obtaining the main axis reference of each segment, it is necessary to construct a local mapping unit area with the control point pair as the center. Each mapping unit area should have sufficient image data corresponding to the segment and good spatial closure to facilitate subsequent point cloud generation and texture mapping; Therefore, for the kth fragment, the spatial area center of the corresponding local mapping unit is recorded as: , and set its spatial bounding box parameters as the long axis direction as the direction vector, the length is set to L (covering distance); the vertical direction , the width W is set by environmental constraints, usually consistent with the lateral width of the road; the height direction is set to H, estimated according to the vertical range of the image viewing angle; The boundary region of the mapping unit is constructed as follows: , where Box represents the spatial enclosing area defined by the center and direction parameters; after the structure is constructed, each mapping unit area will become a spatial carrier for independent image reconstruction, and the processing of image data, point cloud data, and texture mapping are all limited to the mapping unit; In the actual reconstruction process, if the mapping units are completely separated, texture discontinuity, structural fracture or splicing error is very likely to occur at the reconstruction seams. Therefore, a reconstruction redundant band design mechanism is introduced to reserve a certain spatial overlap area between adjacent mapping units for geometric alignment and feature fusion. Assume that the overlapping length of the mapping unit is , then make an equidistant extension on the original mapping unit length L, and define the effective area of the mapping unit as: , in the process of image frame attribution, if the center point of the image frame falls into , then the frame will be assigned to two mapping units for redundant mapping and used for feature fusion processing between the two mapping units; Overlap adjacent segments to a certain extent in the length direction. The images in the overlapping area will participate in the reconstruction process of two mapping units at the same time. Adjacent mapping units will have a common structure in the reconstruction results, which can be used for boundary matching, posture calibration and error fusion in the subsequent stitching stage. After completing the division of spatial mapping units, the system needs to structurally encapsulate the spatial parameters, image data, directional constraints and reconstruction configuration of each mapping unit to form an independent mapping task package. For each mapping unit, a task unit is constructed: ,in, is a set of image frames belonging to the mapping unit; is the set of RTK trajectory points corresponding to the image frame; is the main direction vector; It is the center position of the mapping unit, which is used for path indexing and local coordinate system initialization.
[0025] Starting from the control point drive, a mapping unit division method with directional constraints, structural continuity and redundant robustness is constructed. Each segment has a clear spatial boundary, controllable calculation range and precise image attribution logic, thereby improving the spatial organization ability and construction performance of the overall system. Figure 1 Consistency.
[0026] Step 3: Perform structural feature-guided road element enhancement reconstruction, that is, identify and enhance the reconstruction of areas with structural constraints and semantic significance in the road scene (such as lane lines, edges, curbs, signboards, etc.), and use image structural feature extraction and mapping resources combined with directional consistency constraints and spatial boundary geometry guidance strategies to enhance the reconstruction quality of key areas. The specific steps are as follows: The regional information with road semantics and geometric structure significance is extracted from the image frames assigned to the local mapping unit as the enhancement object for subsequent spatial reconstruction. The semantic segmentation and structural edge joint extraction mechanism is adopted to identify the typical road structure area in each image frame. Assume that the image frame is , where m=1, 2, …, M, is used to represent the image frame number in the local mapping unit; Introducing a pre-trained road scene structure semantic segmentation model , classify the image at the pixel level and obtain the semantic label map: ; At the same time, the Canny operator is used for gradient edge detection to obtain the image edge structure diagram: ,in, It is the Gaussian kernel scale parameter for edge detection, which is used to adjust the smoothness of edge response; Combine the semantic map with the edge map to construct a multi-structure joint mask: ;in, For structural label areas (such as lane lines, road edges, obstacles), It is the edge response area.
[0027] This mask is used to indicate the areas in the image that are most valuable for the structural representation of the mapping process.
[0028] For example, suppose the system captures a frame of image, which is taken on a main road in the city. The picture contains multiple white lane lines, a solid yellow line, curbs and sidewalk boundaries. The image has typical road linear elements with obvious color and structure contrast. The system calls the structured semantic segmentation model to perform pixel-by-pixel label inference on the image. The model identifies areas such as lane lines, solid lines, sidewalk edges, arrow signs, etc. in the image, and generates a corresponding semantic label map. At the same time, the Canny edge detection algorithm is used to extract grayscale gradient change areas from the image to enhance the recognition of structural contours in texture blurred areas. The above two are logically intersected to retain areas that are significant in both semantics and edges to generate a structural mask area, whose corresponding image contains linear structures such as lane lines and road boundaries.
[0029] After image feature extraction, it is necessary to accurately map the salient areas of these two-dimensional images to three-dimensional space. Using the camera pose and spatial orientation information obtained by RTK registration, the high-confidence structural areas in each image are projected in three dimensions to obtain a set of structural guide points. These projection results fall on specific coordinate points in space, representing the corresponding positions of the structurally significant areas in the image in space. The specific steps are as follows: For each image frame, the pixel-level mask is projected into a set of points in three-dimensional space using the camera's internal and external parameters; Convert the two-dimensional image coordinates into normalized camera coordinate system vectors, and then map them into three-dimensional space projection directions through camera extrinsics; Combined with the RTK trajectory points corresponding to the frame, line-of-sight space points can be generated, and the three-dimensional points after projection of all significant structures are recorded as a three-dimensional point set, whose spatial distribution is used to indicate the areas where the sampling density needs to be enhanced in the point cloud reconstruction stage.
[0030] In the point cloud generation stage, an adaptive sampling mechanism is introduced for the significant structure area to improve the geometric restoration accuracy. The dense point cloud generation algorithm is assumed to take the image pair and its corresponding pose information, feature matching points, etc. as input, and the spatial position of the structure enhancement point area is converted into Introduce the reconstruction pipeline as a sampling reference constraint to construct a spatial density adjustment factor For structural enhancement: ,in, Represents the candidate position for reconstruction of the current point cloud, Enhanced impact range thresholds for structures; is the response slope control factor, which is used to control the speed of density change, and q represents the spatial position coordinates of the structural reinforcement point area; The closer the value of the structural space density adjustment factor is to 1, the closer the area is to the structural feature point, and the reconstruction algorithm needs to increase the sampling density.
[0031] After the initial generation of the point cloud, geometric optimization and texture refinement are performed on the structural area to enhance the edge clarity and surface consistency of the reconstruction results. For each point cloud in the mapping unit, a local neighborhood subset close to the structural guide point is extracted, and geometric boundary repair and texture mapping optimization operations are performed. A local optimization strategy based on the reprojection error is adopted. The optimization goal is to minimize the multi-view reprojection error in the structural area, so that the point cloud reconstruction has more image consistency in the semantically important areas.
[0032] Since each tile may be generated by different path segments, its image acquisition has direction changes, RTK error disturbances, and local view occlusion problems, resulting in the reconstruction results of different tiles being in different coordinate systems and unable to be directly spliced into a continuous road space model. For example, in a road scene, each local module (mapping unit) often corresponds to a linear driving path, and its boundary position is usually located at the intersection of the front and rear image reconstruction areas. However, due to factors such as image quality, attitude error, and RTK positioning jitter, adjacent tiles are not naturally seamlessly connected in geometry; Therefore, tile stitching and global alignment guided by structural overlap are needed to ensure that the three-dimensional structure of the entire road scene has geometric continuity, semantic consistency and coordinate uniformity. It is necessary to find spatial structural overlap areas between tile boundaries as stitching anchor points. The spatial structure not only refers to the overlap of geometric point clouds, but also refers to the spatial consistency clues formed by structural features (such as lane line edges, sidewalk corners, and road signs).
[0033] Step 4: perform local tile stitching and global alignment of the road scene space. The specific steps are as follows: Between each pair of adjacent tiles, we first extract the structural salient points in the boundary area, including lane corners, curb edges, sudden corners or traffic sign outlines, to form the splicing anchor points, i.e., the boundary structural feature point set of the kth and k+1th tiles. They are: ,in, , Respectively represent the three-dimensional coordinate vectors of the i-th and j-th structural feature points in the k-th block; Determine whether there are significant structural coincidence point pairs through nearest neighbor matching: , forming a candidate matching set: ,in, is the set of candidate matching structure point pairs of blocks k and k+1; If satisfied: , it is considered that the two blocks have good spatial connectivity at the structural boundary and can be further aligned. is the minimum matching point pair number threshold, which is used to determine whether there is valid structural overlap; Since each tile is constructed based on local control points, its reconstructed coordinate system is a relative coordinate system. Due to factors such as direction estimation fluctuations, RTK attitude disturbances, and inconsistent image acquisition, different tiles may have coordinate offsets or attitude rotation errors in the global space. Therefore, structural coincidence points are used for initial rigid body registration (attitude alignment) to unify adjacent tiles into a continuous road principal axis coordinate system. Input the set of structural point pairs, that is, the set of matching points obtained using the previous substep: In obtaining the set of structural coincidence points Finally, the rigid body transformation is constructed to determine the rigid body space transformation matrix of tile k+1 relative to tile k: ,in, represents the rigid body transformation group in three-dimensional space (including rotation and translation), is the rotation matrix in three-dimensional space, indicating the change of attitude direction; t is the translation vector, indicating the displacement of the coordinate origin; By minimizing the coincidence error function: , and obtain the transformation relationship from block k+1 to block k coordinate system. Then transform all point clouds in block k+1 into: ,in, is the original 3D coordinate of the jth point cloud point in tile k+1; It means that the points in the k+1 block are mapped to the new coordinates in the block k coordinate system after rigid body transformation, so as to achieve coordinate consistency; After completing the local rough alignment, in order to eliminate the accumulated errors at the tile level (especially in long paths or loop closure scenarios), a global alignment optimization mechanism is used to make all tiles have a consistent coordinate system in the global road space; Transform the coordinates of each of the N blocks in the system into: ,in, is the global pose (attitude and position) of the kth tile, which is used to describe how the tile is placed in three-dimensional space; Describe the orientation / posture change of tile k in space; Represents the position offset of tile k relative to the global coordinate system, that is, the displacement of the tile coordinate origin relative to the global origin; Construct a global optimization problem to minimize the structural point coincidence error between all tiles: ,in, , Represent the coordinates of the structural points of tiles i and j respectively; , Respectively represent the global pose transformation of tiles i and j; SY is the index set of all tile pairs with overlapping structural points; By minimizing the coincidence error of the structural points between the above-mentioned blocks, the system automatically solves the optimal transformation of each block, making all the coincident structural points as consistent as possible in the global coordinates (with the minimum coincidence error), thereby achieving posture consistency, high-precision stitching and path continuity construction, thereby solving the problem of local reconstruction error accumulation, and establishing a multi-source fusion alignment framework with RTK trajectory as the initial benchmark and image structural points as the constraint core. The optimal transformation result is used as the input for road scene modeling, and the overall road space survey is carried out.
[0034] For example, a section of a curved road in a suburban city is being modeled with high precision in 3D for road expansion design and municipal maintenance. The survey team deployed a road survey system, using the vehicle-mounted RTK-GNSS system to record the vehicle's position and posture in real time, and simultaneously equipped with multiple high-definition wide-angle cameras to collect images on both sides of the road; the system automatically divides the entire road into multiple image tiles (tiles 1, 2, 3, etc.), each tile covering a space of about 10 meters; The boundary area of tiles 1 and 2 is the same curve. The guardrails, curbs and other structures captured by the camera should overlap. However, due to the vibration, tilt and image reconstruction error of the vehicle during driving, the guardrail point clouds in tiles 1 and 2 are misaligned by tens of centimeters, and the attitude direction is also slightly deviated (for example, one side is slightly tilted up and the other side is slightly tilted down). If alignment is not performed, the entire road skeleton will not be coherent, resulting in spatial fractures and increased measurement errors. In the boundary area between Tile 1 and Tile 2, the system automatically identifies structural points such as the top of the guardrail and the intersection of the white line on the ground, and establishes a point pair set; The coordinate transformation matrix of all tiles is included in the optimization variable, and the goal is to minimize the structural point error between all tile pairs; The result after optimization and solution is that block 1 is fine-tuned and translated to the left, and block 2 is rotated about 0.5 degrees downward, so that the boundary guardrail structure points are aligned in three-dimensional space. The boundaries of all block paths are continuous, and adjacent structures are naturally integrated. The constructed three-dimensional road model path is continuous and seamless, with a natural posture, and the error converges to the centimeter level.
[0035] The threshold information in this embodiment is pre-set by professionals and will not be explained in detail here. Some parameter English letters in the embodiments have the same situation, but different meanings are explained when used, which will not be explained one by one here.
[0036] The present invention collects high-precision RTK trajectory data during road inspections, synchronizes it with image frames, builds a one-to-one correspondence between images and spatial positions, designs continuity detection and anomaly elimination mechanisms for common multipath effects and pseudo-jump point problems in urban roads, extracts physically reasonable and spatially stable control point sequences, builds a road topology skeleton structure based on control points and direction vectors, and divides the entire road space into multiple local mapping units (Tiles) on this basis. Each unit has a main axis direction and a boundary range, and redundant areas are built by setting spatial overlap zones to ensure structural connectivity between tiles. The system further extracts semantic structural information such as lane lines, curbs, and signs from the image, fuses edge detection to build a structural mask, and maps it to three-dimensional space for guiding density distribution adjustment in point cloud reconstruction. Finally, local pose alignment of tile boundaries is achieved through structural point matching, a global error minimization model is constructed, and the poses and positions of all tiles are uniformly optimized to achieve continuous splicing and structural consistency of the road scene spatial model, thereby improving the mapping efficiency in road space scenes.
[0037] The above formulas are all dimensionless and numerical calculations. The formula is a formula that is closest to the actual situation obtained by collecting a large amount of data and performing software simulation. The preset parameters in the formula are set by technicians in this field according to actual conditions.
[0038] The above embodiments may be implemented in whole or in part by software, hardware, firmware or any other combination. When implemented by software, the above embodiments may be implemented in whole or in part in the form of a computer program product.
[0039] Those of ordinary skill in the art will appreciate that the modules and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.
[0040] In addition, each functional module in each embodiment of the present application may be integrated into one processing module, or each module may exist physically separately, or two or more modules may be integrated into one module.
[0041] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art who is familiar with the present technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.
[0042] Finally: The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the protection scope of the present invention.
Claims
1. A road scene spatial rapid survey system combining RTK and image reconstruction, characterized by: It includes a road trajectory acquisition module, a road mapping and analysis module, a structure enhancement module, and a road survey module, and each module is connected by signals; The road trajectory acquisition module is used to collect road inspection trajectories and synchronize image frames with trajectory points, remove abnormal jump points, extract continuous control points, and establish a road topology skeleton based on control points and direction vectors; The road mapping analysis module is used to divide the road space into continuous mapping units based on control points and direction information, set the spatial range and main axis direction and regional attribution mapping for the mapping units, and construct independent mapping task packages to map overlapping areas; The structure enhancement module is used to extract the structural semantic area and edge information in the image, construct a structure mask and map it to the three-dimensional space, and mark it as a structure enhancement point; The road survey module is used to identify the overlapping areas of tile boundary structures and establish anchor points, complete local rigid body registration, and then build a global error minimization model to uniformly optimize the posture and position of all tiles and conduct overall road space survey.
2. The road scene spatial rapid survey system combining RTK and image reconstruction according to claim 1, characterized in that: It is used to collect road inspection tracks and synchronize image frames with track points. The specific steps are as follows: The three-dimensional spatial positioning data collected in the road scene is used as RTK trajectory data. The RTK trajectory data corresponds to the real-time spatial position of each frame of image collection in the road scene. The collection object is the installation position of the image acquisition terminal, including vehicle-mounted cameras, mobile work platforms, handheld devices, and trajectory scanning devices; the collection path is all the spatial points passed along the road topology path during the road inspection process; The coverage area includes motor vehicle lanes, sidewalks, bicycle lanes, curbs, lane line intersections; turning points, intersections, ramps, intersections, under viaducts, culverts, and overpasses; Perform RTK trajectory point collection as a trajectory point sequence. The RTK trajectory point includes the three-dimensional positioning coordinates corresponding to the collection position and the corresponding timestamp; The collected image frame data includes image frames and image frame collection time; For each image frame, find the RTK track point closest to the timestamp and perform time binding between the image and the track point; Set the time matching tolerance threshold. If the absolute value of the difference between the image frame acquisition time and the timestamp of the closest RTK track point is less than or equal to the matching tolerance threshold, the frame image is bound to an RTK track point and the pairing data is retained.
3. The road scene space rapid survey system combining RTK and image reconstruction according to claim 2, characterized in that: Remove abnormal jump points, extract continuous control points, and establish a road topology skeleton based on control points and direction vectors. The specific steps include: Perform RTK trajectory smoothing and pseudo-jump point elimination, and convert three consecutive trajectory points into , the direction vector is calculated as ,in, It is from the point Pointing Point The direction vector, It is from the point Pointing Point The direction vector of Structural direction change angle: ,in, Represents the Euclidean norm of a vector; represents the turning angle of the trajectory at the i-th point; Calculate the spatial distance between adjacent trajectory points : ,like , there is a drastic space jump, is the distance threshold; The pseudo jump point determination rules are set as follows: If the trajectory point satisfies both: and , it is determined to be an abnormal jump point and recorded as a pseudo jump point; is the set turning angle threshold; If the pseudo jump point is only isolated, it is directly removed. If the trajectory points before and after the jump point are continuous, linear interpolation or spline reconstruction is used to fill the gap; Summarize the interpolation points to form the final set of control points: ,in, is a three-dimensional Euclidean real number space, is the kth control point; n is the total number of extracted control points; Each control point from the trajectory point set or interpolation point is used as a space partition anchor point and skeleton graph node.
4. The road scene spatial rapid survey system combining RTK and image reconstruction according to claim 3, characterized in that: Each control point from the trajectory point set or interpolation point is used as a space partition anchor point and skeleton graph node. The specific steps are as follows: Construct a direction vector for any pair of consecutive k-th and k+1-th control points , C is the control point set, and the unit direction vector is defined as: ,in, , represents the main axis direction of the kth space segment; Construct a set of attitude vectors and a set of direction vectors: , each direction vector With corresponding control points Bind to form a posture reference pair: , the pose information is used as the orientation reference for the partition, and as the main vector for initializing the camera pose estimation for image reconstruction, and as the pose constraint input during the tile stitching process; Construct the road skeleton graph structure and define the road skeleton graph as: : ; A set of directed edges: ; Each edge records the following attribute tuple: ,in, is the side length, which indicates the spatial distance between control point pairs; is the direction angle, which indicates the turning angle between road segments.
5. The road scene spatial rapid survey system combining RTK and image reconstruction according to claim 4, characterized in that: It is used to divide the road space into continuous mapping units based on control points and direction information, set the spatial range and main axis direction and area attribution mapping for the mapping units, and construct independent mapping task packages for overlapping area mapping, including the following steps: Construct a spatial principal axis reference set based on the output control point sequence and the corresponding direction vector set. Each pair of adjacent control points is regarded as a segment interval, and the direction vector represents the basic segment principal axis direction. After obtaining the main axis reference of each segment, for the kth segment, the spatial area center of the corresponding local mapping unit is recorded as: ; The boundary region of the mapping unit is constructed as follows: , where Box represents the spatial enclosing area defined by the center and direction parameters; is the main direction vector; its length is L; its vertical direction is , the width W is set by the environmental constraints; the height direction is set to H; After the structure is constructed, each mapping unit area is used as a spatial carrier for independent image reconstruction; According to the overlapping length of the mapping unit, , make an equidistant extension on the original mapping unit length L, and determine the effective area of the mapping unit as: ; In the process of image frame attribution, if the center point of the image frame falls into the valid area of the adjacent mapping unit, the image frame will be assigned to two mapping units for redundant mapping; After completing the division of spatial mapping units, the spatial parameters, image data, direction constraints and reconstruction configuration of each mapping unit are structured and packaged to form an independent mapping task package; Construct a task unit for each mapping unit: ,in, is a set of image frames belonging to the mapping unit; is the set of RTK trajectory points corresponding to the image frame; is the main direction vector; is the center position of the mapping unit.
6. The road scene space rapid survey system combining RTK and image reconstruction according to claim 5, characterized in that: It is used to extract the structural semantic area and edge information in the image, construct a structural mask to map it to the three-dimensional space, and mark it as a structural enhancement point. The specific steps are as follows: Extract regional information with road semantics and geometric structure significance from image frames assigned to local mapping units, and use a joint semantic segmentation and structural edge extraction mechanism to identify road structure regions in each image frame; Perform pixel-level structural classification on the image to obtain a semantic label map; At the same time, the Canny operator is used for gradient edge detection to obtain the image edge structure map; Combine the semantic label map with the edge structure map to construct a multi-structure joint mask; After image feature extraction, the camera pose and spatial orientation information obtained by RTK registration is used to perform three-dimensional projection of the high-confidence structure area in each image to obtain a set of structural guide points, which are used as structural enhancement points for structural enhancement.
7. The road scene space rapid survey system combining RTK and image reconstruction according to claim 6, characterized in that: After image feature extraction, the camera pose and spatial orientation information obtained by RTK registration are used to perform three-dimensional projection of the high-confidence structure area in each image to obtain a set of structural guide points, which are used as structural enhancement points for structural enhancement. The specific steps are as follows: For each image frame, the pixel-level mask is projected into a set of points in three-dimensional space using the camera's internal and external parameters; Convert the two-dimensional image coordinates into normalized camera coordinate system vectors, and map them into three-dimensional space projection directions through camera external parameters; Combine the RTK trajectory points corresponding to the image frame to generate the line of sight space points, and record the three-dimensional points after all significant structures are projected as a three-dimensional point set; In the point cloud generation stage, an adaptive sampling mechanism is used for the significant structure area to enhance the spatial position of the structure point area. Use the reconstruction pipeline as a sampling reference constraint to construct a spatial density adjustment factor : ,in, Represents the candidate position for reconstruction of the current point cloud, Enhanced impact range thresholds for structures; is the response slope control factor; q is the spatial position coordinate of the structural reinforcement point area; After the initial generation of the point cloud, the structure area is subjected to geometric optimization and texture refinement. For each point cloud in the mapping unit, a local neighborhood subset close to the structure guide point is extracted, and geometric boundary repair and texture mapping optimization operations are performed. Based on the local optimization strategy of the reprojection error, the optimization goal is to minimize the multi-view reprojection error of the structure area.
8. The road scene space rapid survey system combining RTK and image reconstruction according to claim 7, characterized in that: It is used to identify the overlapping area of the tile boundary structure and establish anchor points to complete local rigid body registration. The specific steps are as follows: Between each pair of adjacent tiles, the structural salient points of the boundary area are extracted, including lane corners, curb edges, sudden corners or traffic sign outlines, to form splicing anchor points. The boundary structural feature point set of the kth and k+1th tiles is Respectively expressed as: ,in, , Respectively represent the three-dimensional coordinate vectors of the i-th and j-th structural feature points in the k-th block; is a three-dimensional Euclidean real number space; Determine whether there are significant structural coincidence point pairs through nearest neighbor matching: ; And form a candidate matching set: ,in, is the set of candidate matching structure point pairs of blocks k and k+1; If satisfied , then the two blocks have good spatial connectivity at the structural boundary, and further registration is performed, where is the minimum matching point pair number threshold; Initial rigid body registration is performed using structural coincident point pairs to unify adjacent tiles into a continuous road principal axis coordinate system; In obtaining the set of structural coincidence points Finally, construct the rigid body transformation and determine the rigid body space transformation matrix of tile k+1 relative to tile k: ,in, represents the rigid body transformation group in three-dimensional space, is the rotation matrix in three-dimensional space, indicating the change of attitude direction; t is the translation vector, indicating the displacement of the coordinate origin; Minimize the coincidence error function: , get the transformation relationship from block k+1 to block k coordinate system; Transform all point clouds in the k+1th tile into: ,in, is the original 3D coordinate of the jth point cloud point in tile k+1; It means that the points in the k+1 block are mapped to the new coordinates in the block k coordinate system after rigid body transformation.
9. The road scene space rapid survey system combining RTK and image reconstruction according to claim 8, characterized in that: Then build a global error minimization model, uniformly optimize the posture and position of all tiles, and conduct an overall road space survey. The specific steps are as follows: After completing the local rough alignment, use the global alignment optimization mechanism to make all tiles have a consistent coordinate system in the global road space; Transform the coordinates of each of the N tiles contained into: ,in, is the global pose of the kth tile; Describe the orientation / posture change of tile k in space; Represents the position offset of tile k relative to the global coordinate system; Minimize the structure point coincidence error between all tiles: ,in, , Represent the coordinates of the structural points of tiles i and j respectively; , Respectively represent the global pose transformation of tiles i and j; SY is the index set of all tile pairs with overlapping structural points; The coincidence error of structural points between blocks is minimized, the optimal transformation of each block is solved, and the optimal transformation result is used as the input of road scene modeling to conduct overall road space survey.
Citation Information
Patent Citations
Multi-sensor fusion road extraction and indexing method based on global and local grid maps
CN111273305A
Road disease intelligent identification method and system based on image identification technology
CN119068268A
Road traffic anomaly detection and related equipment based on holographic perception
CN119540832A
News scene three-dimensional reconstruction and visualization method based on multi-source remote sensing data
CN119904592A
Monocular image-based method for updating road signs and markings
WO2021197341A1
Cited By
Multi-source data driven real estate modeling method and system
CN121053316A
Field survey data real-time modeling method and system based on edge calculation
CN121095495A
Road network disease image recognition method and system based on deep learning
CN122020351A