An analytical method for generating VR panoramas of railways by integrating 3D laser point clouds

By using multi-sensor collaborative acquisition and a depth compensation model, the problem of lack of spatial measurement in railway scenarios in VR panoramic technology has been solved, achieving efficient spatial analysis and measurement, and improving the construction efficiency and analysis capabilities of railway virtual scenarios.

CN122312975APending Publication Date: 2026-06-30CHINA RAILWAY DESIGN GRP CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CHINA RAILWAY DESIGN GRP CO LTD
Filing Date
2026-03-24
Publication Date
2026-06-30

AI Technical Summary

Technical Problem

Existing VR panoramic technology lacks accurate spatiotemporal references and geometric information in railway scenarios, making it impossible to perform spatial measurements and analysis. Furthermore, existing image and point cloud fusion methods suffer from limited field of view coverage, insufficient spatial continuity, and spatiotemporal inconsistencies, making it difficult to meet the needs of efficient modeling and rapid updates for complex railway scenarios.

Method used

By employing multi-sensor collaborative acquisition technology, and through time and space synchronization, combined with semantic segmentation networks and depth compensation models, we can achieve efficient fusion of panoramic images and 3D laser point clouds to construct a VR panorama with precise spatial coordinates.

Benefits of technology

It enables spatial analysis and measurement while providing immersive browsing, improving the flexibility and efficiency of data acquisition, enhancing the spatial cognition and interpretation capabilities of virtual railway scenes, and improving the construction efficiency and analysis capabilities of complex railway scenes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122312975A_ABST
    Figure CN122312975A_ABST
Patent Text Reader

Abstract

This invention discloses an analytical method for generating railway VR panoramas by integrating 3D laser point clouds, comprising: S1, collaboratively acquiring panoramic images and 3D laser point clouds for the railway environment; S2, mapping pixel coordinates in the panoramic image to spherical angle coordinates to initially construct the panoramic sphere; for the 3D laser point cloud, calculating its spherical angle coordinates inversely; achieving the binding of geometric and color information, and saving the binding relationship to the panoramic sphere; S3, extracting image semantic pairs from the panoramic image; extracting spatial instance objects from the 3D laser point cloud; using these to obtain cross-modal semantic objects through semantic alignment; and then constructing a railway scene feature model; S4, constructing and optimizing the depth compensation model of the railway VR panorama for depth compensation of the panoramic sphere; and then combining the railway scene feature model to perform geometric constraint correction on key scene objects in the panoramic sphere, thereby constructing an analytical railway VR panorama.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of virtual reality, and more specifically to an analytical method for generating VR panoramas of railways that integrates three-dimensional laser point clouds. Background Technology

[0002] my country's railway system is rapidly developing towards intelligence, digitalization, and informatization. As a crucial component of the national comprehensive transportation system, the planning, construction, operation, and maintenance of railways place higher demands on high-precision, high-efficiency, and highly realistic virtual environments. Against this backdrop, Virtual Geographic Environments (VGEs) technology is gradually becoming one of the key technologies supporting the full lifecycle management of railway engineering. Virtual railway scenarios, as a specific application of VGEs in the railway field, possess highly visualized, interactive, simulated, and analyzable characteristics, providing strong digital support for railway design, construction, operation and maintenance, disaster simulation, and emergency drills. Virtual railway scenarios not only help improve the digitalization level of railway engineering but also play an important role in complex environment simulation, multi-source data fusion, and intelligent decision support. However, due to the thousands of kilometers of railway lines and the complex and varied terrain, traditional methods for constructing virtual railway scenarios often face challenges such as massive data volumes, high computational resource consumption, and long production cycles, resulting in high costs.

[0003] VR panoramas, combining emerging technologies such as Virtual Reality (VR) and panoramic video, offer a novel approach to modeling and analyzing virtual railway scenes. They provide users with an immersive 360° perspective, allowing them to understand complex railway scenarios through high-fidelity interaction. Compared to traditional methods of constructing virtual railway scenes, VR panoramas eliminate the need for complex 3D reconstruction processes. By using spherical projection, they can quickly convert panoramic data from the real-world scene into a 3D scene, significantly enhancing the realism and real-time nature of the virtual railway scene. VR panoramas primarily rely on panoramic camera capture and multi-view image stitching technology to map the optical images of the real scene into two-dimensional spherical images to construct the virtual scene. This approach is characterized by rapid production, strong texture realism, and low cost. Especially when dealing with ultra-long-distance, cross-regional railway scenarios, VR panoramas achieve large-scale coverage at low cost, avoiding the high manpower and computing power required for segment-by-segment detailed processing in traditional modeling. Furthermore, through diverse panoramic acquisition platforms, VR panoramas can collect and transmit environmental information comprehensively and across scales, quickly and completely reflecting the spatiotemporal processes of scene elements and phenomena. This is of great significance for improving users' understanding of railway scenes.

[0004] However, due to the lack of precise spatiotemporal references and geometric information, such visual VR panoramas are limited to viewing only and cannot address the issues requiring measurement and analysis. Image-based depth restoration methods can construct VR panoramas with geometric information by inverting three-dimensional spatial structures from two-dimensional images. These methods mainly include monocular depth estimation methods based on deep learning and multi-view photogrammetric reconstruction methods. The former uses neural networks to predict depth from a single panoramic image, while the latter reconstructs dense point clouds through the principle of parallax. Although VR panoramas constructed by these two methods possess basic spatial measurement capabilities, their low geometric accuracy and poor environmental adaptability cannot meet the needs of high-precision spatial analysis and measurement in complex railway scenarios.

[0005] In contrast, the method of fusing panoramic images with LiDAR point clouds can simultaneously provide high-fidelity texture information and high-precision spatial geometric information. However, existing fusion methods mainly rely on feature matching between 2D images and point clouds. Their data models are usually based on the correspondence between images with limited viewpoints and local projection relationships, resulting in problems such as limited field of view coverage, insufficient spatial continuity, and poor global geometric consistency. In linear infrastructure scenarios such as railways, due to factors such as drastic changes in viewpoint, large spatial scale spans, and complex occlusion, existing 2D image and point cloud fusion methods struggle to achieve a unified representation of large-scale continuous scenes, and the fusion results often only achieve local texture mapping, making it difficult to support immersive spatial analysis and measurement applications. Furthermore, existing image and point cloud fusion processes typically employ offline post-processing, i.e., images and point clouds are acquired independently before registration and fusion processing. This type of method lacks a unified time reference and sensor motion constraints, making it susceptible to factors such as differences in acquisition time and dynamic environmental changes, leading to spatiotemporal inconsistencies between cross-modal data, thereby reducing registration accuracy and fusion efficiency. Meanwhile, post-processing registration usually relies on a large number of iterative optimization calculations, which are complex, computationally expensive, and have poor real-time performance. This makes it difficult to meet the application requirements of efficient modeling and rapid updating of large-scale railway scenes, and further affects the spatial accuracy and analytical reliability of virtual scenes. Summary of the Invention

[0006] To address the problems existing in the prior art, this invention provides an efficient analytical method for generating VR panoramas of railways by fusing 3D laser point clouds.

[0007] Therefore, the present invention adopts the following technical solution:

[0008] An analytical method for generating VR panoramas of railways by integrating 3D laser point clouds includes the following steps:

[0009] S1, for the railway environment, perform coordinated acquisition of panoramic images and three-dimensional laser point clouds, wherein the coordination includes time synchronization and spatial synchronization;

[0010] S2, map the pixel coordinates in the panoramic image to spherical angle coordinates to initially complete the construction of the panoramic sphere; for the three-dimensional laser point cloud, calculate its spherical angle coordinates; based on the consistency of spherical angles, match the spatial points of the three-dimensional laser point cloud with the pixel units of the panoramic image to realize the binding of geometric information and color information, and save the binding relationship to the panoramic sphere;

[0011] S3. First, a semantic segmentation network is used to extract visual features with clear semantic meanings from the panoramic image obtained in S1, forming image semantic objects of image modalities with semantic labels. Simultaneously, stable geometric features are extracted from the 3D laser point cloud obtained in S1, and spatial instance objects of point cloud modalities with hierarchical structures are further identified. Then, the image semantic objects and spatial instance objects are semantically aligned to construct cross-modal semantic objects representing the same scene entity in different modalities. Finally, a semantically guided graph matching algorithm is used to perform feature matching on the features of the two modalities in the cross-modal semantic objects within the semantically consistent object range. A railway scene feature model containing the semantic information and spatial structural relationships of each object in the scene is constructed.

[0012] S4. Using the relative depth obtained by visual estimation from the panoramic image obtained in S1 and the real spatial depth provided by the three-dimensional laser point cloud obtained in S1, a depth compensation model for the railway VR panorama is constructed and optimized. The depth compensation model obtained by optimization is used to perform depth compensation on the panoramic sphere obtained in S2. Then, the geometric constraint correction of key scene objects in the panoramic sphere is performed by combining the railway scene feature model, and then VR technology is used to construct an analytical railway VR panorama.

[0013] In step S1 above: multiple sensors are integrated on the collaborative acquisition platform for collaborative acquisition of panoramic images and three-dimensional laser point clouds. The multiple sensors include a panoramic camera, a lidar, and an inertial measurement unit.

[0014] The time synchronization method in step S1 above is as follows: For the panoramic camera, LiDAR, and inertial measurement unit, a hardware-triggered multi-sensor time synchronization acquisition mechanism is constructed to achieve strict timestamp alignment; wherein, the multi-sensor time synchronization acquisition mechanism is achieved by building a low-level hardware synchronization system based on an STM32 microcontroller and a TTL to USB module to realize the data acquisition time synchronization between the panoramic camera and the LiDAR.

[0015] The STM32 microcontroller is set as the global master clock source. By writing high-frequency timing logic, periodic TTL pulse signals are sent to the industrial camera through the GPIO interface, forcing the camera to perform exposure acquisition at a specific physical moment.

[0016] Simultaneously, the periodic TTL pulse signal is transmitted in real time to the host computer computing platform via a TTL-to-USB module and captured as a unified time reference tag, ensuring that the acquired panoramic image and the three-dimensional laser point cloud are synchronized at the microsecond level in the time dimension.

[0017] The method for spatial synchronization in step S1 above is as follows:

[0018] Intrinsic parameter calibration of the panoramic camera was performed to obtain the imaging model and distortion parameters of the panoramic camera, ensuring the geometric consistency of the panoramic image. Based on this, the transformation formula between the camera coordinate system and the radar coordinate system established through joint calibration of the panoramic camera and LiDAR is as follows:

[0019] ;

[0020] In the formula, ( , , () is a point in the camera coordinate system, , , () represents a point in the corresponding lidar coordinate system. For rotation matrix, It is a translation matrix;

[0021] For panoramic images and 3D laser point clouds with the same timestamp, a rapid matching of multi-sensor data is achieved by using temporal semantic constraints and joint calibration methods, thus realizing spatial synchronization.

[0022] The method for achieving fast matching of multi-sensor data using temporal semantic constraints and joint calibration is as follows:

[0023] Acquire temporal semantic information, combine it with attitude angle, angular velocity and acceleration information provided by the inertial measurement unit, establish prior constraints for sensor motion, make initial estimates of the pose changes of multiple sensors and use them as initial values ​​for optimization;

[0024] Then, corner features in the panoramic image and plane and edge features in the point cloud are extracted as constraints by a joint calibration algorithm. The algorithm quickly converges and solves the relative position and attitude relationship of multiple sensors. The rotation matrix and translation vector between the two are solved by a nonlinear optimization method, and the panoramic image and the three-dimensional laser point cloud are unified to the same spatial coordinate system. The temporal semantic information includes a unified timestamp and acquisition frequency control.

[0025] In step S2 above:

[0026] The pixel coordinates in the two-dimensional panoramic image are represented by an equidistant cylindrical projection model. Mapped to spherical angular coordinates ( The preliminary construction of the panoramic sphere is complete, and the mapping relationship is as follows:

[0027] , ;

[0028] in: Indicates the horizontal direction angle; Indicates the vertical direction angle; This is a scale factor between angles and pixels, used to establish a linear correspondence between spherical angles and panoramic image coordinates;

[0029] Through mapping operations, each pixel of the panoramic image corresponds to a unique spatial orientation on the panoramic sphere;

[0030] For 3D laser point clouds, Cartesian coordinates are used ( ) to spherical coordinates ( The conversion formula for inversely calculating spherical angle coordinates () ):

[0031] , , ;

[0032] in, Indicates along direction ( The spatial radius of ) when When taking a unit value, it is used to express spatial direction; when When the actual distance value measured from the 3D laser point cloud is taken, it represents the corresponding actual 3D spatial coordinates.

[0033] In step S3 above:

[0034] The visual features include the outline of the railway track, platform boundaries, overhead contact lines, building facades, and surrounding facilities;

[0035] The geometric features include planes, linear structures, and corner points;

[0036] The spatial instance objects include the track surface and the tower column;

[0037] The semantically guided graph matching algorithm uses the image semantic objects and spatial instance objects in the cross-modal semantic objects obtained by semantic alignment as graph nodes of the image modality and point cloud modality, respectively. Each graph node contains spatial location, geometric features and semantic labels, and their spatial relationships are used as edges.

[0038] In the graph matching process, node matching calculates similarity based on the semantic label consistency of nodes and local geometric features, filters candidate corresponding node pairs, and quickly narrows down the matching range; relation matching further compares the spatial relationship consistency between nodes and verifies whether the edges between corresponding node pairs maintain the same geometric constraints, thereby realizing feature matching within the semantically consistent object range and constructing a railway scene feature model that includes the semantic information and spatial structural relationship of each object in the scene.

[0039] In step S4 above:

[0040] The depth compensation model is as follows:

[0041] ;

[0042] In the formula, The fusion depth of the panoramic sphere; The relative depth is obtained by visual estimation using the panoramic image obtained from S1; Provides the true spatial depth for the 3D laser point cloud obtained by S1; , The weighting coefficients are and satisfy the following conditions: + =1, used to balance geometric accuracy and visual continuity;

[0043] In regions with dense point clouds and stable structures, geometric accuracy is the primary consideration. In areas where point clouds are sparse or missing, visual continuity is the primary consideration. ;

[0044] The key scene objects include the track surface, facade structure, and towers.

[0045] The collaborative data acquisition platform includes handheld, quadruped robot-mounted, and UAV-borne data acquisition platforms, among which:

[0046] For areas accessible by human intervention, a handheld data collection platform enables lightweight, rapid deployment, and highly mobile data collection.

[0047] For areas with complex terrain or where stable observation by humans is difficult, a quadruped robot-mounted data acquisition platform is used to enhance autonomous movement and continuous data acquisition from multiple perspectives in complex environments.

[0048] For large-scale scenes and areas with significant elevation differences, a drone-mounted data acquisition platform is used to acquire high-altitude overhead, side, and overall structural data.

[0049] Compared with the prior art, the present invention has the following beneficial effects:

[0050] 1. While traditional VR panoramas offer high-fidelity visual performance, they lack spatial references and geometric structures, making it impossible to perform spatial measurements such as distance, area, and volume. This invention constructs a VR panorama with precise spatial coordinates by integrating three-dimensional laser point clouds, enabling users to perform spatial analysis and measurement while immersing themselves in the experience, thus achieving "what you see is what you can measure".

[0051] 2. In view of the characteristics of railway scenarios such as large spatial scale, complex terrain and heterogeneous sensors, this invention solves the problem of low registration accuracy caused by asynchronous data acquisition in traditional mode by adopting a hardware-triggered multi-sensor time synchronization acquisition mechanism and rapid matching of spatiotemporal semantic constraints, which significantly improves the flexibility and efficiency of data acquisition.

[0052] 3. In existing technologies, image and point cloud fusion often remains at the level of texture mapping or geometric alignment. This invention solves the problem of inconsistency between visual and physical scales through a multi-level fusion strategy of "data-feature-depth", enabling the generated analytical VR panorama to have stronger spatial cognition and interpretation capabilities.

[0053] 4. Innovatively combining multi-sensor collaboration, cross-modal multi-level fusion, and VR panoramic generation technology, an analytical VR panoramic generation method for complex railway scenarios is formed, which significantly improves the efficiency and analytical capabilities of railway virtual scene construction. Attached Figure Description

[0054] Figure 1 This is a flowchart of the analytical railway VR panorama generation method of the present invention;

[0055] Figure 2 This is a schematic diagram of the collaborative acquisition of panoramic images and laser point clouds in this invention;

[0056] Figure 3 This is a schematic diagram of the integrated collaborative data acquisition platform adapted to the perception of complex railway scenarios in this invention;

[0057] Figure 4 This is a schematic diagram illustrating the data-level fusion of panoramic images and three-dimensional laser point clouds based on spherical projection in this invention;

[0058] Figure 5 This is a schematic diagram of semantically guided panoramic and point cloud feature matching in this invention;

[0059] Figure 6 This is a schematic diagram of depth compensation and scale correction in this invention. Detailed Implementation

[0060] The technical solution of the invention will be clearly and completely described below with reference to the accompanying drawings and embodiments. Obviously, the following embodiments are only some embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0061] Example

[0062] An analytical method for generating VR panoramas of railways that integrates 3D laser point clouds, such as... Figure 1 As shown, it includes the following steps:

[0063] S1, panoramic image and 3D laser point cloud collaborative acquisition, specifically:

[0064] A multi-sensor system, including a panoramic camera, LiDAR, and inertial measurement unit, is integrated into the collaborative acquisition platform for the collaborative acquisition of panoramic images and 3D laser point clouds.

[0065] In view of the characteristics of railway environment, such as complex spatial structure, drastic terrain undulation, and significant differences in human accessibility, such as Figure 3 As shown, the collaborative data acquisition platform includes handheld, quadruped robot-mounted, and UAV-mounted acquisition platforms. For areas accessible by human intervention, the handheld acquisition platform enables lightweight, rapid deployment, and highly mobile data acquisition. For areas with complex terrain or where stable observation by humans is difficult, the quadruped robot-mounted acquisition platform enhances autonomous movement and multi-view continuous data acquisition capabilities in complex environments. For large-scale scenes and areas with significant elevation differences, the UAV-mounted acquisition platform enables high-altitude overhead, side-view, and overall structural acquisition.

[0066] The collaborative acquisition includes time synchronization and spatial synchronization, such as... Figure 2 As shown, where:

[0067] (1) For sensors such as panoramic cameras, lidar, and inertial measurement units, a hardware-triggered multi-sensor time synchronization acquisition mechanism is constructed to achieve strict timestamp alignment.

[0068] The multi-sensor time synchronization acquisition mechanism addresses the challenge of synchronizing data acquisition between industrial panoramic cameras and LiDAR clock sources by building a low-level hardware synchronization system based on an STM32 microcontroller and a TTL-to-USB module. The STM32 microcontroller is set as the global master clock source. High-frequency timing logic is programmed to send periodic TTL pulse signals to the industrial camera via the GPIO interface, forcing the camera to perform exposure acquisition at specific physical moments. Simultaneously, this pulse signal is transmitted in real-time to the host computer platform via the TTL-to-USB module, where it is captured as a unified time reference tag. This hardware-level triggering mechanism avoids communication delays and random jitter caused by software triggering, ensuring microsecond-level synchronization between the acquired panoramic image data (with strong visual realism) and the 3D LiDAR point cloud data (with accurate spatial geometric information) in the time dimension.

[0069] (2) In order to solve the problem of pose transformation of heterogeneous sensors in space and achieve spatial synchronization, the rigid transformation matrix between the lidar coordinate system and the panoramic camera coordinate system is solved.

[0070] First, the intrinsic parameters of the panoramic camera are calibrated to obtain the imaging model and distortion parameters of the panoramic camera, ensuring the geometric consistency of the panoramic image. Based on this, the transformation formula between the camera coordinate system and the radar coordinate system, established through joint calibration of the panoramic camera and LiDAR, is as follows:

[0071] ;

[0072] In the formula, ( , , () is a point in the camera coordinate system, , , () represents a point in the lidar coordinate system. For rotation matrix, It is a translation matrix.

[0073] For panoramic images and 3D laser point clouds with the same timestamp, a fast matching of multi-sensor data is achieved using temporal semantic constraints and joint calibration methods. The specific operation is as follows:

[0074] First, the temporal semantic information includes a unified timestamp and acquisition frequency control, which are used to time-align the data acquired by the panoramic camera and LiDAR, ensuring that the images and point clouds involved in the registration originate from the scene state at the same moment, thereby reducing registration errors caused by motion changes or viewpoint shifts. At the same time, by combining the attitude angle, angular velocity and acceleration information provided by the inertial measurement unit (IMU), prior constraints on sensor motion are established, and the initial estimation of the pose changes of multiple sensors is used as the initial value for optimization, which can improve the stability and convergence efficiency of spatial registration.

[0075] Secondly, corner features (such as building outline corners) in panoramic images and planar and edge features (such as road boundary lines, fixed features, and 3D contours) in point clouds are extracted by a joint calibration algorithm as constraints. The algorithm quickly converges and solves the relative position and attitude relationship of multiple sensors. The rotation matrix and translation vector between the two are solved by a nonlinear optimization method. This is used to unify independent visual data and geometric data into the same spatial coordinate system, forming a unified spatial reference architecture. This enables rapid matching of panoramic and point clouds with spatiotemporal semantic constraints, achieving spatial synchronization.

[0076] Compared with traditional methods that rely on post-processing for data alignment, this invention establishes a unified spatiotemporal reference during the data acquisition stage by using time-synchronous acquisition and motion state constraints at the sensor level. This avoids the accumulation of time inconsistencies and pose estimation errors in offline registration, improves the accuracy and efficiency of cross-modal data fusion, and enhances the real-time performance and robustness of the registration process in complex railway scenarios.

[0077] S2, data-level fusion of panoramic images and 3D laser point clouds based on spherical projection:

[0078] like Figure 4 As shown, based on the two-dimensional panoramic image and three-dimensional laser point cloud acquired synchronously by S1, a unified spherical coordinate reference frame is constructed to achieve consistent expression of directional and spatial information.

[0079] First, an equidistant cylindrical projection model is used to represent the pixel coordinates in the two-dimensional panoramic image. Mapped to spherical angular coordinates ( The preliminary construction of the panoramic sphere is complete, and the mapping relationship is as follows:

[0080] , ;

[0081] in: Indicates the horizontal direction angle; Indicates the vertical direction angle; This is a scale factor between angles and pixels, used to establish a linear correspondence between spherical angles and panoramic image coordinates. Through mapping operations, each pixel of the panoramic image corresponds to a unique spatial orientation on the panoramic sphere.

[0082] For 3D laser point clouds, Cartesian coordinates are used ( ) to spherical coordinates ( The conversion formula for inversely calculating spherical angle coordinates () ):

[0083] , , ;

[0084] in, Indicates along direction ( The spatial radius of ). When When taking a unit value, it is used to express spatial direction; when When the actual distance value measured from the 3D laser point cloud is taken, it represents the corresponding actual 3D spatial coordinates.

[0085] Based on the consistency of spherical angles, point cloud spatial points and corresponding pixel units are matched to achieve the binding of geometric information (X,Y,Z) and color information (R,G,B), complete data-level fusion, and save the binding relationship to the panoramic sphere.

[0086] S3, semantically guided panoramic and point cloud feature matching:

[0087] like Figure 5 As shown, firstly, a semantic segmentation network is used to extract visual features with clear semantic meanings from the panoramic image obtained from S1, such as railway track outlines, platform boundaries, catenary, building facades, and surrounding facilities, forming image semantic objects with semantic labels in the image modality. At the same time, stable geometric features such as planes, linear structures, and corners are extracted from the 3D laser point cloud obtained from S1, and spatial instance objects with hierarchical structures such as track surfaces and tower columns are further identified. Subsequently, the above image semantic objects and spatial instance objects are semantically aligned according to the consistency of semantic labels, constructing cross-modal semantic objects that represent the same scene entity in both modalities.

[0088] Subsequently, a semantically guided graph matching algorithm is constructed. The extracted cross-modal semantic objects (spatially aligned image semantic objects and spatial instance objects) are used as graph nodes in the image modality and point cloud modality, respectively. Each node contains spatial location, geometric features, and semantic labels, with spatial relationships (such as parallelism, perpendicularity, and connectivity) as edges. During graph matching, node matching calculates similarity based on the consistency of semantic labels and local geometric features, filtering candidate corresponding node pairs to quickly narrow the matching range. Relationship matching further compares the consistency of spatial relationships between nodes, verifying whether the edges between corresponding node pairs maintain the same geometric constraints (such as whether parallel relationships correspond and whether connection topologies are consistent). Through the semantically guided graph matching algorithm, accurate correspondences are established under spatial constraints, and feature matching is performed only within the scope of semantically consistent objects, thereby achieving the fusion and correspondence of cross-modal features. This constructs a railway scene feature model (capable of reflecting the spatial relationships between objects) containing semantic information and spatial structural relationships of each object within the scene. This significantly improves the accuracy and robustness of cross-modal feature matching in railway scenes under complex lighting and viewpoint changes.

[0089] S4 generates analytical railway VR panoramas through depth compensation and scale correction, such as... Figure 6 As shown, the specific operation is as follows:

[0090] To address the issue that panoramic images only possess relative depth information and lack true spatial scale, the relative depth obtained from the panoramic image obtained by S1 is estimated visually. The true spatial depth provided by the 3D laser point cloud obtained from S1 Construct a depth compensation model for railway VR panorama and optimize the solution:

[0091] ;

[0092] In the formula, For the fusion depth of the panoramic sphere, , The weighting coefficients are and satisfy the following conditions: + =1, used to balance geometric accuracy and visual continuity; in dense point cloud and structurally stable regions, geometric accuracy takes precedence. In areas where point clouds are sparse or missing, visual continuity is the primary consideration. .

[0093] The depth compensation model obtained after solving the problem is used to perform depth compensation on the panoramic sphere obtained in S2. Then, combined with the railway scene feature model obtained in S3, geometric constraint corrections are performed on key scene objects such as track surfaces, elevation structures, and towers in the panoramic sphere. VR technology is then used to construct a railway VR panorama with realistic spatial proportions and depth information. This railway VR panorama possesses both excellent immersive visual effects and high-precision spatial measurement and analysis capabilities.

Claims

1. A method for generating an analytical railway VR panorama fusing three-dimensional laser point clouds, characterized in that, Includes the following steps: S1, for the railway environment, perform coordinated acquisition of panoramic images and three-dimensional laser point clouds, wherein the coordination includes time synchronization and spatial synchronization; S2, map the pixel coordinates in the panoramic image to spherical angle coordinates to initially complete the construction of the panoramic sphere; for the three-dimensional laser point cloud, calculate its spherical angle coordinates; based on the consistency of spherical angles, match the spatial points of the three-dimensional laser point cloud with the pixel units of the panoramic image to realize the binding of geometric information and color information, and save the binding relationship to the panoramic sphere; S3. First, a semantic segmentation network is used to extract visual features with clear semantic meanings from the panoramic image obtained in S1, forming image semantic objects of image modalities with semantic labels. Simultaneously, stable geometric features are extracted from the 3D laser point cloud obtained in S1, and spatial instance objects of point cloud modalities with hierarchical structures are further identified. Then, the image semantic objects and spatial instance objects are semantically aligned to construct cross-modal semantic objects representing the same scene entity in different modalities. Finally, a semantically guided graph matching algorithm is used to perform feature matching on the features of the two modalities in the cross-modal semantic objects within the semantically consistent object range. A railway scene feature model containing the semantic information and spatial structural relationships of each object in the scene is constructed. S4. Using the relative depth obtained by visual estimation from the panoramic image obtained in S1 and the real spatial depth provided by the three-dimensional laser point cloud obtained in S1, a depth compensation model for the railway VR panorama is constructed and optimized. The depth compensation model obtained by optimization is used to perform depth compensation on the panoramic sphere obtained in S2. Then, the geometric constraint correction of key scene objects in the panoramic sphere is performed by combining the railway scene feature model, and then VR technology is used to construct an analytical railway VR panorama.

2. The analytical railway VR panorama generation method according to claim 1, characterized in that, In step S1: Multiple sensors are integrated on the collaborative acquisition platform for collaborative acquisition of panoramic images and 3D laser point clouds. The multiple sensors include a panoramic camera, a lidar, and an inertial measurement unit.

3. The analytical railway VR panorama generation method according to claim 2, characterized in that, The time synchronization method in step S1 is as follows: For the panoramic camera, LiDAR, and inertial measurement unit, a hardware-triggered multi-sensor time synchronization acquisition mechanism is constructed to achieve strict timestamp alignment; wherein, the multi-sensor time synchronization acquisition mechanism is achieved by building a low-level hard synchronization system based on an STM32 microcontroller and a TTL to USB module to realize the data acquisition time synchronization between the panoramic camera and the LiDAR.

4. The analytical railway VR panorama generation method according to claim 3, characterized in that: The STM32 microcontroller is set as the global master clock source. By writing high-frequency timing logic, periodic TTL pulse signals are sent to the industrial camera through the GPIO interface, forcing the camera to perform exposure acquisition at a specific physical moment. Simultaneously, the periodic TTL pulse signal is transmitted in real time to the host computer computing platform via a TTL-to-USB module and captured as a unified time reference tag, ensuring that the acquired panoramic image and the three-dimensional laser point cloud are synchronized at the microsecond level in the time dimension.

5. The analytical railway VR panorama generation method according to claim 4, characterized in that, The method for spatial synchronization in step S1 is as follows: Intrinsic parameter calibration of the panoramic camera was performed to obtain the imaging model and distortion parameters of the panoramic camera, ensuring the geometric consistency of the panoramic image. Based on this, the transformation formula between the camera coordinate system and the radar coordinate system established through joint calibration of the panoramic camera and LiDAR is as follows: ; In the formula, ( , , () is a point in the camera coordinate system, , , () represents a point in the corresponding lidar coordinate system. Let be a rotation matrix. It is a translation matrix; For panoramic images and 3D laser point clouds with the same timestamp, a rapid matching of multi-sensor data is achieved by using temporal semantic constraints and joint calibration methods, thus realizing spatial synchronization.

6. The analytical railway VR panorama generation method according to claim 5, characterized in that, The method for achieving fast matching of multi-sensor data using temporal semantic constraints and joint calibration is as follows: Acquire temporal semantic information, combine it with attitude angle, angular velocity and acceleration information provided by the inertial measurement unit, establish prior constraints for sensor motion, make initial estimates of the pose changes of multiple sensors and use them as initial values ​​for optimization; Then, corner features in the panoramic image and plane and edge features in the point cloud are extracted as constraints by a joint calibration algorithm. The algorithm quickly converges and solves the relative position and attitude relationship of multiple sensors. The rotation matrix and translation vector between the two are solved by a nonlinear optimization method, and the panoramic image and the three-dimensional laser point cloud are unified to the same spatial coordinate system. The temporal semantic information includes a unified timestamp and acquisition frequency control.

7. The analytical railway VR panorama generation method according to claim 6, characterized in that, In step S2: The pixel coordinates in the two-dimensional panoramic image are represented by an equidistant cylindrical projection model. Mapped to spherical angular coordinates ( The preliminary construction of the panoramic sphere is complete, and the mapping relationship is as follows: , ; in: Indicates the horizontal direction angle; Indicates the vertical direction angle; This is a scale factor between angles and pixels, used to establish a linear correspondence between spherical angles and panoramic image coordinates; Through mapping operations, each pixel of the panoramic image corresponds to a unique spatial orientation on the panoramic sphere; For 3D laser point clouds, Cartesian coordinates are used ( ) to spherical coordinates ( The conversion formula for inversely calculating spherical angle coordinates () ): , , ; in, Indicates along direction ( The spatial radius of ) when When taking a unit value, it is used to express spatial direction; when When the actual distance value measured from the 3D laser point cloud is taken, it represents the corresponding actual 3D spatial coordinates.

8. The analytical railway VR panorama generation method according to claim 7, characterized in that, In step S3: The visual features include the outline of the railway track, platform boundaries, overhead contact lines, building facades, and surrounding facilities; The geometric features include planes, linear structures, and corner points; The spatial instance objects include the track surface and the tower column; The semantically guided graph matching algorithm uses the image semantic objects and spatial instance objects in the cross-modal semantic objects obtained by semantic alignment as graph nodes of the image modality and point cloud modality, respectively. Each graph node contains spatial location, geometric features and semantic labels, and their spatial relationships are used as edges. In the graph matching process, node matching calculates similarity based on the semantic label consistency of nodes and local geometric features, filters candidate corresponding node pairs, and quickly narrows down the matching range; relation matching further compares the spatial relationship consistency between nodes and verifies whether the edges between corresponding node pairs maintain the same geometric constraints, thereby realizing feature matching within the semantically consistent object range and constructing a railway scene feature model that includes the semantic information and spatial structural relationship of each object in the scene.

9. The analytical railway VR panorama generation method according to claim 8, characterized in that, In step S4: The depth compensation model is as follows: ; In the formula, The fusion depth of the panoramic sphere; The relative depth is obtained by visual estimation using the panoramic image obtained from S1; Provides the true spatial depth for the 3D laser point cloud obtained by S1; , The weighting coefficients are and satisfy the following conditions: + =1, used to balance geometric accuracy and visual continuity; In regions with dense point clouds and stable structures, geometric accuracy is the primary consideration. In areas where point clouds are sparse or missing, visual continuity is the primary consideration. ; The key scene objects include the track surface, facade structure, and towers.

10. The analytical railway VR panorama generation method according to claim 2, characterized in that: The collaborative data acquisition platform includes handheld, quadruped robot-mounted, and UAV-borne data acquisition platforms, among which: For areas accessible by human intervention, a handheld data collection platform enables lightweight, rapid deployment, and highly mobile data collection. For areas with complex terrain or where stable observation by humans is difficult, a quadruped robot-mounted data acquisition platform is used to enhance autonomous movement and continuous data acquisition capabilities in complex environments. For large-scale scenes and areas with significant elevation differences, a drone-mounted data acquisition platform is used to acquire high-altitude overhead, side, and overall structural data.