A method for axis determination and axis compensation based on fusion of vision and laser

By employing a vision-laser fusion method for axle positioning and compensation, and using 2D cameras and laser sensors to collect data in parallel, the axle positioning and compensation processes are performed after matching and alignment. This solves the problem of continuity and high precision in axle positioning under complex conditions of long train formations, and achieves complete acquisition of the axle coordinates of the entire train.

CN122636729APending Publication Date: 2026-08-25NANJING KINGYOUNG INTELLIGENT SCI & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610886017.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-18
Publication Date
2026-08-25

AI Technical Summary

Technical Problem

Under complex operating conditions of long-formation trains, single vision or laser positioning methods are prone to positioning interruption or data loss due to interference from factors such as lighting, dirt, corrosion or obstruction. The lack of reasonable means to supplement data leads to the interruption of the entire train axle positioning operation.

Method used

A vision- and laser-fusion-based axle positioning and compensation method is adopted. Axle data is collected in parallel by 2D cameras and laser sensors to form a vision and laser positioning axle coordinate dataset. The axle coordinates are matched and aligned according to the carriage number, and axle positioning and compensation are performed using preset priority rules to generate a complete axle positioning dataset for the entire train carriage.

Benefits of technology

It achieves continuous, complete, and high-precision acquisition of the axle coordinates of the entire train, solving the problem of easy failure of positioning by a single sensor under complex working conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122636729A_ABST
    Figure CN122636729A_ABST
Patent Text Reader

Abstract

The application discloses a kind of based on vision and laser fusion's fixed axle method of supplementing axle, comprising: in the process that robot moves along train car side, 2D camera and laser sensor are triggered in parallel, vehicle axle related data are synchronously collected and visual vehicle axle coordinate data and laser vehicle axle coordinate data are generated, form visual positioning vehicle axle coordinate data set and laser positioning vehicle axle coordinate data set;Visual positioning vehicle axle coordinate data set and laser positioning vehicle axle coordinate data set are matched and aligned according to car number, form double-source data matrix aligned according to car;Double-source data matrix is handled according to the preset priority rule section by section car axle, generates the complete vehicle axle positioning data set of whole train car.This application establishes three-level fixed axle mechanism of vision priority, laser supplement, interpolation complement, solves the positioning failure of single sensor in long marshalling train complex working condition, data discontinuity problem, realizes the high-precision acquisition of whole train axle coordinate.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent detection technology for rail transit, and in particular to a fixed-axis compensation method based on vision and laser fusion. Background Technology

[0002] In the rail transit operation and maintenance system, the train axle, as the core running gear that bears the entire load of the vehicle and transmits rotational motion to the wheels, requires regular inspection of its structural integrity and operational status as a crucial link in ensuring train safety. Whether it's routine maintenance such as ultrasonic flaw detection and wheelset dimension measurement, or safety inspections such as axle temperature monitoring and bearing condition assessment, accurate acquisition of the axle's spatial position is a prerequisite for all operations. Only when the axle's spatial coordinates are accurately determined can automated testing equipment move precisely to the target position to perform subsequent specialized testing tasks. The accuracy of axle positioning directly impacts the reliability of the entire train maintenance process.

[0003] Currently, there are two main types of axle positioning methods in the rail transit field:

[0004] One approach is a visual positioning method based on optical images. This method relies on fixed or vehicle-mounted industrial cameras to capture visible light images of the axle area from the side of the train. Image processing algorithms are then used to identify the axle's characteristic shape and calculate its spatial position. However, the actual environment of railway stations and maintenance workshops is harsh, with issues such as varying lighting, axle surface stains and corrosion, and component obstructions, which can easily lead to failures in feature extraction or template matching.

[0005] The second method is a geometric positioning method based on laser ranging. This method uses lidar to acquire 3D point cloud data of the axle and surrounding structures, and then uses point cloud processing algorithms to extract the cylindrical surface geometric features of the axle, calculate the axis direction, and the coordinates of the axle center. However, in practical applications, when the axle surface has severe rust pits, wear, or foreign object obstruction, the point cloud data can become abnormally scattered or incomplete, leading to deviations in the geometric feature extraction algorithm or complete failure due to the inability to obtain complete axle contour information.

[0006] Long-formation trains typically consist of dozens or even hundreds of carriages. When stopping at stations or maintenance depots, complex conditions arise, such as variable stopping positions and flexible changes in carriage spacing. This necessitates that axle positioning methods not only achieve high-precision positioning at single points but also possess continuous and complete positioning capabilities across the entire train. However, single-sensor or laser positioning methods are prone to positioning interruptions or data loss due to interference from factors such as lighting, dirt, corrosion, or obstructions. When a single sensor fails to locate a carriage, the lack of reasonable data replenishment methods based on prior knowledge and mathematical approaches can lead to the interruption of axle positioning operations for the entire train.

[0007] In summary, under current technological conditions, how to overcome the inherent limitations of single-sensor positioning methods and achieve continuous, complete, and high-precision output of axle positioning data for the entire train under complex working conditions where the stopping positions of long trains are not fixed and the spacing between carriages changes flexibly is a key technical problem that urgently needs to be solved in the field of intelligent detection of rail transit. Summary of the Invention

[0008] The present invention aims to solve the above problems by providing a fixed-axis axle compensation method based on vision and laser fusion, which solves the problems of continuity, integrity and high precision in axle positioning under complex working conditions.

[0009] The present invention solves the aforementioned problem by employing the following technical solution: a method for fixing and compensating for axis based on vision and laser fusion, comprising the following steps:

[0010] As the robot moves along the side of the train, the 2D camera and the laser sensor are triggered in parallel to collect axle-related data and generate visual axle coordinate data and laser axle coordinate data, forming visual positioning axle coordinate dataset and laser positioning axle coordinate dataset, respectively; the 2D camera and the laser sensor are mounted on the robot.

[0011] The visual positioning axle coordinate dataset and the laser positioning axle coordinate dataset are matched and aligned according to the carriage number to form a dual-source data matrix aligned by carriage.

[0012] The dual-source data matrix is ​​processed car by car according to a preset priority rule to generate a complete axle positioning dataset for the entire train car.

[0013] The formation of the visual positioning axle coordinate dataset and the laser positioning axle coordinate dataset includes a first branch and a second branch:

[0014] The first branch works through the 2D camera, triggers the 2D camera to capture images of the train side at a preset frequency, identifies the axle position based on the train side images and calculates the axle coordinates in the global coordinate system, generates the visual axle coordinate data, and collects it to form the visual positioning axle coordinate dataset.

[0015] The second branch operates through the laser sensor, driving the laser sensor to continuously scan the side of the train to obtain three-dimensional point cloud data. Based on the three-dimensional point cloud data, it extracts the geometric features of the axle and calculates the coordinates of the axle in the global coordinate system, generating the laser axle coordinate data, which is then collected to form the laser positioning axle coordinate dataset.

[0016] The first branch specifically includes:

[0017] The 2D camera receives the hard trigger signal generated by the microcontroller and the robot's real-time global coordinate information, captures images of the train's side, and binds the train's side images with the real-time global coordinate information at the time of capture to form image data frames with spatial location tags.

[0018] For the image data frames with spatial location labels, a target detection method is used to extract multi-scale features of the image, which are then matched with pre-trained axle target features to filter out image regions containing axles.

[0019] For the selected image regions containing axles, calculate the pixel coordinates of the axle center in the image coordinate system, and convert them into X-direction and Y-direction deviation values ​​of the axle relative to the 2D camera using camera intrinsic parameters.

[0020] Using the real-time global coordinate information bound in the image data frame with spatial location labels, the X-direction deviation value and Y-direction deviation value are converted into the absolute coordinates of the axle in the global coordinate system, generating visual axle coordinate data with timestamps.

[0021] The second branch specifically includes:

[0022] The laser sensor is driven to continuously scan the side surface of the train to obtain a dense three-dimensional point cloud data stream. Each data point in the three-dimensional point cloud data stream contains coordinate information and reflection intensity information in the global coordinate system.

[0023] Perform point cloud filtering on the three-dimensional point cloud data stream to remove noise points and outliers, and obtain denoised point cloud data;

[0024] Clustering and segmentation processing is performed on the denoised point cloud data to separate the point cloud regions of different objects and obtain multiple point cloud clusters.

[0025] Geometric features are extracted from each of the multiple point cloud clusters, and the extracted geometric features are matched with a preset axle laser template, wherein the axle laser template defines the geometric parameters and spatial position constraints of the axle target;

[0026] When the geometric features of a certain point cloud cluster and the matching deviation of the axle laser template are within a preset threshold range, it is determined that the point cloud cluster corresponds to an axle. The centroid coordinates of the point cloud cluster are calculated and used as the coordinates of the axle in the global coordinate system, generating laser axle coordinate data with timestamps.

[0027] Forming the dual-source data matrix aligned by carriage includes:

[0028] Read the visual positioning axle coordinate dataset and the laser positioning axle coordinate dataset;

[0029] The timestamp information of each data record in the two datasets is extracted. Based on the timestamp sequence of the visual positioning axle coordinate dataset, time interpolation processing is performed on the laser positioning axle coordinate dataset to achieve alignment of the two datasets in the time dimension.

[0030] Confirm that the axle coordinates in both datasets have been transformed to the same global coordinate system, thus completing the spatial benchmark unification;

[0031] Based on prior structural knowledge of long-formation trains, the coordinate data of the axles contained in the same carriage in the visual positioning axle coordinate dataset and the laser positioning axle coordinate dataset are associated with the corresponding carriage number. A structured data table is constructed with the carriage number as the row index and the visual axle coordinate data and the laser axle coordinate data as the two data fields, forming the carriage-aligned dual-source data matrix.

[0032] The generated complete axle positioning dataset for all train carriages includes:

[0033] According to the carriage number order, traverse the data rows corresponding to each carriage in the dual-source data matrix aligned by carriage;

[0034] For the currently traversed carriage, check whether the visual axle coordinate data corresponding to the carriage is valid;

[0035] If the visual axle coordinate data is valid, then the visual axle coordinate data is extracted as the final axle positioning data of the carriage and written into the record position corresponding to the carriage in the final axle positioning data sequence.

[0036] The preset priority rules also include:

[0037] If the visual axle coordinate data corresponding to the current carriage is invalid, then further check whether the laser axle coordinate data corresponding to the carriage is valid;

[0038] If the laser axle coordinate data is valid, then the laser axle coordinate data is extracted as the final axle positioning data of the carriage and written into the record position corresponding to the carriage in the final axle positioning data sequence.

[0039] The preset priority rules also include: if both the visual axle coordinate data and the laser axle coordinate data corresponding to the current carriage are invalid, then axle compensation processing is triggered; the axle compensation processing includes the following steps:

[0040] Read the pre-configured carriage length parameters and obtain the final axle positioning data of the carriages that are adjacent to the current carriage in front and behind and have valid final axle positioning data;

[0041] Based on the carriage length parameter and the final axle positioning data of the adjacent carriages, a linear interpolation algorithm is used to calculate the estimated axle coordinates of the current carriage.

[0042] The estimated axle coordinates are used as the final axle positioning data for the carriage, and written into the corresponding record position of the carriage in the final axle positioning data sequence.

[0043] The robot is equipped with a navigation module. As the robot moves along the side of the train, the navigation module continuously acquires the robot's real-time global coordinate information and performs trajectory correction based on the real-time global coordinate information to maintain a preset detection distance between the robot and the side of the train.

[0044] The target detection method is based on the YOLOv8 network.

[0045] The beneficial effect of this invention is that by establishing a three-level fixed-axis compensation mechanism of vision priority, laser supplementation, and interpolation completion, it effectively solves the problems of easy failure of positioning and discontinuous data of single sensor under the complex working conditions of long train formations, and realizes continuous, complete and high-precision acquisition of the axle coordinates of the entire train. Attached Figure Description

[0046] Figure 1 This is a flowchart of the overall solution of the present invention.

[0047] Figure 2 This is a flowchart of the first branch of the present invention.

[0048] Figure 3 This is a flowchart of the second branch of the present invention.

[0049] Figure 4 This is a flowchart of the process of forming a dual-source data matrix aligned by carriage according to the present invention.

[0050] Figure 5 This is a flowchart illustrating how the present invention generates a complete axle positioning dataset for all train carriages. Detailed Implementation

[0051] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0052] The implementation environment of this invention includes a robot, which can be an AGV intelligent inspection robot. The robot is equipped with a navigation module, a 2D camera and a laser sensor. The navigation module is used to obtain the robot's real-time coordinates in the global coordinate system. Different navigation sensors output different real-time coordinates, such as magnetic navigation coordinates, QR code coordinates, laser SLAM coordinates, etc.

[0053] like Figure 1 As shown, a data processing flow for a fixed-axis compensation method based on vision and laser fusion includes the following steps:

[0054] S1: During the robot's movement along the side of the train, the 2D camera and laser sensor are triggered in parallel to collect axle-related data and generate visual axle coordinate data and laser axle coordinate data, respectively forming a visual positioning axle coordinate dataset and a laser positioning axle coordinate dataset.

[0055] S2: Match and align the visual positioning axle coordinate dataset and the laser positioning axle coordinate dataset according to the carriage number to form a dual-source data matrix aligned by carriage.

[0056] S3: Perform axle positioning and axle compensation processing on each carriage according to the preset priority rules to generate a complete axle positioning dataset for the entire train carriage.

[0057] This embodiment describes in detail the data processing flow of a fixed-axis compensation method based on vision and laser fusion, specifically including:

[0058] S1: During the robot's movement along the side of the train, the 2D camera and laser sensor are triggered in parallel to collect axle-related data and generate visual axle coordinate data and laser axle coordinate data, respectively forming a visual positioning axle coordinate dataset and a laser positioning axle coordinate dataset.

[0059] In this scheme, the robot first autonomously travels to the starting position in the long train inspection area according to a preset navigation path, completing initial pose calibration so that the detection field of view of its side-mounted 2D camera and laser sensor faces the side of the train. According to the preferred configuration of this embodiment, the detection distance between the robot and the side of the train is set to 0.5m to 0.7m. This distance range takes into account both the complete coverage of the axle targets by the sensor's field of view and the safety collision avoidance requirements during equipment operation. After positioning, the robot's navigation module is activated, driving the robot body to move at a preset speed along the side of the train at a constant speed. For example, this preset moving speed can be set to 1m / s. This speed parameter has a co-design relationship with the subsequent trigger acquisition frequency of the 2D camera, aiming to ensure that each axle can be completely included in the image acquisition range when the robot is moving at high speed, avoiding missed detections. During the movement, the navigation module continuously acquires the robot's real-time global coordinate information and performs closed-loop trajectory correction based on this real-time global coordinate information, keeping the detection distance between the robot and the side of the train stable.

[0060] As the robot moves, the parallel data acquisition and processing flow of the two sensors is simultaneously triggered. This parallel working mechanism is the core architecture for achieving redundant perception in this solution, consisting of two independent and synchronously running branches: the first branch and the second branch.

[0061] S1.1: The first branch works through a 2D camera, which triggers the 2D camera to capture images of the train side at a preset frequency. Based on the images of the train side, the axle position is identified and the coordinates of the axle in the global coordinate system are calculated to generate visual axle coordinate data, which is then collected to form a visual positioning axle coordinate dataset.

[0062] In the first branch, the microcontroller continuously generates hard-triggered pulse signals at a preset frequency. These signals are output to the 2D camera to control its image acquisition timing. According to a further improvement in this embodiment, the frequency of the hard-triggered signal is preferably set to 20Hz, matching the robot's moving speed of 1m / s, ensuring a sufficiently dense image sampling interval along the train's direction of travel, achieving complete image capture for each axle. Simultaneously, the navigation module sends the robot's real-time global coordinate information to the 2D camera. This information specifically includes the robot's X-direction position value, Y-direction position value, and yaw angle attitude data in the global coordinate system.

[0063] Specifically, after receiving the hard trigger signal and real-time global coordinate information, the 2D camera performs a complete image acquisition and processing cycle, such as... Figure 2 As shown, this cycle further includes the following processing steps.

[0064] S1.1.1: The 2D camera receives the hard trigger signal generated by the microcontroller and the robot's real-time global coordinate information, captures images of the train's side, and binds the train's side images with the real-time global coordinate information at the time of capture to form image data frames with spatial location labels.

[0065] In this processing step, the 2D camera responds to the rising edge of the hard trigger signal, performs an exposure operation, and captures a high-resolution image of the train's side within the current field of view. After capture, the image processing unit inside the 2D camera binds and encapsulates the captured train side image with the real-time global coordinate information received simultaneously. Specifically, it writes the global coordinate value and attitude angle information corresponding to the capture moment into the metadata field of the image file, thus forming an image data frame with a spatial location tag. This spatial location tag records the robot's precise pose in the global coordinate system of the detection site at the moment of image acquisition, providing a spatial reference for subsequent axle global coordinate conversion.

[0066] S1.1.2: For image data frames with spatial location labels, use a YOLO-based target detection method to extract multi-scale features of the image, match them with pre-trained axle target features, and filter out image regions containing axles.

[0067] In this processing stage, the airborne data processing unit performs axle recognition processing on image data frames with spatial location labels. In a preferred embodiment, the axle recognition processing adopts a target detection method based on YOLOv8. First, the data processing unit loads a pre-trained YOLOv8 model, which has completed transfer learning on a dataset containing labeled samples of axle boxes and surrounding bolts. Its backbone network adopts a CSPDarknet structure, the neck network adopts a PAN-FPN structure, and the detection head is a decoupled bounding box and category prediction branch. The data processing unit receives the image data frames to be detected in real time, uniformly scales the image size to the model's preset input size (e.g., 640×640 pixels), and performs normalization processing. Subsequently, the processed image tensor is fed into the YOLOv8 model for forward inference: the model extracts multi-scale feature maps through the backbone network, fuses semantic information of different scales through the neck network, and finally outputs multiple candidate detection boxes by the detection head. Each detection box contains the category confidence of the axle box or bolt, and the bounding box coordinates (center point x, y, width w, height h). For the original output, the data processing unit uses the Non-Maximum Suppression (NMS) algorithm, setting an IoU threshold (e.g., 0.5) and a confidence threshold (e.g., 0.6) to remove redundant and low-confidence detection boxes, obtaining the finally identified axle box target region (denoted as Box_axle) and four bolt target regions (denoted as Box_bolt1~Box_bolt4). When the number of detected axle boxes is not 1 or the number of bolts is not 4, the current image frame is considered incomplete, and the frame is skipped or the process waits for the next frame. Provided the detection is valid, the data processing unit calculates the center point coordinates of the axle box detection boxes: Let the upper left corner coordinates of the axle box detection box be (x1, y1) and the lower right corner coordinates be (x2, y2), then the axle box center C_axle = ((x1+x2) / 2, (y1+y2) / 2). Similarly, the center point coordinates of the four bolt detection boxes are calculated, denoted as C_bolt1~C_bolt4. To comprehensively calculate the axle box center, the data processing unit employs a weighted average or geometric method: It calculates the average position of all four bolt center points, C_bolts_avg = ((∑x_i) / 4, (∑y_i) / 4), then merges the axle box center C_axle with the average bolt center C_bolts_avg, and obtains the final axle box center C_final = 0.6*C_axle + 0.4*C_bolts_avg according to preset weights (e.g., axle box weight 0.6, bolt weight 0.4). If bolt identification is incomplete in a frame but axle box identification is valid, only the axle box center can be used as a temporary output. This comprehensive calculation method utilizes the global position of the axle box itself and the local fine structure of the bolts, improving the robustness and accuracy of center positioning.

[0068] S1.1.3: For the selected image regions containing axles, calculate the pixel coordinates of the axle center in the image coordinate system, and convert them into the X-direction deviation value and Y-direction deviation value of the axle relative to the 2D camera by combining the camera intrinsic parameters.

[0069] In this processing step, the data processing unit precisely locates the selected image regions containing the axle. Specifically, within the image region defined by the effective matching point pairs, the data processing unit further extracts the contour boundary of the axle through edge detection or morphological analysis, and determines the pixel coordinates of the axle center in the image coordinate system based on contour fitting. After obtaining the pixel coordinates, the data processing unit calls the camera intrinsic parameter matrix obtained from the factory calibration or online calibration of the 2D camera. This intrinsic parameter matrix includes the camera focal length parameter and principal point coordinate parameter, converting the pixel coordinates of the axle center into the physical deviation value of the axle relative to the optical center of the 2D camera, obtaining the deviation value in the horizontal X direction and the deviation value in the vertical Y direction.

[0070] S1.1.4: Using the real-time global coordinate information bound in the image data frame with spatial location labels, the X-direction deviation value and Y-direction deviation value are converted into the absolute coordinates of the axle in the global coordinate system, generating visual axle coordinate data with timestamps.

[0071] In this processing step, the data processing unit extracts the bound real-time global coordinate information from the metadata field of the currently processed image data frame to obtain the robot's pose parameters in the global coordinate system at the time of capture. Subsequently, based on the robot's geodetic coordinates and yaw angle contained in these pose parameters, the data processing unit constructs a homogeneous coordinate transformation matrix from the camera's local coordinate system to the global coordinate system. The X-direction and Y-direction deviation values ​​are then converted into absolute coordinate values ​​of the axle in the global coordinate system using this transformation matrix. These absolute coordinate values ​​constitute a visual axle coordinate data record. The data processing unit appends a current timestamp to this record, generating timestamped visual axle coordinate data. This complete process of image acquisition, feature matching, deviation calculation, and global coordinate conversion is executed cyclically, generating a timestamped visual axle coordinate data record with each execution. All data records are then aggregated to form a visual positioning axle coordinate dataset.

[0072] S1.2: The second branch works through a laser sensor, which continuously scans the side of the train to obtain three-dimensional point cloud data. Based on the three-dimensional point cloud data, it extracts the geometric features of the axle and calculates the coordinates of the axle in the global coordinate system, generating laser axle coordinate data, which is then collected to form a laser positioning axle coordinate dataset.

[0073] In the second branch, a laser sensor continuously scans the train's side surface during robot movement. This laser sensor can be a line-scan lidar or an area-array lidar, emitting laser detection signals to the train's side at a high sampling frequency and receiving reflected echoes. By measuring the laser's time of flight or phase difference, the spatial distance to the target point is calculated, thereby continuously acquiring a three-dimensional point cloud data stream of the train's side surface.

[0074] Specifically, each frame scan by the laser sensor generates a set of point cloud data containing a large number of data points, such as... Figure 3 As shown, the processing flow of this branch further includes the following steps.

[0075] S1.2.1: Drive the laser sensor to continuously scan the side surface of the train to obtain a dense three-dimensional point cloud data stream. Each data point in the three-dimensional point cloud data stream contains coordinate information and reflection intensity information in the global coordinate system.

[0076] In this processing stage, the laser sensor performs reciprocating or spiral scanning on the train's side surface at a preset scanning frequency and angular resolution to acquire a dense three-dimensional point cloud data stream. Each data point in this three-dimensional point cloud data stream contains at least three coordinate dimensions and one intensity dimension: the X, Y, and Z coordinate values ​​in the global coordinate system, and the reflection intensity information representing the reflectivity of the target surface. The global coordinate values ​​are obtained through coordinate transformation and fusion calculation based on the raw distance and angle information measured by the laser sensor, combined with the laser sensor's external parameter calibration results on the robot, and the robot's real-time pose information provided by the navigation module.

[0077] S1.2.2: Perform point cloud filtering on the 3D point cloud data stream to remove noise points and outliers, and obtain denoised point cloud data.

[0078] In this processing stage, the data processing unit performs point cloud preprocessing operations on the acquired 3D point cloud data stream. In a preferred embodiment, noise removal is achieved using pass-through filtering and statistical filtering algorithms. Pass-through filtering compares the X, Y, and Z values ​​of each data point in the point cloud with the threshold values ​​set by prior knowledge of the external environment. If the values ​​exceed the threshold, the point cloud is deemed invalid and removed. The statistical filtering algorithm calculates the average distance of each data point to its K nearest neighbors. Assuming that the average distance between all points and their neighbors follows a Gaussian distribution, the mean and standard deviation of the global average distance are calculated. When the average distance of a data point exceeds the range of the mean plus N times the standard deviation, the data point is identified as a noise point or an outlier and removed from the point cloud data. After this filtering process, isolated noise points and measurement anomalies in the original point cloud are effectively removed, resulting in denoised point cloud data with more uniform density and clearer structure.

[0079] S1.2.3: Perform clustering and segmentation processing on the denoised point cloud data to separate the point cloud regions of different objects and obtain multiple point cloud clusters.

[0080] In this processing stage, the data processing unit performs target segmentation on the denoised point cloud data. In a preferred embodiment, a clustering segmentation algorithm based on Euclidean distance is used. This algorithm uses a preset spatial distance threshold as the criterion to perform a neighborhood search for each data point in the denoised point cloud data: starting from any unclassified point, it searches for nearest neighbors within a radius of the distance threshold, assigning the found nearest neighbors to the same cluster, and then using the newly assigned point as the seed point to continue the search outward until the cluster can no longer include any new nearest neighbors. After extracting a cluster, the above process is repeated for the remaining unclassified points until all data points are assigned to their corresponding clusters. Through this clustering segmentation process, the originally continuous point cloud data is divided into multiple independent point cloud clusters, each corresponding to an independent object or object component in the scene, such as a wheel, axle, bogie, or vehicle sidewall.

[0081] S1.2.4: Extract geometric features from multiple point cloud clusters one by one, and match the extracted geometric features with the preset axle laser template, wherein the axle laser template defines the geometric parameters and spatial position constraints of the axle target.

[0082] In this processing stage, the data processing unit performs geometric feature extraction and template matching operations on each of the multiple point cloud clusters obtained after clustering and segmentation. For each point cloud cluster, the data processing unit extracts its geometric feature parameters, which include, but are not limited to: the three-dimensional coordinates of the centroid of the point cloud cluster, the dimensions (length, width, and height) of the minimum circumscribed cuboid of the point cloud cluster, the direction vector of the principal axis of the point cloud cluster, and the radius value and axis direction of the cylinder obtained after fitting the point cloud cluster to a cylinder. Subsequently, the data processing unit compares and matches the extracted geometric feature parameters with a preset axle laser template. This axle laser template defines the range of geometric parameters and spatial position constraints that a standard axle target should meet: the range of geometric parameters includes the typical cylinder diameter and length range of the axle, and the spatial position constraints include that the axle axis direction should be approximately perpendicular to the track extension direction, and the height of the axle center relative to the track surface should be within a reasonable range.

[0083] S1.2.5: When the geometric feature of a certain point cloud cluster and the matching deviation of the axle laser template are within a preset threshold range, it is determined that the point cloud cluster corresponds to an axle, the centroid coordinates of the point cloud cluster are calculated as the coordinates of the axle in the global coordinate system, and laser axle coordinate data with timestamps are generated.

[0084] In this processing step, the data processing unit makes a judgment based on the geometric feature matching results from the previous step. When the deviation values ​​of the geometric feature parameters of a point cloud cluster and the corresponding parameters in the axle laser template are all within a preset threshold range, the point cloud cluster is determined to correspond to a real train axle. For example, when the absolute deviation between the cylinder fitting radius value of the point cloud cluster and the axle radius value defined by the template is less than the preset radius tolerance, and the angular deviation between the axial direction of the point cloud cluster and the directional constraint defined by the template is less than the preset angular tolerance, the matching is considered successful. After successful matching, the data processing unit calculates the three-dimensional coordinates of the centroid of the point cloud cluster, determines the centroid coordinates as the coordinate values ​​of the axle in the global coordinate system, and generates a laser axle coordinate data record. The current timestamp information is appended to this record to generate laser axle coordinate data with a timestamp. The complete process of point cloud acquisition, filtering and denoising, clustering and segmentation, geometric feature extraction and template matching is continuously executed. Each successful identification of an axle generates a laser axle coordinate data with a timestamp. All data records are collected to form a laser positioning axle coordinate dataset.

[0085] S2: Match and align the visual positioning axle coordinate dataset and the laser positioning axle coordinate dataset according to the carriage number to form a dual-source data matrix aligned by carriage.

[0086] When the robot travels along the side of the train to the preset detection endpoint, the navigation module sends a positioning signal to the robot's control unit. The control unit then generates a stop command, driving the robot to stop moving. Simultaneously, it sends stop acquisition commands to the 2D camera and laser sensor, causing both sensors to synchronously terminate data acquisition and coordinate calculation. At this point, the acquisition of both the visual positioning axle coordinate dataset and the laser positioning axle coordinate dataset is complete, and both datasets are fully stored in the robot's local storage unit.

[0087] Furthermore, the control unit initiates a coordinate data comparison and alignment process. The purpose of this process is to unify the axle coordinate data acquired independently by the two sensors at different sampling times and using different operating mechanisms, under the same spatiotemporal reference, and then structurally merge them according to the vehicle compartment to which the axle belongs, such as... Figure 4 As shown, the specific processing steps include the following.

[0088] S2.1: Read the visual positioning axle coordinate dataset and the laser positioning axle coordinate dataset.

[0089] In this processing step, the control unit reads all records from the visual positioning axle coordinate dataset and the laser positioning axle coordinate dataset from the local storage unit. Each visual positioning axle coordinate data record includes at least: the X, Y, and Z coordinate values ​​of the axle in the global coordinate system, the data acquisition timestamp, and the data source identifier. Each laser positioning axle coordinate data record includes at least: the X, Y, and Z coordinate values ​​of the axle in the global coordinate system, the data acquisition timestamp, and the data source identifier. The control unit performs integrity verification on the read data records, checking the continuity of the timestamps and the validity of the coordinate values. If abnormal records are found, they are marked to ensure that subsequent processing steps are based on a reliable data foundation.

[0090] S2.2: Extract the timestamp information of each data record in the two datasets. Using the timestamp sequence of the visual positioning axle coordinate dataset as a reference, perform time interpolation processing on the laser positioning axle coordinate dataset to achieve alignment of the two datasets in the time dimension.

[0091] In this processing stage, the control unit extracts the timestamp information of each data record in the visual positioning axle coordinate dataset to form a visual reference time series; simultaneously, it extracts the timestamp information of each data record in the laser positioning axle coordinate dataset to form a laser time series. Since the 2D camera in the first branch is triggered to acquire data at a fixed frequency of 20Hz, the visual reference time series has a uniform time sampling interval. However, the laser sensor in the second branch operates in continuous scanning mode, and the time interval of its laser time series is determined by the scanning frequency of the laser sensor, which may not completely coincide with the sampling time points of the visual reference time series. To achieve a one-to-one correspondence between the two datasets in the time dimension, the control unit uses the visual reference time series as the reference axis and performs time interpolation processing on the laser time series: for each time sampling point in the visual reference time series, it finds the two data records in the laser time series with the smallest absolute time difference from that sampling point. Based on the coordinate values ​​and time values ​​of these two data records, a linear interpolation method is used to calculate the estimated laser axle coordinate value corresponding to that visual sampling moment. After this time interpolation process, each record in the visual positioning axle coordinate dataset obtains a time-aligned laser axle coordinate corresponding value, and the two datasets achieve alignment and matching in the time dimension.

[0092] S2.3: Confirm that the axle coordinates in both datasets have been transformed to the same global coordinate system, thus completing the spatial reference unification.

[0093] In this processing step, the control unit performs a spatial reference consistency check on all axle coordinate data in both datasets. Specifically, the control unit verifies whether the axle coordinates in the visual positioning axle coordinate dataset and the laser positioning axle coordinate dataset are represented in the same global coordinate system. This global coordinate system typically uses a preset reference control point of the detection site as the origin, the track extension direction as the X-axis, and the direction perpendicular to the track plane upwards as the Z-axis. If the detection finds a coordinate system deviation between the two datasets, the data is unified to the same global coordinate system through coordinate transformation based on the sensor installation extrinsic parameter calibration data and the robot's real-time pose data.

[0094] S2.4: Based on the prior structural knowledge of long-formation trains, the coordinate data of the axles contained in the same carriage in the visual positioning axle coordinate dataset and the laser positioning axle coordinate dataset are associated with the corresponding carriage number. A structured data table is constructed with the carriage number as the row index and the visual axle coordinate data and the laser axle coordinate data as the two data fields, forming a dual-source data matrix aligned by carriage.

[0095] In this processing stage, the control unit, based on prior structural knowledge of long-formation trains, performs car assignment on the axle coordinate data after time alignment and spatial standardization. This prior structural knowledge specifically includes the standard car length parameters for this train model, the number of axles in a single car, the standard wheelbase parameters between axles, and the typical range of connection gaps between adjacent cars. For example, the control unit can sort all axle coordinate points according to the projection values ​​of the axle coordinates along the track direction (X-axis direction), and set a sliding matching window based on the standard car length parameters. Axles that are spatially continuous and whose spacing conforms to the wheelbase characteristics are grouped into the same car, and the axle groups of adjacent cars should exhibit spacing characteristics consistent with the car connection gaps. Through this assignment operation, each detected axle is assigned a unique car number. Based on this, the control unit constructs a structured data table: using the carriage number as the row index, two data fields are established for each carriage. The first data field stores the visual axle coordinate data corresponding to that carriage, and the second data field stores the laser axle coordinate data corresponding to that carriage. If a carriage has no valid axle coordinate data under a certain sensor channel, the corresponding data field is set to null. This structured data table is a dual-source data matrix aligned to the carriages.

[0096] S3: For the dual-source data matrix aligned by carriage, perform axle positioning and axle compensation processing on each carriage according to the preset priority rules to generate a complete axle positioning dataset for the entire train carriage.

[0097] After obtaining the dual-source data matrix aligned to the carriages, the robot's control unit initiates the axle positioning and compensation process. The core task of this process is to select the optimal value from the available axle coordinate data for each carriage according to a preset data source priority rule, and to use mathematical interpolation to fill in missing data in the special case where all available data sources fail, thereby ensuring the completeness and continuity of the final output axle positioning data for the entire train. Figure 5 As shown, the processing flow specifically includes the following steps.

[0098] S3.1: Following the order of the carriage numbers, sequentially traverse the data rows corresponding to each carriage in the dual-source data matrix aligned by carriage.

[0099] In this processing stage, the control unit reads each row of data from the dual-source data matrix aligned to the carriage number, starting from the first carriage and proceeding sequentially until the last carriage is reached. For each carriage, the control unit obtains the corresponding visual axle coordinate data domain and laser axle coordinate data domain from the matrix, which serve as inputs for subsequent hierarchical decision processing.

[0100] S3.2: For the currently traversed carriage, check whether the visual axle coordinate data corresponding to the carriage is valid.

[0101] In this processing step, the control unit performs a validity check on the visual axle coordinate data of the currently traversed carriages. Specific criteria for validity check include: the value of the data field is not empty; the coordinate value is within a preset valid coordinate range based on the detection area and train geometry; and the timestamp of the data record has reasonable temporal continuity with the timestamps of valid data records in adjacent carriages. If all the above criteria are met, the visual axle coordinate data of that carriage is considered valid.

[0102] S3.3: If the visual axle coordinate data is valid, extract the visual axle coordinate data as the final axle positioning data of the carriage and write it into the record position corresponding to the carriage in the final axle positioning data sequence.

[0103] In this processing step, when the previous step determines that the visual axle coordinate data is valid, the control unit executes the first-priority axle positioning operation. The control unit extracts the complete coordinate values ​​of the visual axle coordinate data corresponding to the carriage from the dual-source data matrix, determines it as the final axle positioning data for that carriage, and writes the coordinate values ​​into the record position corresponding to the carriage number in the final axle positioning data sequence. At the same time, the data source is marked as visual positioning in this record. The visual positioning method has a better ability to capture detailed features of the axle than laser sensors, and therefore it is set as the highest priority data source.

[0104] S3.4: If the visual axle coordinate data corresponding to the current carriage is invalid, then further check whether the laser axle coordinate data corresponding to the carriage is valid.

[0105] In this processing step, if the visual axle coordinate data of the current carriage is determined to be invalid, the control unit enters the second priority determination process. The control unit performs a validity determination on the laser axle coordinate data corresponding to the carriage. The determination criteria are similar to those for visual data, including that the data field is not empty, the coordinate values ​​are within a preset valid range, and the timestamps have reasonable continuity.

[0106] S3.5: If the laser axle coordinate data is valid, extract the laser axle coordinate data as the final axle positioning data of the carriage and write it into the record position corresponding to the carriage in the final axle positioning data sequence.

[0107] In this processing step, when the previous step determines that the laser axle coordinate data is valid, the control unit executes a second-priority axle compensation operation. The control unit extracts the complete coordinate values ​​of the laser axle coordinate data corresponding to the carriage from the dual-source data matrix, determines it as the final axle positioning data for the carriage, and writes the coordinate values ​​into the record position corresponding to the carriage number in the final axle positioning data sequence. At the same time, the data source is marked as laser positioning in the record. In an optional embodiment, the laser sensor achieves positioning through point cloud geometric feature extraction. Its working performance is less affected by factors such as ambient lighting conditions, carriage surface stains, and surface occlusion. Even when the 2D camera fails to provide visual positioning due to insufficient lighting or stain occlusion, the laser sensor can still provide reliable positioning data supplementation, and therefore it is set as the second-priority data source.

[0108] S3.6: If both the visual axle coordinate data and the laser axle coordinate data corresponding to the current carriage are invalid, then the axle compensation process is triggered. The axle compensation process is as follows: read the pre-configured carriage length parameter, obtain the final axle positioning data of the carriages that are adjacent to the current carriage before and after and have valid final axle positioning data; based on the carriage length parameter and the obtained final axle positioning data of the adjacent carriages before and after, use a linear interpolation algorithm to calculate the estimated axle coordinate value of the current carriage; use the estimated axle coordinate value as the final axle positioning data of the carriage, and write it into the record position corresponding to the carriage in the final axle positioning data sequence.

[0109] In this processing step, if both the visual axle coordinate data and the laser axle coordinate data of the current carriage are determined to be invalid, meaning that neither sensor channel has successfully located the axle at the carriage position, the control unit triggers the third-priority axle compensation process. This axle compensation process is based on the prior assumption that the structure of long-formation train carriages has regularity and continuity, and generates estimated axle coordinate values ​​for the missing carriages through mathematical interpolation.

[0110] Specifically, the control unit first reads the pre-configured car length parameters. These parameters can be retrieved from a parameter database based on the actual train formation and model being inspected, or they can be manually set by the operator before the inspection based on train information. Furthermore, they can be dynamically adjusted during the inspection process based on on-site feedback. Subsequently, the control unit searches forward within the generated final axle positioning data sequence for the nearest car with valid final axle positioning data, obtaining its final axle positioning data. Simultaneously, it searches backward for the nearest car with valid final axle positioning data, obtaining its final axle positioning data. In a preferred embodiment, this forward and backward search process can specify a maximum search span. If no valid car is found within the maximum search span, the search range is automatically expanded or an end-point extrapolation method is used.

[0111] Based on the acquired final axle positioning data of the forward and backward effective carriages, and the pre-configured carriage length parameters, the control unit uses a linear interpolation algorithm to calculate the estimated axle coordinates of the currently missing carriage. For example, this linear interpolation calculation process can be expressed as: Xcomplement equals Xfront plus ((Xback minus Xfront) divided by Ninterval) multiplied by noffset, where Xcomplement represents the estimated axle X-coordinate of the carriage to be supplemented, Xfront represents the axle X-coordinate value of the forward effective carriage, Xback represents the axle X-coordinate value of the backward effective carriage, Ninterval represents the number of carriages between the forward and backward effective carriages, and noffset represents the carriage offset of the carriage to be supplemented relative to the forward effective carriage. The estimated Y and Z coordinates can be calculated in the same way or directly set based on the track plane height.

[0112] After calculating the estimated axle coordinates, the control unit determines these estimated axle coordinates as the final axle positioning data for the current carriage and writes them into the record position corresponding to the carriage number in the final axle positioning data sequence. Simultaneously, the data source is marked as interpolated in this record. According to a further improvement in this embodiment, the axle coordinate estimates obtained using the above linear interpolation method have a deviation from the actual axle position under actual operating conditions controlled within an allowable range of ±10mm. This accuracy meets the accuracy requirements for axle positioning data in subsequent train maintenance and safety inspection operations.

[0113] The control unit sequentially executes the hierarchical decision-making and processing operations S3.2 to S3.6 for each car in the dual-source data matrix according to the car number order. When all cars have been processed, the recorded position of each car in the final axle positioning data sequence is assigned a valid axle coordinate value, and there is no missing data. This complete axle coordinate sequence constitutes the complete axle positioning dataset for the entire train. This complete axle positioning dataset also includes metadata information for this inspection operation, including but not limited to the inspection timestamp, train number, robot number, total number of train cars, and the data source identifier of the final axle positioning data for each car, to facilitate subsequent data traceability and multi-dimensional statistical analysis. This complete axle positioning dataset can be uploaded to the central database of the rail transit inspection management system via the robot's communication module for use by downstream tasks such as train maintenance scheduling and safety status assessment. This axle fixing and replacement operation is now complete.

[0114] The beneficial effects of this invention are as follows: By constructing a fixed-axis axle compensation mechanism that prioritizes vision, supplements with laser, and utilizes interpolation, it solves the problems of easy failure and poor data continuity in existing single-sensor positioning methods under complex working conditions. At the data acquisition level, a 2D camera and a laser sensor work in parallel. The 2D camera provides high-precision visual positioning under normal operating conditions, while the laser sensor provides supplementary positioning when vision is affected by lighting, dirt, etc. At the data processing level, the fixed-axis axle compensation strategy, based on carriage number alignment and hierarchical priority decision rules, prioritizes the visual data with the highest positioning accuracy. When visual positioning fails, it automatically and seamlessly switches to laser data. When both sensors fail, it relies on prior parameters of the train carriage structure to generate estimated coordinates that meet engineering accuracy requirements through linear interpolation. This achieves zero-loss and fully continuous output of axle positioning data across the entire train. The overall solution achieves full automation from autonomous positioning and mobile inspection of the inspection robot, to dual-channel parallel perception, coordinate alignment, fixed-axis axle compensation, and the generation of a complete dataset, without human intervention. This provides an efficient and reliable intelligent axle positioning solution for large-scale, routine maintenance and safety inspection of rail transit trains.

[0115] The preferred embodiments of the present invention have been described in detail above. However, the present invention is not limited to the specific details in the above embodiments. Within the scope of the technical concept of the present invention, various equivalent transformations can be made to the technical solutions of the present invention, and these equivalent transformations all fall within the protection scope of the present invention.

Claims

1. A method for fixed-axis compensation based on vision and laser fusion, characterized in that, include: As the robot moves along the side of the train, the 2D camera and the laser sensor are triggered in parallel to collect axle-related data and generate visual axle coordinate data and laser axle coordinate data, forming visual positioning axle coordinate dataset and laser positioning axle coordinate dataset, respectively; the 2D camera and the laser sensor are mounted on the robot. The visual positioning axle coordinate dataset and the laser positioning axle coordinate dataset are matched and aligned according to the carriage number to form a dual-source data matrix aligned by carriage. The dual-source data matrix is ​​processed car by car according to a preset priority rule to generate a complete axle positioning dataset for the entire train car.

2. The method according to claim 1, characterized in that, The formation of the visual positioning axle coordinate dataset and the laser positioning axle coordinate dataset includes a first branch and a second branch: The first branch works through the 2D camera, triggers the 2D camera to capture images of the train side at a preset frequency, identifies the axle position based on the train side images and calculates the axle coordinates in the global coordinate system, generates the visual axle coordinate data, and collects it to form the visual positioning axle coordinate dataset. The second branch operates through the laser sensor, driving the laser sensor to continuously scan the side of the train to obtain three-dimensional point cloud data. Based on the three-dimensional point cloud data, it extracts the geometric features of the axle and calculates the coordinates of the axle in the global coordinate system, generating the laser axle coordinate data, which is then collected to form the laser positioning axle coordinate dataset.

3. The method according to claim 2, characterized in that, The first branch specifically includes: The 2D camera receives the hard trigger signal generated by the microcontroller and the robot's real-time global coordinate information, captures images of the train's side, and binds the train's side images with the real-time global coordinate information at the time of capture to form image data frames with spatial location tags. For the image data frames with spatial location labels, a target detection method is used to extract multi-scale features of the image, which are then matched with pre-trained axle target features to filter out image regions containing axles. For the selected image regions containing axles, calculate the pixel coordinates of the axle center in the image coordinate system, and convert them into X-direction and Y-direction deviation values ​​of the axle relative to the 2D camera using camera intrinsic parameters. Using the real-time global coordinate information bound in the image data frame with spatial location labels, the X-direction deviation value and Y-direction deviation value are converted into the absolute coordinates of the axle in the global coordinate system, generating visual axle coordinate data with timestamps.

4. The method according to claim 2, characterized in that, The second branch specifically includes: The laser sensor is driven to continuously scan the side surface of the train to obtain a dense three-dimensional point cloud data stream. Each data point in the three-dimensional point cloud data stream contains coordinate information and reflection intensity information in the global coordinate system. Perform point cloud filtering on the three-dimensional point cloud data stream to remove noise points and outliers, and obtain denoised point cloud data; Clustering and segmentation processing is performed on the denoised point cloud data to separate the point cloud regions of different objects and obtain multiple point cloud clusters. Geometric features are extracted from each of the multiple point cloud clusters, and the extracted geometric features are matched with a preset axle laser template, wherein the axle laser template defines the geometric parameters and spatial position constraints of the axle target; When the geometric features of a certain point cloud cluster and the matching deviation of the axle laser template are within a preset threshold range, it is determined that the point cloud cluster corresponds to an axle. The centroid coordinates of the point cloud cluster are calculated and used as the coordinates of the axle in the global coordinate system, generating laser axle coordinate data with timestamps.

5. The method according to claim 1, characterized in that, Forming the dual-source data matrix aligned by carriage includes: Read the visual positioning axle coordinate dataset and the laser positioning axle coordinate dataset; The timestamp information of each data record in the two datasets is extracted. Based on the timestamp sequence of the visual positioning axle coordinate dataset, time interpolation processing is performed on the laser positioning axle coordinate dataset to achieve alignment of the two datasets in the time dimension. Confirm that the axle coordinates in both datasets have been transformed to the same global coordinate system, thus completing the spatial benchmark unification; Based on prior structural knowledge of long-formation trains, the coordinate data of the axles contained in the same carriage in the visual positioning axle coordinate dataset and the laser positioning axle coordinate dataset are associated with the corresponding carriage number. A structured data table is constructed with the carriage number as the row index and the visual axle coordinate data and the laser axle coordinate data as the two data fields, forming the carriage-aligned dual-source data matrix.

6. The method according to claim 1, characterized in that, The generated complete axle positioning dataset for all train carriages includes: According to the carriage number order, traverse the data rows corresponding to each carriage in the dual-source data matrix aligned by carriage; For the currently traversed carriage, check whether the visual axle coordinate data corresponding to the carriage is valid; If the visual axle coordinate data is valid, then the visual axle coordinate data is extracted as the final axle positioning data of the carriage and written into the record position corresponding to the carriage in the final axle positioning data sequence.

7. The method according to claim 6, characterized in that, The preset priority rules also include: If the visual axle coordinate data corresponding to the current carriage is invalid, then further check whether the laser axle coordinate data corresponding to the carriage is valid; If the laser axle coordinate data is valid, then the laser axle coordinate data is extracted as the final axle positioning data of the carriage and written into the record position corresponding to the carriage in the final axle positioning data sequence.

8. The method according to claim 7, characterized in that, The preset priority rules also include: if both the visual axle coordinate data and the laser axle coordinate data corresponding to the current carriage are invalid, then axle compensation processing is triggered; the axle compensation processing includes the following steps: Read the pre-configured carriage length parameters and obtain the final axle positioning data of the carriages that are adjacent to the current carriage in front and behind and have valid final axle positioning data; Based on the carriage length parameter and the final axle positioning data of the adjacent carriages, a linear interpolation algorithm is used to calculate the estimated axle coordinates of the current carriage. The estimated axle coordinates are used as the final axle positioning data for the carriage, and written into the corresponding record position of the carriage in the final axle positioning data sequence.

9. The method according to claim 1, characterized in that, The robot is equipped with a navigation module. As the robot moves along the side of the train, the navigation module continuously acquires the robot's real-time global coordinate information and performs trajectory correction based on the real-time global coordinate information to maintain a preset detection distance between the robot and the side of the train.

10. The method according to claim 3, characterized in that, The target detection method is based on the YOLOv8 network.