Anti-shake data acquisition and obstacle identification method for rail transit vehicle

CN121191135BActive Publication Date: 2026-09-08CHINA ACADEMY OF RAILWAY SCI CORP LTD +1
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202511459538.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-10-13
Publication Date
2026-09-08
Estimated Expiration
2045-10-13

AI Technical Summary

Technical Problem

[0002]在轨道交通车辆运行过程中,由于车辆高速行驶及轨道振动,车载摄像头采集到的图像常因抖动产生模糊和位移,导致后续障碍物检测精度下降,给行车安全带来风险,因此,如何在抖动环境下获得高质量、稳定的图像数据,并对车辆前方的障碍物进行实时、准确的识别,一直是业界关注的热点问题

Benefits of technology

1.利用同步与错时拍摄相结合采集连续四帧原始图像,再通过预先构建的图层和精细的像素方格划分,为后续的局部像素级对齐提供统一参考坐标;采用全局运动估计和局部特征匹配,结合创新设计的非线性运动估计与局部对齐公式,有效补偿车辆抖动所产生的图像偏移和失真,使得各帧图像能够实现高精度对齐和无缝融合,从而获得高质量的稳定图像。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121191135B_ABST
    Figure CN121191135B_ABST
Patent Text Reader

Abstract

The application discloses a method for anti-shake data acquisition and obstacle identification of a rail transit vehicle, relates to the technical field of rail transit vehicle data acquisition and processing, and adopts synchronous and staggered shooting to collect four continuous original images through a binocular camera, performs global motion estimation in combination with collected sensor data, realizes local pixel-level alignment through preset layer and pixel grid division, generates a stable image through weighted fusion, optimizes image local features based on a multi-head attention mechanism, maps the stable image to a high-dimensional vector space through a deep encoder, fuses motion correction information, finally, generates three-dimensional point cloud through a deep reconstruction network combined with binocular disparity, monitors three-dimensional coordinate changes in real time, realizes obstacle identification and early warning, can obtain high-quality images in a severe vibration environment of the vehicle, improves data acquisition accuracy and real-time performance and accuracy of obstacle detection, and effectively guarantees rail transit operation safety.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data acquisition and processing technology for rail transit vehicles, specifically a method for anti-vibration data acquisition and obstacle recognition for rail transit vehicles. Background Technology

[0002] During the operation of rail transit vehicles, due to the high speed of the vehicle and the vibration of the track, the images captured by the on-board camera are often blurred and displaced due to shaking, which leads to a decrease in the accuracy of subsequent obstacle detection and poses a risk to driving safety. Therefore, how to obtain high-quality and stable image data in a shaking environment and to identify obstacles in front of the vehicle in real time and accurately has always been a hot issue of concern in the industry.

[0003] Chinese invention patent application CN117876992A discloses an obstacle detection method, mainly targeting autonomous vehicles. It detects foreground and background obstacles by analyzing point cloud data in the current frame and then fuses the two data to reduce the risk of missed detections. This method focuses on using point cloud data for obstacle detection. However, in the scenario of rail transit where vehicle vibration is severe, relying solely on point cloud data may result in insufficient accuracy and poor adaptability to high dynamic environments. At the same time, it lacks deep fusion of images captured by cameras and sensor data, and cannot fully compensate for the impact of vehicle shaking.

[0004] Chinese invention patent application CN117058649A discloses an obstacle recognition method for unmanned delivery vehicles based on panoramic image stitching. This method preprocesses the field-of-view images captured by the unmanned vehicle, extracts features, calculates the homography matrix using the Random Sample Consensus (RANSAC) algorithm, fuses and stitches the images to form a panoramic image, and then performs obstacle recognition. Although this scheme has achieved certain results in terms of the overall quality of the panoramic image and the accuracy of obstacle recognition, its image stitching technology is prone to image mis-stitching and inaccurate local alignment in the high-speed and shaking environment of rail transit, resulting in unstable obstacle detection results. In addition, this method has low dependence on vibration compensation, making it difficult to meet the requirements of high-precision obstacle recognition under violent motion conditions. Summary of the Invention

[0005] The purpose of this application is to overcome the shortcomings of the prior art and propose a method for anti-shake data acquisition and obstacle recognition for rail transit vehicles to solve the above-mentioned problems.

[0006] This application acquires continuous images through synchronous and staggered shooting methods, and combines gyroscope and accelerometer data to perform global and local motion estimation, thereby achieving accurate pixel-level alignment and stable image fusion, making up for the shortcomings of single point cloud or image stitching methods that are susceptible to jitter in dynamic environments.

[0007] This application effectively reduces image shift and noise caused by vehicle vibration through global motion estimation, local alignment, and image fusion formulas, obtaining high-quality and stable images and providing accurate data for subsequent obstacle detection.

[0008] This application introduces a depth encoder to map a stable image to a high-dimensional vector space, and uses binocular parallax combined with a depth reconstruction network to generate a high-precision 3D point cloud, thereby reconstructing the real 3D environment in front of the vehicle.

[0009] This application utilizes a multi-head attention model and time series analysis to monitor continuous frame 3D data in real time, quickly identify obstacles with abnormal movement or inactivity, and trigger early warning measures in a timely manner to ensure driving safety.

[0010] The system of this application automatically performs self-calibration and error correction during operation, dynamically adjusts the parameters of each module, and ensures high accuracy and real-time performance during long-term operation, adapting to different environmental conditions.

[0011] The purpose of this application is achieved through the following technical solution: a method for anti-vibration data acquisition and obstacle recognition for rail transit vehicles, comprising the following steps: S1, Image and Sensor Data Acquisition Using a binocular camera, four consecutive raw images are captured by a combination of synchronous and staggered shooting, while sensor data corresponding to each raw image is also collected. S2, Layer Construction and Pixel Grid Division The four original images were placed into pre-created layers 1, 2, 3 and 4 respectively. Each layer was subdivided into pixel squares of a fixed size, so that the pixel squares at the same position in each layer corresponded one-to-one. At the same time, layer 5 was pre-created to store the stable images generated later. S3, Preliminary Global Motion Estimation Based on the sensor data collected in S1, global motion parameters are estimated for each frame of image to obtain global motion parameters and provide prior information for subsequent local pixel-level alignment. S4, Local pixel-level alignment By combining the global motion parameters obtained from S3 with the local feature matching within each pixel square in the layer, the displacement of the corresponding pixel square in each layer is iteratively optimized to achieve precise pixel-level alignment and obtain local alignment and correction parameters. S5, Image fusion to generate stable images Based on the alignment result of S4, the pixel squares in layers 1 to 4 are fused using a weighted fusion or median filtering method to generate a stable image, which is then stored in layer 5. S6, Multi-head Attention Optimization Processing Using a multi-head attention model based on Transformer or ViT, the four original images in S1 are preprocessed and divided into fixed-size patches. The encoded sensor data is used as input to calculate the displacement difference between each patch, thereby further guiding local cropping, stitching and fusion, and optimizing the quality and low noise output of the stable image in S5. S7, High-dimensional vector space mapping The stable image generated in S5 is input into the depth encoder. During the mapping process, the local alignment and correction parameters obtained in S4 are fused to convert the stable image into a high-dimensional feature vector and output a discriminative image feature vector. S8, 3D Reconstruction and Point Cloud Generation Based on binocular visual disparity information and image feature vectors of consecutive frames, a depth reconstruction network is used to generate depth maps and 3D coordinate data of each object in the scene through depth regression and disparity estimation techniques, and a 3D environment model in front of the vehicle is reconstructed in the form of point cloud data. S9, Real-time 3D Coordinate Monitoring and Obstacle Recognition Using another model based on a multi-head attention mechanism, the three-dimensional coordinate data generated in continuous frames in S8 is processed in time series. Through differential calculation and Kalman filtering algorithm, the displacement, velocity and acceleration of objects are monitored in real time. Combined with the preset middle area of ​​the track and the vehicle's travel range, obstacles are identified for abnormally stationary or suddenly moving objects, triggering early warning measures.

[0012] In some possible implementations, in S1, the binocular camera uses the following acquisition method: synchronous shooting: the left and right cameras simultaneously capture the first frame image; staggered shooting: within a predetermined time interval, the left and right cameras respectively capture the second frame image; the combination of synchronous shooting and staggered shooting constitutes four consecutive original images, providing temporal and spatial redundancy information for subsequent S2 to S9.

[0013] In this implementation, the combination of synchronous and staggered shooting can obtain sufficient redundant data in the vibrating environment of rail transit, ensuring that enough effective image information can be captured even under extreme shaking conditions. At the same time, the sensor data provides a reliable basis for motion compensation.

[0014] In some possible implementations, the sensor data collected in S1 includes gyroscope data and accelerometer data. By corresponding to the image acquisition time, a preliminary estimate of the global motion parameters of each frame of the image is achieved, and prior information is provided for local pixel-level alignment in S3 and S4.

[0015] In this implementation, the data corresponds precisely to the image acquisition time, providing prior parameters for subsequent global motion estimation and local alignment.

[0016] In some possible implementations, the pre-established layers in S2 include 5 blank layers, wherein: layers 1 to 4 are used to store the acquired original images; layer 5 is used to store the stable images generated after processing by S4 and S5; each layer is subdivided into fixed-size pixel squares to ensure that pixel squares at the same position correspond one-to-one in each layer, thereby achieving precise alignment and fusion.

[0017] In this implementation, layers are pre-established and subdivided into pixel grids, enabling precise alignment of different frame images in a unified coordinate system. This reduces the computational complexity of global alignment and improves alignment accuracy, providing an accurate reference basis for subsequent image fusion.

[0018] In some possible implementations, when processing the image using a Transformer or ViT model in S6, the following sub-steps are included: preprocessing each frame of the image; dividing the preprocessed image into fixed-size patches; using the patches and encoded sensor data as multimodal inputs, calculating the displacement differences between each patch through a multi-head attention mechanism to guide local cropping, stitching, and fusion, thereby optimizing the stable image generated in S5.

[0019] In this implementation, the multi-head attention mechanism fully utilizes the detailed information between image patches and sensor data, which can effectively capture local motion changes and further compensate for possible local alignment deficiencies. Through multimodal data fusion, the local correction parameters are more accurate, and the optimized stable image has significant improvements in detail clarity and noise suppression, providing higher quality input for subsequent feature extraction.

[0020] In some possible implementations, in step S7, the depth encoder uses a residual network, a convolutional autoencoder, or a ViT encoding part to map the stable image in step S5 to a high-dimensional vector space, and fuses the local alignment and correction parameters obtained in step S4 during the mapping process, so that the output image feature vector contains both spatial structure information and motion correction information, providing accurate input for the three-dimensional reconstruction in step S8.

[0021] In this implementation, the depth encoder can effectively extract high-level semantic information from images, and by fusing local alignment parameters, the output feature vector is significantly improved in terms of noise resistance and accuracy. The high-dimensional feature vector provides a richer and more accurate description for subsequent 3D reconstruction and obstacle recognition, further enhancing the robustness and real-time performance of the entire system.

[0022] In some possible implementations, in step S8, the 3D reconstruction adopts the principle of binocular vision, utilizes the parallax information between the left and right camera images and the image feature vector obtained in step S7, and uses depth regression and parallax estimation techniques through a depth reconstruction network to generate depth maps and 3D coordinate data of each object in the scene, and reconstructs the 3D environment model in front of the vehicle in the form of point cloud data.

[0023] In this implementation, point cloud data provides high-precision and rich spatial information for subsequent obstacle detection; combined with the principle of binocular vision, it enables accurate reconstruction of the 3D environment model in front of the vehicle even in shaking environments. It fully considers the contributions of image details and motion correction information under vehicle shaking conditions and achieves adaptive compensation for depth estimation through nonlinear mapping.

[0024] In some possible implementations, the real-time three-dimensional coordinate monitoring in S9 includes the following sub-steps: using another model based on a multi-head attention mechanism to process the three-dimensional coordinate data generated in the continuous frames in S8, automatically capturing the displacement changes of each object in the local and global ranges; performing differential calculation on the three-dimensional coordinate data of the continuous frames, and combining Kalman filtering or other time series prediction algorithms for smoothing, to calculate the velocity and acceleration information of the object in real time.

[0025] In this implementation, the three-dimensional coordinate data of consecutive frames are processed in real time, enabling obstacles to be detected in a very short time.

[0026] In some possible implementations, obstacle recognition in S9 further includes the following sub-steps: fusing and analyzing the real-time monitored three-dimensional coordinate change information of the object with the preset track middle area and vehicle travel range; when continuous frame data indicates that there are motion features in a certain predefined area, the object is identified as an obstacle, and the motion features include, but are not limited to, at least one of the following: the object is abnormally still, the object suddenly accelerates, or the object changes direction abruptly; once an obstacle is identified, the system immediately triggers early warning measures and notifies the vehicle control system through image marking, sound alarm, or data transmission to take automatic braking, deceleration, or other safety measures.

[0027] In this implementation, the combination of multi-head attention model and Kalman filtering improves the accuracy of motion information extraction in complex dynamic environments; by integrating preset region information, irrelevant motion is effectively filtered out, the false alarm rate is reduced, and the reliability of safety warnings is ensured.

[0028] In some possible implementations, a self-calibration and error correction step (S10) is also included. This step is automatically executed during the stable driving phase of the vehicle. Specifically, it includes: performing joint calibration on the collected sensor data using static environmental data to correct errors caused by temperature drift, sensor noise, or equipment aging; dynamically adjusting the model parameters through a feedback mechanism based on the processing results of each module in S4, S7, S8, and S9 to ensure that the overall data processing flow always maintains high accuracy and real-time performance; and periodically updating the global motion parameters and local alignment parameters within each module based on the calibration results, thereby further improving the stability and accuracy of the image stabilization data acquisition and obstacle recognition system.

[0029] In this implementation, self-calibration and error correction ensure that the system can automatically compensate for sensor drift and model errors during long-term operation, maintaining overall high accuracy. Through dynamic feedback adjustment, the entire data processing flow can adapt to different environmental changes and vehicle status changes, improving the stability, robustness, and reliability of the anti-shake data acquisition and obstacle recognition system, and ensuring the safe operation of rail transit vehicles.

[0030] The beneficial effects of this application are: 1. By combining synchronous and staggered shooting to acquire four consecutive raw images, and then using pre-constructed layers and fine pixel grid division, a unified reference coordinate is provided for subsequent local pixel-level alignment; global motion estimation and local feature matching are adopted, combined with an innovatively designed nonlinear motion estimation and local alignment formula, to effectively compensate for image offset and distortion caused by vehicle shaking, so that each frame of images can achieve high-precision alignment and seamless fusion, thereby obtaining high-quality stable images.

[0031] 2. By using a multi-head attention mechanism based on Transformer or ViT, nonlinear adaptive correction is performed on image patches. Sensor data is used to assist in adjustment, further optimizing local details. The innovatively designed multi-head attention correction formula can dynamically balance the differences between local regions and global average features, thereby effectively suppressing local noise and residual errors caused by jitter, ensuring that the image is clearer and smoother in detail.

[0032] 3. A deep encoder (e.g., residual network, convolutional autoencoder, ViT coding part, etc.) is used to map the stable image to a high-dimensional vector space, and local alignment correction information is fused during the mapping process, so that the output image feature vector contains rich spatial structure information and motion compensation information at the same time; the innovative feature fusion formula can adaptively filter regions with large alignment errors, enhance the discriminativeness of the overall features, and provide accurate and robust input data for subsequent 3D reconstruction and obstacle recognition.

[0033] 4. Based on the principle of binocular vision, using disparity information between images and high-dimensional feature vector data, a high-precision depth map and three-dimensional coordinate data are generated through a depth reconstruction network; using an innovatively designed depth regression and disparity estimation model, stable images can be converted into three-dimensional point cloud data that reflects the actual scene, thereby accurately reconstructing a real three-dimensional environment model in front of the vehicle.

[0034] 5. Through multi-head attention mechanism and time series processing (including differential calculation and Kalman filtering methods), the system tracks the three-dimensional coordinate changes of each object in continuous frames in real time and calculates the motion state (displacement, velocity, acceleration); by fusing preset track area information, when continuous frame data indicates that there is an object that is abnormally stationary or suddenly accelerates or changes direction in a certain area, the system can quickly identify the obstacle and immediately trigger an alarm, transmitting safety warning information to the vehicle control system to ensure driving safety.

[0035] 6. The system automatically collects static environmental data during the stable driving phase of the vehicle and uses a joint calibration algorithm to correct the gyroscope and camera data, compensating for errors caused by temperature drift, sensor noise, or equipment aging; through a feedback mechanism, the parameters of each module are dynamically adjusted to ensure that the overall data processing flow always maintains high precision and real-time performance, so that the system can still stably and reliably collect anti-shake data and identify obstacles during long-term operation. Attached Figure Description

[0036] Figure 1 For the steps of this application Figure 1 ; Figure 2 For the steps of this application Figure 2 ; Figure 3 For the steps of this application Figure 3 ; Figure 4 For the steps of this application Figure 4 ; Figure 5 This is a diagram of the apparatus used in this application. Detailed Implementation

[0037] The technical solutions of this application will be clearly and completely described below with reference to the embodiments. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are within the scope of protection of this application.

[0038] It should be noted that the directional concepts of "left", "right", "up", "down", "front", "back", "inside", and "outside" in the following scheme are all relative directions, and will not be listed one by one here.

[0039] Example 1: like Figures 1 to 4 As shown, the focus of this embodiment is to achieve stable image generation by acquiring multiple frames of image data, constructing fine layers, and aligning pixels under conditions of high-speed vehicle travel and strong vibration, thus providing high-quality basic data for subsequent obstacle recognition and 3D reconstruction.

[0040] S1: Image and sensor data acquisition.

[0041] The system uses a high-resolution binocular camera to acquire image data by combining synchronous and staggered shooting.

[0042] Simultaneous shooting: The left and right cameras capture the first frame image at the same time, ensuring that the images have the same timestamp and overlapping viewpoints.

[0043] Staggered shooting: Within a predetermined, extremely short time interval, the left and right cameras respectively capture the second frame of the image.

[0044] The combination of these two acquisition methods forms four consecutive frames of raw image data, providing temporal and spatial redundancy information for subsequent steps.

[0045] Simultaneously with image acquisition, the system also collects sensor data corresponding to each frame of the image and records vehicle motion information (e.g., angular velocity, acceleration, displacement changes, etc.).

[0046] These data correspond precisely to the image acquisition time, providing prior parameters for subsequent global motion estimation and local alignment.

[0047] By combining synchronous and staggered shooting, sufficient redundant data can be obtained in the severely vibrating environment of rail transit, ensuring that enough effective image information can be captured even under extreme shaking conditions. At the same time, the sensor data provides a reliable basis for motion compensation.

[0048] S2: Layer construction and pixel grid division.

[0049] Five blank layers are created in the system beforehand: Layers 1 through 4: These are used to store the four original images captured in step S1.

[0050] Layer 5: Used to store the stable image generated after subsequent processing.

[0051] Pixel grid division: Subdivide each layer by dividing it into fixed-size pixel squares (e.g., 16×16 pixels).

[0052] The pixel squares at the same position in each layer correspond one-to-one, forming a unified reference coordinate system, which facilitates subsequent image alignment and fusion.

[0053] By pre-establishing layers and subdividing them into pixel grids, images from different frames can be precisely aligned in a unified coordinate system. This reduces the computational complexity of global alignment and improves alignment accuracy, providing an accurate reference basis for subsequent image fusion.

[0054] S3: Preliminary global motion estimation.

[0055] Global motion parameter estimation: The sensor data collected in S1 includes gyroscope data and accelerometer data. Using the gyroscope data and accelerometer data, global motion parameters are estimated for each frame of image to obtain global motion parameters.

[0056] The calculations include: Translation: Estimate the overall translational change of the image based on acceleration data and time intervals.

[0057] Rotation angle: The rotation angle of the camera is calculated using gyroscope data.

[0058] Prior information transmission: The aforementioned global motion parameters are passed as prior information to step S4 to provide an initial estimate for subsequent local pixel-level alignment, thereby narrowing the local search range.

[0059] Global motion estimation provides accurate initial values ​​for local alignment, which can reduce the computational burden of local matching and provide effective compensation when vehicle vibration is large, thereby reducing the accumulation of errors caused by vibration.

[0060] The global motion estimation formula is used to calculate the global translation and rotation parameters for each frame of the image. Let the sensor input for each frame be:

[0061] in, Indicates the magnitude of acceleration. It represents angular velocity.

[0062] Define the global motion parameter vector as follows:

[0063]

[0064] in, These are constant parameters obtained through system calibration.

[0065] This formula uses a combination of square root, logarithm, and power functions to nonlinearly map the acceleration and angular velocity information in the sensor into translation and rotation parameters, effectively reducing the impact of jitter noise on global motion estimation and providing robust prior information for subsequent local alignment.

[0066] S4: Local pixel-level alignment.

[0067] Local feature extraction and matching: Within the pixel squares of each layer in S2, local features are extracted using either the Scale-invariant feature transform (SIFT) algorithm or the Oriented Fast and Rotated BRIEF algorithm.

[0068] By comparing the local features of corresponding pixel squares in different layers, feature matching relationships can be obtained.

[0069] Iterative optimization alignment: Based on the global motion parameters provided by S3, iterative optimization algorithms (such as least squares method, RANSAC, etc.) are used to refine the matching results of each pixel square.

[0070] Calculate the precise local displacement between each layer to ensure that the corresponding pixel areas are accurately aligned.

[0071] Local pixel-level alignment effectively compensates for local image distortion caused by vehicle vibration, ensuring high-precision matching of each frame in the same coordinate system, providing a foundation for the generation of stable images, and reducing seam and ghosting problems in subsequent image fusion.

[0072] Calculate the local displacement correction for each pixel square in the layer, assuming it is in image I. i In this context, a pixel at position (x, y) belongs to a local region, and its corresponding local neighborhood is denoted as N(x, y). The local alignment displacement vector is defined as Δp. i (x,y)=(Δx i (x,y),Δy i (x,y)), the formula is expressed as:

[0073] in, Scaling factor For image midpoint pixel values, This is the average value of the pixels in all corresponding regions. Standard deviation, It is a small constant to prevent the denominator from being zero. Based on local gradient direction The calculated weight vector.

[0074] This formula normalizes the deviation of pixels in a local area and assigns weights based on the local gradient direction to achieve a fine displacement estimate for each pixel square. Its adaptive adjustment capability can compensate for local image distortion caused by vehicle vibration and ensure alignment accuracy.

[0075] S5: Image fusion generates a stable image.

[0076] Applications of fusion algorithms: Based on the alignment results of step S4, the corresponding pixels in layers 1 to 4 are merged.

[0077] The fusion methods can be: Weighted fusion: Weights are assigned based on local feature matching scores or image sharpness, and a weighted average is performed on the pixel squares of each frame; Median filtering: Take the median value for each corresponding pixel square to filter out random errors and noise.

[0078] Stable image output: The fusion result is stored as a stable image in a pre-created layer 5.

[0079] The output stable images have high clarity and stability, providing high-quality basic data for subsequent obstacle recognition or 3D modeling steps.

[0080] By fusing multiple precisely aligned image frames, this embodiment achieves significant results in eliminating vehicle vibration and random noise. The generated stable image has high clarity and rich detail, which can greatly improve the accuracy and robustness of subsequent processing algorithms.

[0081] Adaptive fusion is performed on the pixels of each frame after local alignment. Let I be the image of each frame after correction in step S4. i ′=(i=1,2,3,4), the fusion result F(x,y) at pixel position (x,y) is defined as:

[0082] in, and The horizontal and vertical displacements are respectively calculated by the local alignment formula (2). For in position The corresponding reference pixel value is usually the median of four frame pixel values. and The parameters are adjusted to control the sensitivity of weight decay.

[0083] This fusion formula introduces adaptive exponential weights, which non-linearly weight the deviations of each frame's pixels from the reference value. This more accurately suppresses residual noise and eliminates edge ghosting caused by imperfect local alignment. The weighting mechanism ensures that the generated stable image has higher clarity and detail retention, laying a solid foundation for subsequent processing.

[0084] By implementing steps S1 to S5, this embodiment achieves high-precision image acquisition and anti-shake processing in environments with high-speed rail transit vehicles and significant vibration. Data redundancy: Synchronous and staggered shooting ensures data redundancy in time and space, enabling the capture of effective images even under extreme vibration conditions.

[0085] Precise alignment: The pre-established layers and pixel grid divisions provide a unified reference system, and combined with global and local motion parameter estimation, pixel-level precise alignment is achieved.

[0086] Image stability: By iteratively optimizing alignment and weighted / median fusion, the blur and noise caused by vehicle vibration can be effectively eliminated, thus improving the quality of stable images.

[0087] Subsequent processing foundation: The generated high-quality and stable images provide a solid data foundation for subsequent multi-head attention optimization, depth coding and 3D reconstruction, thereby improving the accuracy and real-time performance of the overall obstacle recognition system.

[0088] The above is Example 1, which demonstrates the method and process for achieving high-precision image stabilization data acquisition in the vibration environment of rail transit vehicles.

[0089] Example 2: like Figures 1 to 4 As shown, in Example 1, a stable image (stored in layer 5) was obtained by performing layer construction, pixel-level alignment and fusion on the four original images collected. However, due to local detail errors and residual noise in the vehicle shaking environment, in order to further optimize the image quality and extract more discriminative feature information, this example introduces multi-head attention optimization processing and high-dimensional vector mapping on the basis of the stable image.

[0090] S6: Multi-head attention optimization processing.

[0091] Input preprocessing: Using the four frames of raw image data acquired in step S1 of Example 1, each frame of image is preprocessed first.

[0092] Preprocessing includes: Normalization: Normalizes the pixel values ​​of an image to ensure that the input data is within a uniform numerical range.

[0093] Denoising: Image denoising algorithms (such as mean filtering, bilateral filtering, etc.) are used to reduce noise.

[0094] Color correction: Corrects color deviations in images to ensure consistency in color, brightness, and other aspects between different frames.

[0095] Patch partitioning and multimodal data construction: Each preprocessed frame image is divided into several patches according to a fixed size (such as 16×16 or 32×32 pixels).

[0096] Meanwhile, the sensor data collected in step S1 (after preliminary motion parameter estimation and encoding) is processed and used as auxiliary information to form a multimodal input with the image patch.

[0097] Multi-head attention mechanism calculation: Using a multi-head attention model based on Transformer or ViT, the patches of each frame image and the corresponding sensor data are input into the model; the model calculates the displacement difference between each patch through the multi-head attention mechanism to capture subtle motion information in local areas.

[0098] Based on the calculation results, the model automatically generates relative displacement correction parameters for each patch between frames, thereby guiding local cropping, splicing, and fusion.

[0099] Optimize the stable image output; using the above local correction results, further optimize the local details of the stable image generated in step S5 of Example 1; by finely cropping and stitching the local areas, eliminate residual local misalignment and noise, and output a more refined and low-noise stable image.

[0100] By employing a multi-head attention mechanism to fully utilize the detailed information between image patches and sensor data, local motion changes can be effectively captured, further compensating for any potential local alignment deficiencies in Example 1.

[0101] By fusing multimodal data, the local correction parameters are more accurate, and the optimized stable image shows significant improvements in detail clarity and noise suppression, providing higher quality input for subsequent feature extraction.

[0102] Assuming each frame of the image is preprocessed and divided into several fixed-size patches, the feature representation of each patch in the i-th frame and the j-th patch is as follows:

[0103] in, This represents the h-th head in a multi-head sequence. Simultaneously, the sensor motion code for the corresponding frame is:

[0104] Define the average feature of the j-th patch in the h-th header across all frames as:

[0105] Design an adaptive correction factor to reflect the deviation between individual patch features and the global average, and incorporate sensor data to assist in adjustment:

[0106] in, The scaling parameter in header h, For sensor data The modulation factor is obtained by mapping through a nonlinear activation function.

[0107] Next, the features of each patch are corrected:

[0108] in, is the learning rate factor for the head h, used to control the correction magnitude.

[0109] Finally, the aggregated output of the corrected patch features of each frame at the h-th header is:

[0110] After further weighted integration, the final output of this patch is obtained:

[0111] in, The weights of each head.

[0112] This formula uses an exponential function to nonlinearly amplify the difference between local patch features and average features, while introducing sensor data for modulation, making the correction of each patch more adaptive. The multi-branch extracts information from different angles, and the final weighted fusion output effectively reduces local deviations caused by vehicle vibration, improving the accuracy and robustness of subsequent processing (such as high-dimensional mapping and 3D reconstruction).

[0113] S7: High-dimensional vector space mapping.

[0114] The depth encoder input is the stable image optimized by step S6 (stored in layer 5) as input to the depth encoder.

[0115] The deep encoder can use a residual network, a convolutional autoencoder, or a ViT encoding part. The specific model can be selected according to the actual application environment, and a pre-trained model can be used as the initial weight.

[0116] Local alignment parameters are fused; during the encoding process, the local alignment and correction parameters obtained in step S4 of Example 1 are fused.

[0117] This fusion enables the encoder to not only extract global features of a stable image, but also embed anti-shake correction information, making the feature vector more discriminative and robust.

[0118] High-dimensional vector mapping and output: After processing by a deep encoder, the stable image is mapped to a high-dimensional vector space, generating a compact and discriminative image feature vector.

[0119] The output high-dimensional feature vector not only reflects the spatial structure information of the stable image, but also contains motion compensation information, laying an accurate foundation for subsequent 3D reconstruction (S8) and obstacle recognition based on this feature.

[0120] Deep encoders can effectively extract high-level semantic information from stable images, and by fusing local alignment parameters, the output feature vectors are greatly improved in terms of noise resistance and accuracy.

[0121] High-dimensional feature vectors provide a richer and more accurate description for subsequent 3D reconstruction and obstacle recognition, further enhancing the robustness and real-time performance of the entire system.

[0122] Suppose that the stable image generated in step S5 is divided into N patches after preprocessing, and the normalized pixel feature vector I is taken for each patch j. j (For example, obtained through local averaging) and the local alignment displacement vector Δ obtained from step S4. j (This can be expressed as a local translation error), and the mapping formula is: First, construct the composite descriptor for each patch:

[0123] in, This is the attenuation parameter, used to suppress the influence of patches with large alignment errors. To integrate weights, The exponents are nonlinear amplification parameters, all determined through system calibration. The exponents are calculated element by element to form a nonlinear mapping.

[0124] To highlight patches with better alignment, we define a weighting factor:

[0125] in, Controlling the weight decay rate results in a smaller alignment error. Smaller (smaller) corresponds to higher weight.

[0126] Then, the high-dimensional feature vector V of the entire image is obtained by weighted summation of the descriptors of each patch:

[0127] Finally, V is normalized to obtain the final high-dimensional feature vector:

[0128] Through steps S6 and S7 of Example 2, the system achieves the following based on the stable image of Example 1: Detail optimization and noise reduction: Multi-head attention optimization effectively extracts and corrects motion differences between local patches, resulting in more refined and less noisy image fusion results.

[0129] High-dimensional vector mapping fuses image and motion correction information into compact feature vectors, providing accurate and rich data input for subsequent 3D reconstruction and obstacle recognition.

[0130] Example 2 significantly improves the image processing workflow in terms of local alignment, detail restoration, and feature extraction, further enhancing the overall accuracy and robustness of rail transit vehicles in the anti-shake data acquisition and obstacle recognition system.

[0131] The above embodiment 2 achieves multi-head attention optimization and high-dimensional vector space mapping based on embodiment 1 through steps S6 and S7, providing high-quality input for subsequent 3D reconstruction and obstacle detection, and has significant practical application effects.

[0132] Example 3: like Figures 1 to 4 As shown, S8: 3D reconstruction and point cloud generation.

[0133] Input data preparation; This step uses the image feature vector output in step S7 of Example 2. This data contains image details and fused motion correction information after multi-head attention optimization. Meanwhile, the disparity information in the left and right images captured by the binocular camera (based on known camera intrinsic and extrinsic parameters) is used as an auxiliary depth estimation basis.

[0134] Deep reconstruction networks; processing input data using deep reconstruction networks (e.g., deep regression models based on convolutional neural networks, specialized disparity estimation networks, etc.): Depth Regression: The network generates depth maps of objects in the scene using depth regression technology based on high-dimensional feature vectors and binocular disparity information; Disparity Estimation: By combining the disparity information from the left and right images, the details in the depth map are further corrected to ensure the accuracy of depth estimation.

[0135] During network training, a large amount of real orbital environment data and synthetic data can be used for pre-training, giving the model a strong generalization ability.

[0136] 3D coordinate data calculation and point cloud generation: Using depth maps and camera calibration parameters, the depth information of each pixel in the image is converted into 3D coordinates to form point cloud data; Point cloud data focuses on the environment in front of the vehicle, reflecting the position, shape, and relative distance of each object in three-dimensional space; through depth regression and disparity estimation techniques, it ensures that the generated depth map and three-dimensional coordinate data accurately reflect the actual scene; Point cloud data provides high-precision and rich spatial information for subsequent obstacle detection; combined with the principle of binocular vision, it can still accurately reconstruct the three-dimensional environment model in front of the vehicle even in a shaking environment.

[0137] The formula for 3D reconstruction and point cloud generation organically integrates the high-dimensional image feature vector output from step S7 with binocular parallax information to generate a depth map and then convert it into 3D coordinate data, ultimately constructing a point cloud. The formula design fully considers the contribution of image details and motion correction information under vehicle vibration conditions, and achieves adaptive compensation for depth estimation through nonlinear mapping.

[0138] Set at pixel position Location: The parallax value is measured directly by the binocular camera; The high-dimensional image feature vector output from step S7; is the reference feature vector obtained through system calibration; f is the camera focal length, and B is the baseline distance of the binocular cameras; For characteristic correction coefficients, This is a non-linear scaling factor.

[0139] The depth estimation formula is:

[0140] In the denominator, It reflects the depth information of traditional disparity estimation, while A high-dimensional feature vector is introduced as a compensation term to compensate for the difference between the vector and the reference template, enabling adaptive adjustment of vehicle vibration and local motion compensation; exponential The nonlinear response of the overall mapping is used to adjust the depth estimation to better match the actual scene distribution.

[0141] Point cloud generation and 3D coordinate transformation formulas: Using the above depth Combined with the camera's intrinsic parameters (principal point coordinates) and focal length ), calculate pixels Corresponding three-dimensional coordinates :

[0142] in, The parameters are adjusted to attenuate areas where features deviate significantly from the reference values, thereby further improving the accuracy of 3D reconstruction.

[0143] Based on the traditional pinhole model, the following was introduced The attenuation factor non-linearly adjusts the horizontal and vertical coordinates of pixels, resulting in better local alignment (i.e., Smaller regions have higher weights; Finally, by combining the three-dimensional coordinates of each pixel, point cloud data reflecting the environment in front of the vehicle can be formed, providing high-precision and rich spatial information.

[0144] S9: Real-time 3D coordinate monitoring and obstacle recognition.

[0145] Processing of continuous frame 3D coordinate data; using the point cloud data generated from continuous frames in step S8, and processing it through another model based on a multi-head attention mechanism: Multi-head attention extraction of motion information: This model automatically captures the changes in the three-dimensional coordinates of each region in continuous frames and identifies the differences in local and global displacements; Differential calculation: Differentiate the three-dimensional coordinate data of continuous frames to obtain the displacement changes of each object in a short period of time.

[0146] Motion state smoothing and velocity / acceleration calculation; combining Kalman filtering or other time series prediction algorithms to smooth the differential data and eliminate noise effects; The real-time velocity and acceleration information of each object is calculated, reflecting the dynamic characteristics of the object's motion state.

[0147] Obstacle recognition and early warning; real-time monitoring of the object's three-dimensional coordinates and motion parameters is fused and analyzed with the preset track center area and vehicle travel range; The criteria for judgment include: if an object in a predefined area shows abnormal stillness in consecutive frames (which may be an obstacle falling on the track); or if it exhibits abnormal motion characteristics such as sudden acceleration or abrupt change in direction, then the object is identified as an obstacle.

[0148] Once an obstacle is identified, the system immediately triggers warning measures, such as marking the image on a stable view, issuing an audible alarm, or transmitting data to the vehicle control system to take automatic braking, deceleration, or other safety measures.

[0149] Real-time processing of three-dimensional coordinate data from consecutive frames enables obstacles to be detected in a very short time. The combination of multi-head attention model and Kalman filter improves the accuracy of motion information extraction in complex dynamic environments; by integrating pre-set region information, irrelevant motion is effectively filtered out, the false alarm rate is reduced, and the reliability of safety warnings is ensured.

[0150] Multi-head attention-weighted motion difference extraction: Let the three-dimensional coordinates of region j in the i-th frame be:

[0151] The original difference between each frame is defined as:

[0152] For each attention head Let the average coordinates of region j in the previous frame be:

[0153] Constructing a multi-head weighted difference:

[0154] in, Assigning weights to each head, To adjust the parameters.

[0155] This formula utilizes the local statistical information of each attention head to perform nonlinear weighting on the motion difference of consecutive frames, which can adaptively filter noise and highlight the true signal of motion changes within the region.

[0156] Dynamic Kalman filter update status: Define state vector (Position and velocity). An innovative adaptive Kalman update formula is employed:

[0157] Among them, the measurement matrix (usually taken) Kalman gain Nonlinear adjustment is adopted:

[0158] in, It is a short matrix with constant gain. This refers to the sensitivity parameter.

[0159] This formula extracts motion differences from multi-head attention and uses an exponential function to dynamically adjust the Kalman gain, thereby making the state update more adaptable to motion changes in a high-dynamic jitter environment.

[0160] Real-time motion anomaly detection and obstacle recognition: Based on the updated state, the real-time velocity and acceleration are first calculated using finite difference:

[0161] Score for abnormal motion:

[0162] in, This represents the local motion standard deviation of region j around frame i. and To adjust the parameters, The speed is predicted based on historical data (obtained using simple time series forecasting or filtering). Ultimately, if... If a preset threshold is set, then region j will be identified as an obstacle in frame i, triggering a warning (e.g., image marking, sound alarm, or data transmission to the vehicle control system).

[0163] S10: Self-calibration and error correction.

[0164] Static environment data acquisition: During the stable driving phase of the vehicle, the system automatically acquires static environment data. At this time, the image and sensor data are stable and there is no significant motion disturbance.

[0165] Joint calibration algorithm; using static environmental data, jointly calibrate the collected gyroscope data and camera data: correct errors caused by temperature drift, sensor noise or equipment aging, etc.; output corrected global motion parameters and local alignment parameters.

[0166] Feedback mechanism and dynamic parameter adjustment: Based on the feedback results of S4 (local pixel-level alignment), S7 (high-dimensional vector mapping), S8 (3D reconstruction), and S9 (3D monitoring), the parameters within each module are dynamically adjusted through the feedback mechanism. The motion compensation coefficients, local alignment parameters, and some weights of the deep reconstruction network are dynamically updated; the calibration parameters are updated regularly to ensure the accuracy and real-time performance of each module during long-term operation.

[0167] Self-calibration and error correction ensure that the system can automatically compensate for sensor drift and model errors during long-term operation, maintaining overall high accuracy; through dynamic feedback adjustment, the entire data processing flow can adapt to different environmental changes and vehicle status changes. Improve the stability, robustness, and reliability of anti-shake data acquisition and obstacle recognition systems to ensure the safe operation of rail transit vehicles.

[0168] Example 3 further transforms the high-quality images and feature vectors obtained in steps S1 to S7 into three-dimensional point cloud data, and monitors the motion state of objects in the three-dimensional environment in real time, maintaining the long-term accuracy of the system through dynamic calibration; Precise 3D Reconstruction: Utilizing binocular visual parallax information and high-dimensional features, accurately reconstruct the 3D environment in front of the vehicle, providing precise information for safety decisions.

[0169] Real-time obstacle detection: Dynamic monitoring and fusion analysis of continuous frame data can quickly identify obstacles that are moving abnormally or stationary, and trigger early warning measures in a timely manner.

[0170] Improved system stability: The self-calibration and error correction mechanism enables the system to maintain high-precision operation under different working conditions, effectively coping with sensor drift and environmental interference.

[0171] Through Example 3, the entire anti-shake data acquisition and obstacle recognition system, from image acquisition, stabilization processing, feature extraction to 3D reconstruction and real-time early warning, has constructed a complete, efficient and robust technical solution, which effectively improves the safety and reliability of rail transit vehicles in vibration environments.

[0172] Figure 5 The device diagram of this application shows that the anti-shake data acquisition and obstacle recognition device 500 for rail transit vehicles includes: The image and sensor data acquisition module 501 is used to acquire four consecutive frames of raw images by using a binocular camera and combining synchronous and staggered shooting, while simultaneously acquiring sensor data corresponding to each frame of raw images. The layer construction and pixel grid division module 502 is used to place the four original images captured into the pre-established layers 1, 2, 3 and 4 respectively. Each layer is subdivided into pixel grids of fixed size so that the pixel grids at the same position in each layer correspond one-to-one. At the same time, layer 5 is pre-established to store the stable image generated later. The preliminary global motion estimation module 503 is used to estimate the global motion parameters of each frame of image based on the sensor data acquired in the image and sensor data acquisition module 501, obtain the global motion parameters, and provide prior information for subsequent local pixel-level alignment. The local pixel-level alignment module 504 is used to combine the global motion parameters obtained by the preliminary global motion estimation module 503 with the local feature matching within each pixel square in the layer, to iteratively optimize the displacement of the corresponding pixel square in each layer, thereby achieving accurate pixel-level alignment and obtaining local alignment and correction parameters. The image fusion to generate a stable image module 505 is used to fuse the pixel squares in layers 1 to 4 using a weighted fusion or median filtering method based on the alignment result of the local pixel-level alignment module 504, generate a stable image, and store the stable image in layer 5. The multi-head attention optimization processing module 506 is used to preprocess the four original images in S1 using a multi-head attention model based on Transformer or ViT, divide them into fixed-size patches, and combine the encoded sensor data as input to calculate the displacement difference between each patch, thereby further guiding local cropping, stitching and fusion, and optimizing the quality and low noise output of the stable image in the image fusion generating stable image module 505; The high-dimensional vector space mapping module 507 is used to input the stable image generated in the image fusion to the stable image generation module 505 into the depth encoder. During the mapping process, the local alignment and correction parameters obtained by the local pixel-level alignment module 504 are fused to convert the stable image into a high-dimensional feature vector and output a discriminative image feature vector. The 3D reconstruction and point cloud generation module 508 is used to generate depth maps and 3D coordinate data of each object in the scene based on binocular vision disparity information and image feature vectors of continuous frames, using a depth reconstruction network, depth regression and disparity estimation techniques, and reconstructing the 3D environment model in front of the vehicle in the form of point cloud data. The real-time 3D coordinate monitoring and obstacle recognition module 509 is used to perform time series processing on the 3D coordinate data generated by the continuous frames in the 3D reconstruction and point cloud generation module 508 using another model based on a multi-head attention mechanism. Through differential calculation and Kalman filtering algorithm, it monitors the displacement, velocity and acceleration of objects in real time. Combined with the preset middle area of ​​the track and the vehicle's travel range, it identifies obstacles such as abnormally stationary or suddenly moving objects and triggers early warning measures.

[0173] In some possible implementations, the binocular camera in the image and sensor data acquisition module 501 adopts the following acquisition methods: synchronous shooting: the left and right cameras simultaneously capture the first frame image; staggered shooting: within a predetermined time interval, the left and right cameras respectively capture the second frame image; the combination of synchronous shooting and staggered shooting constitutes four consecutive original images, providing temporal and spatial redundancy information for the subsequent layer construction and pixel grid division module 502 to the real-time three-dimensional coordinate monitoring and obstacle recognition module 509.

[0174] In some possible implementations, the sensor data acquired in the image and sensor data acquisition module 501 includes gyroscope data and accelerometer data. By corresponding to the image acquisition time, a preliminary estimate of the global motion parameters of each frame of the image is achieved, and prior information is provided for the local pixel-level alignment in the preliminary global motion estimation module 503 and the local pixel-level alignment module 504.

[0175] In some possible implementations, the layers pre-established in the layer construction and pixel grid division module 502 include five blank layers, wherein: layers 1 to 4 are used to store the acquired original images; layer 5 is used to store the stable image generated after processing by the local pixel-level alignment module 504 and the image fusion to generate a stable image module 505; each layer is subdivided into pixel squares of a fixed size to ensure that pixel squares at the same position correspond one-to-one in each layer, thereby achieving precise alignment and fusion.

[0176] In some possible implementations, the multi-head attention optimization processing module 506 is used to preprocess each frame of image; divide the preprocessed image into patches of fixed size; use the patches and encoded sensor data as multimodal inputs; calculate the displacement difference between each patch through the multi-head attention mechanism; guide local cropping, stitching and fusion; and thus optimize the stable image generated in the image fusion and stable image generation module 505.

[0177] In some possible implementations, in the high-dimensional vector space mapping module 507, the depth encoder uses a residual network, a convolutional autoencoder, or a ViT encoding part to map the stable image in the image fusion to stable image generation module 505 to a high-dimensional vector space. During the mapping process, the local alignment and correction parameters obtained by the local pixel-level alignment module 504 are fused, so that the output image feature vector contains both spatial structure information and motion correction information, providing accurate input for the 3D reconstruction and point cloud generation module 508.

[0178] In some possible implementations, the 3D reconstruction and point cloud generation module 508 uses the principle of binocular vision for 3D reconstruction. It utilizes the disparity information between the left and right camera images and the image feature vectors obtained by the high-dimensional vector space mapping module 507. Through a depth reconstruction network, it uses depth regression and disparity estimation techniques to generate depth maps and 3D coordinate data of each object in the scene, and reconstructs the 3D environment model in front of the vehicle in the form of point cloud data.

[0179] In some possible implementations, the real-time 3D coordinate monitoring and obstacle recognition module 509 is used to process the 3D coordinate data generated by the 3D reconstruction and point cloud generation module 508 in continuous frames using another model based on a multi-head attention mechanism, automatically capturing the displacement changes of each object in the local and global range; performing differential calculation on the 3D coordinate data of continuous frames, and combining Kalman filtering or other time series prediction algorithms for smoothing, to calculate the velocity and acceleration information of the object in real time.

[0180] In some possible implementations, the real-time three-dimensional coordinate monitoring and obstacle recognition module 509 is used to fuse and analyze the real-time monitored three-dimensional coordinate change information of the object with the preset track middle area and vehicle travel range; when continuous frame data indicates that there are motion features in a certain predefined area, the object is identified as an obstacle, and the motion features include, but are not limited to, at least one of the following: the object is abnormally still, the object suddenly accelerates, or the object changes direction abruptly; once an obstacle is identified, the system immediately triggers early warning measures and notifies the vehicle control system through image marking, sound alarm or data transmission, so as to take automatic braking, deceleration or other safety measures.

[0181] In some possible implementations, the device also includes a self-calibration and error correction module, used to perform joint calibration of the acquired sensor data using static environmental data, correcting errors caused by temperature drift, sensor noise, or equipment aging; based on the processing results of each module in the local pixel-level alignment module 504, high-dimensional vector space mapping module 507, 3D reconstruction and point cloud generation module 508, and real-time 3D coordinate monitoring and obstacle recognition module 509, the model parameters are dynamically adjusted through a feedback mechanism to ensure that the overall data processing flow always maintains high precision and real-time performance; based on the calibration results, the global motion parameters and local alignment parameters within each module are updated periodically, thereby further improving the stability and accuracy of the image stabilization data acquisition and obstacle recognition system.

[0182] The above description is merely a preferred embodiment of this application. It should be understood that this application is not limited to the form disclosed herein and should not be construed as excluding other embodiments. It can be used in various other combinations, modifications, and environments, and can be modified within the scope of the concept described herein through the above teachings or the technology or knowledge in related fields. Modifications and changes made by those skilled in the art that do not depart from the spirit and scope of this application should be protected within the scope of the appended claims.

Claims

1. A method for vibration stabilization data acquisition and obstacle recognition in rail transit vehicles, characterized in that, Includes the following steps: S1, Image and Sensor Data Acquisition Using a binocular camera, four consecutive raw images are captured by a combination of synchronous and staggered shooting, while sensor data corresponding to each raw image is also collected. In step S1, the binocular camera uses the following acquisition method: Simultaneous shooting: The left and right cameras capture the first frame image simultaneously; Staggered shooting: Within a predetermined time interval, the left and right cameras respectively capture the second frame image; The combination of synchronous shooting and staggered shooting constitutes four consecutive original images, providing temporal and spatial redundancy information for subsequent S2 to S9. S2, Layer Construction and Pixel Grid Division The four original images were placed into pre-created layers 1, 2, 3 and 4 respectively. Each layer was subdivided into pixel squares of a fixed size, so that the pixel squares at the same position in each layer corresponded one-to-one. At the same time, layer 5 was pre-created to store the stable images generated later. The pre-established layers in S2 include 5 blank layers, among which: Layer 1 to Layer 4 are respectively used to store the acquired original images; Layer 5 is used to store the stable image generated after processing by S4 and S5; Each layer is subdivided into pixel squares of a fixed size; S3, Preliminary Global Motion Estimation Based on the sensor data collected in S1, global motion parameters are estimated for each frame of image to obtain global motion parameters and provide prior information for subsequent local pixel-level alignment. S4, Local pixel-level alignment By combining the global motion parameters obtained from S3 with the local feature matching within each pixel square in the layer, the displacement of the corresponding pixel square in each layer is iteratively optimized to achieve precise pixel-level alignment and obtain local alignment and correction parameters. S5, Image fusion to generate stable images Based on the alignment result of S4, the pixel squares in layers 1 to 4 are fused using a weighted fusion or median filtering method to generate a stable image, which is then stored in layer 5. S6, Multi-head Attention Optimization Processing Using a multi-head attention model based on Transformer or ViT, the four original images in S1 are preprocessed and divided into fixed-size patches. The encoded sensor data is used as input to calculate the displacement difference between each patch, thereby further guiding local cropping, stitching and fusion, and optimizing the quality and low noise output of the stable image in S5. S7, High-dimensional vector space mapping The stable image generated in S5 is input into the depth encoder. During the mapping process, the local alignment and correction parameters obtained in S4 are fused to convert the stable image into a high-dimensional feature vector and output a discriminative image feature vector. S8, 3D Reconstruction and Point Cloud Generation Based on binocular visual disparity information and image feature vectors of consecutive frames, a depth reconstruction network is used to generate depth maps and 3D coordinate data of each object in the scene through depth regression and disparity estimation techniques, and a 3D environment model in front of the vehicle is reconstructed in the form of point cloud data. S9, Real-time 3D Coordinate Monitoring and Obstacle Recognition Using another model based on multi-head attention mechanism, the three-dimensional coordinate data generated in continuous frames in S8 is processed in time series. Through differential calculation and Kalman filtering algorithm, the displacement, velocity and acceleration of objects are monitored in real time. Combined with the preset middle area of ​​the track and the vehicle travel range, obstacles are identified for abnormally stationary or suddenly moving objects, and early warning measures are triggered. S10. Self-calibration and error correction steps, which are automatically executed during stable vehicle operation, specifically include: The collected sensor data is jointly calibrated using static environmental data to correct errors caused by temperature drift, sensor noise, or equipment aging. Based on the processing results of each module in S4, S7, S8 and S9, the model parameters are dynamically adjusted through a feedback mechanism; The global motion parameters and local alignment parameters within each module are updated periodically based on the calibration results.

2. The method for vibration stabilization data acquisition and obstacle recognition for rail transit vehicles according to claim 1, characterized in that: The sensor data collected in S1 includes gyroscope data and accelerometer data. By corresponding to the image acquisition time, a preliminary estimate of the global motion parameters of each frame of the image is achieved, and prior information is provided for local pixel-level alignment in S3 and S4.

3. The method for vibration stabilization data acquisition and obstacle recognition for rail transit vehicles according to claim 1, characterized in that: In step S6, when processing the image using a Transformer or ViT model, the following sub-steps are included: Preprocess each frame of the image; The preprocessed image is divided into patches of fixed size; Using patches and encoded sensor data as multimodal inputs, a multi-head attention mechanism is used to calculate the displacement differences between patches, guiding local cropping, stitching, and fusion, thereby optimizing the stable image generated in S5.

4. The method for vibration stabilization data acquisition and obstacle recognition for rail transit vehicles according to claim 1, characterized in that: In step S7, the depth encoder uses a residual network, a convolutional autoencoder, or a ViT encoding part to map the stable image in step S5 to a high-dimensional vector space. During the mapping process, the local alignment and correction parameters obtained in step S4 are fused together so that the output image feature vector contains both spatial structure information and motion correction information, providing accurate input for the three-dimensional reconstruction in step S8.

5. The method for vibration stabilization data acquisition and obstacle recognition for rail transit vehicles according to claim 1, characterized in that: In step S8, the three-dimensional reconstruction adopts the principle of binocular vision. It utilizes the parallax information between the left and right camera images and the image feature vector obtained in step S7. Through a depth reconstruction network, it uses depth regression and parallax estimation techniques to generate depth maps and three-dimensional coordinate data of each object in the scene, and reconstructs the three-dimensional environment model in front of the vehicle in the form of point cloud data.

6. The method for vibration stabilization data acquisition and obstacle recognition for rail transit vehicles according to claim 1, characterized in that: The real-time three-dimensional coordinate monitoring in S9 includes the following sub-steps: Another model based on multi-head attention mechanism is used to process the three-dimensional coordinate data generated in continuous frames in S8, and the displacement changes of each object in the local and global range are automatically captured. Differential calculations are performed on the three-dimensional coordinate data of consecutive frames, and smoothing is achieved by combining Kalman filtering or other time series prediction algorithms to calculate the velocity and acceleration information of objects in real time.

7. The method for vibration stabilization data acquisition and obstacle recognition for rail transit vehicles according to claim 1, characterized in that: The obstacle recognition in S9 further includes the following sub-steps: The real-time monitoring of the object's three-dimensional coordinate changes is fused and analyzed with the preset track center area and vehicle travel range. When continuous frame data indicates the presence of motion features within a predefined region, the object is identified as an obstacle. Motion features include: the object is abnormally still, the object suddenly accelerates, and the object's orientation changes abruptly. Once an obstacle is detected, the system immediately triggers a warning mechanism, notifying the vehicle control system via image marking, audible alarm, or data transmission, so that automatic braking, deceleration, or other safety measures can be taken.

Citation Information

Patent Citations

  • Unmanned express delivery vehicle obstacle identification method and system based on panoramic image stitching

    CN117058649A

  • Obstacle detection method, device and equipment and automatic driving vehicle

    CN117876992A

  • Custom layer generation method and device, vehicle and storage medium

    CN115148030A

  • Automatically detecting traffic signals using sensor data

    US20230169780A1