Vision-based rapid identification method for target object in track area

By using four vehicle-mounted phase mechanisms to build a fuzzy database and image calibration model in track detection, and dynamically correcting the error, the problem of insufficient processing delay and matching accuracy caused by image blur in high-speed trains is solved, real-time and accurate track target detection is achieved, and the system's robustness and detection coverage are improved.

CN120259639AActive Publication Date: 2025-07-04CRRC HANGZHOU DIGITAL TECH CO LTD

Patent Information

Application Number
CN202510748205.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-06
Publication Date
2025-07-04
Estimated Expiration
2045-06-06

AI Technical Summary

Technical Problem

During the high-speed train driving, the existing track detection methods face the problems of processing delay caused by image blur, insufficient blurred image matching accuracy, and low efficiency of multi-camera collaborative calibration, which is difficult to meet the real-time detection requirements and improve the reliability and robustness of recognition.

Method used

Four parallel-setting vehicle cameras are adopted to build a fuzzy database and image calibration model to realize real-time image calibration and feature matching, dynamically correct common mode errors and differential mode errors, and improve the detection coverage and robustness of the system.

Benefits of technology

Real-time accurate detection of track target objects in high-speed train environments is achieved, the reliability and safety of detection is improved, the risk of processing delays and missed detection is reduced, and dynamic interference in complex operating environments is adapted to.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120259639A_ABST
    Figure CN120259639A_ABST
Patent Text Reader

Abstract

The invention relates to the field of track detection, in particular to a vision-based rapid identification method for a target object in a track area, which comprises a target object image sampling step, a fuzzy database construction step, a sampled image calibration step and a target object rapid identification step. The problems of image blurring, multi-source errors and real-time performance under high-speed movement are systematically solved; the beneficial effects of efficiently coping with image blurring and dynamic interference, improving system robustness and detection coverage rate through a multi-camera cooperative calibration mechanism, adapting to a complex operation environment through a dynamic error elimination strategy, and improving detection reliability and safety are achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of track detection, and specifically relates to a method for quickly identifying objects in a track area based on vision. Background Art

[0002] With the rapid development of rail transit, track detection technology has become an important guarantee for ensuring the safe operation of trains. In recent years, object recognition methods based on vision have been widely applied to the field of track detection. They collect track images in real time through on-vehicle cameras and identify abnormal objects on the track (such as foreign objects, damage to track components, etc.) in combination with image processing technology. However, during the high-speed operation of trains, the existing vision recognition methods still face the following technical bottlenecks: (1) Image blur leads to processing delay: Due to the high train speed, the images collected by on-vehicle cameras are easily affected by motion blur and vibration interference, with low image resolution and blurred edges. Traditional image enhancement and deblurring algorithms require a large amount of computing resources and are difficult to meet the real-time processing requirements. Especially in complex lighting or dynamic environments, the processing time of existing algorithms increases significantly, resulting in lag in object recognition and inability to timely feedback potential safety hazards. (2) Insufficient matching accuracy for blurred images: Existing methods usually rely on a clear image database for feature matching, but it is difficult to extract features from blurred images, and misjudgment or missed detection is likely to occur during direct matching. For example, the contour of the object in the blurred image is deformed or the color is distorted, resulting in inaccurate feature mapping relationships and seriously affecting the reliability of the recognition results. (3) Low efficiency of multi-camera collaborative calibration: To improve the detection coverage rate, existing technologies often use multi-camera systems, but it is difficult to effectively eliminate the geometric deformation and errors of images caused by differences in installation angles of different cameras or environmental interference (such as vibration, lighting changes). Traditional calibration methods rely on complex hardware adjustments or static calibration and cannot dynamically adapt to the real-time changes during train operation, with deficiencies in both calibration accuracy and efficiency. In addition, there is a lack of targeted correction strategies for the common-mode errors (such as overall offset) and differential-mode errors (such as local anomalies) of multi-camera systems, further reducing the robustness of the system.

[0003] Therefore, to solve the above problems, the present invention proposes a method for quickly identifying objects in a track area based on vision. Summary of the Invention

[0004] Aiming at the deficiencies of the existing technology, the purpose of the present invention is to provide a method for quickly identifying objects in a track area based on vision.

[0005] To achieve the above purpose, the present invention provides the following technical solutions: A method for quickly identifying objects in a track area based on vision, comprising the following steps: An object image sampling step of collecting an image of the track area through an on-vehicle camera to obtain a blurred sampling image of the object and a spatial calibration image; Steps for constructing a fuzzy database: training the database with the collected comparison fuzzy image data, extracting the fuzzy image features and establishing a mapping relationship with specific target objects to obtain a fuzzy database; Steps for calibrating a sampled image: constructing an image calibration model based on the spatial calibration image and a pre-collected standard calibration image, and calibrating the fuzzy sampled image of the target object through the image calibration model to obtain the fuzzy image data of the target object; Steps for quickly identifying a target object: comparing the fuzzy image features of the fuzzy image data with those of the comparison fuzzy image data in the fuzzy database to obtain the target object with a mapping relationship and outputting it as the identified target object.

[0006] As a further improvement of the present invention, the vehicle-mounted camera includes a sampling camera and a calibration camera, which are set as four cameras arranged in parallel, including a left sampling camera, a right sampling camera, a left calibration camera, and a right calibration camera. The left sampling camera and the left calibration camera are arranged on the left front side in the advancing direction of the rail train, and the right sampling camera and the right calibration camera are arranged on the right front side in the advancing direction of the rail train. The left sampling camera and the left calibration camera are mirror-symmetric with the right sampling camera and the right calibration camera along the rail center line. The camera sampling areas of the left calibration camera and the right calibration camera are known, and standard calibration images with a directly overhead view are collected.

[0007] As a further improvement of the present invention, it further includes steps for detecting a calibration camera. When the sampling areas and the sizes of the collected images of the left calibration camera and the right calibration camera are different, determining the abnormal calibration camera and the normal calibration camera through a calibration layer correction strategy, and selecting the normal calibration camera or the calibrated calibration camera as the calibration camera for collecting the spatial calibration image.

[0008] As a further improvement of the present invention, the sampling image calibration step includes constructing an image calibration model, performing pixel simplification processing on the spatial calibration images collected by the calibration camera, pixelating the spatial calibration images to obtain spatial calibration image pixel values and comparing them with pixel thresholds preset according to the standard image features. When the spatial calibration image pixel values are greater than the pixel thresholds, perform an equal-ratio pixel value reduction on the spatial calibration images according to the standard image feature ratio, perform multi-valued processing on the reduced spatial calibration images, divide the pixel points into several preset RGB thresholds according to the preset RGB gradient range values, confirm the image boundaries of the regions corresponding to the standard image features in the sampling area of the calibration camera in the image according to the threshold classification strategy for the multi-valued processed spatial calibration images, and perform masking processing on the spatial calibration images according to the image boundaries to obtain a masking processing result. Construct the calibration layer of the image calibration model according to the masked part and the background part of the masking processing result through the fuzzy image calibration rule. The fuzzy image calibration rule calibrates the angle and size of the image input into the image calibration model by comparing the masking processing result with the size of the standard image features; after processing the blurred sampling image of the target object through the calibration layer, input it into the fuzzy image comparison layer to obtain the blurred image data of the target object through image comparison processing.

[0009] As a further improvement of the present invention, the construction of the image calibration model includes the calibration camera and sampling camera synchronization step. Construct a deconvolution layer through the spatial calibration images of the calibration camera, and establish a geometric mapping relationship between the image pixel values and the actual track features according to the known fixed track features collected by the calibration camera; based on the geometric mapping relationship, perform deconvolution processing on the blurred sampling image of the target object collected by the sampling camera to correct the image deformation caused by the installation angle difference between the sampling camera and the calibration camera, so that the calibrated blurred image features are consistent with the feature mapping relationship in the blurred database, and complete the camera synchronization.

[0010] As a further improvement of the present invention, the threshold classification strategy includes extracting the color features and shape features of the spatial calibration images, setting thresholds for the pixel points of the pixelated spatial calibration images according to the color feature data of the standard image features, and setting the pixel points of the spatial calibration images to several fixed RGB values according to the threshold partition. By comparing the color features of the spatial calibration images and the standard image features, narrow the RGB value difference to obtain the RGB threshold, and extract the pixel points with RGB values greater than the threshold in the spatial calibration images to generate image boundaries.

[0011] As a further improvement of the present invention, it further includes a calibration layer correction step. The calibration layer correction step includes a common-mode error correction step and a differential-mode error correction step. When the error amounts of the left calibration camera and the right calibration camera are the same, the output is a common-mode error; otherwise, the output is a differential-mode error. The common-mode error correction step calculates the error through the overall translation in the spatial calibration images collected by the two calibration cameras; the differential-mode error calculates the error through the differential translation caused by the abnormal local orbital features of the calibration cameras.

[0012] As a further improvement of the present invention, the common-mode error correction step includes: obtaining the spatial calibration images collected by the left calibration camera and the right calibration camera, comparing the overall offset of the spatial calibration images with the pre-stored standard calibration image, calculating the common-mode error coefficient according to the overall offset, generating a reverse compensation matrix according to the common-mode error coefficient, applying the reverse compensation matrix to the blurred sampling image of the sampling camera, and performing translation or scaling compensation on the blurred sampling image to eliminate the common-mode error.

[0013] As a further improvement of the present invention, the differential-mode error correction step includes: extracting the pixel abnormal features in the local areas of the spatial calibration images of the left calibration camera and the right calibration camera, and comparing them with the corresponding areas of the standard calibration image; determining the differential-mode error area according to the comparison result of the difference, performing pixel-level segmentation on the differential-mode error area, extracting the RGB values and spatial coordinates of the abnormal pixel points, and reconstructing the target values of the abnormal pixel points based on the RGB gradient distribution of the neighboring pixels by using the bilinear interpolation algorithm to repair the abnormal area. Taking the repaired calibration image as the input to update the calibration layer parameters.

[0014] As a further improvement of the present invention, it further includes an error elimination strategy. The error elimination strategy includes: dynamically updating the preset pixel threshold and the RGB gradient range threshold in the pixel simplification process according to the historical error data of the calibration camera, setting a feedback loop in the calibration layer, performing a secondary comparison of the image after each correction with the standard calibration image, adjusting the deconvolution layer parameters or the compensation matrix according to the residual error until the error is lower than the preset threshold. When the error of a certain side calibration camera continuously exceeds the tolerance range, it is marked as an abnormal camera, and the standby calibration camera is switched to or the calibration data of the mirror symmetry side is used to complete the correction.

[0015] The beneficial effects of the present invention are: (1) It can efficiently handle image blur and dynamic interference and achieve real-time and accurate detection. Traditional methods are limited by motion blur and vibration interference, with low image resolution and significant processing delays, making it difficult to meet the real-time detection requirements of high-speed trains. The present invention solves this problem and achieves the following beneficial effects: By collecting a large number of comparison fuzzy image data to train the database, directly establish the mapping relationship between the fuzzy image features and the target object, and avoid the feature mismatch problem caused by relying on the clear image database. For example, for the edge deformation features of the track fastener fuzzy image, the database can store the gradient distribution model under different degrees of fuzziness, so that the recognition algorithm can quickly match the corresponding target object.

[0016] Use the spatial calibration images collected by the calibration camera to construct an image calibration model including a calibration layer and a comparison layer, and calibrate the angle and size of the fuzzy sampling image through pixel simplification, multi-valued processing and masking rules.

[0017] (2) Multi-camera collaborative calibration mechanism to improve the system robustness and detection coverage. The existing multi-camera systems have low calibration efficiency due to the installation angle differences and environmental interferences. The present invention significantly optimizes the performance through the symmetric layout and dynamic correction strategy: The left sampling / calibration camera and the right sampling / calibration camera are symmetrically arranged along the center line of the track, covering the full-width detection range of the track (for example, the detection width of a single-side camera is 3.5 meters, and the total for both sides is 7 meters), avoiding the field of view blind area problem of the traditional single-side camera.

[0018] By distinguishing the common-mode error (overall offset) and the differential-mode error (local anomaly), achieve targeted correction. The common-mode error eliminates the overall translation or scaling through the reverse compensation matrix; the differential-mode error repairs the local anomaly through pixel-level interpolation, and the system robustness is significantly enhanced.

[0019] (3) Dynamic error elimination strategy to adapt to complex operating environments. For the dynamic interferences such as vibration and light changes during train operation, the present invention establishes a multi-level error suppression mechanism: Based on the historical error data of the calibration camera, dynamically adjust the pixel threshold and the RGB gradient range.

[0020] Through the secondary comparison between the calibration layer and the standard image (such as the mean square error MSE detection), dynamically adjust the parameters of the deconvolution layer or the compensation matrix. When the error of a single-side camera continuously exceeds the limit, automatically switch to the standby camera or use the mirror-symmetric data (such as the horizontally flipped image of the right camera to replace the left camera) to ensure the uninterrupted operation of the system and improve the fault tolerance rate.

[0021] (4) Improve the detection reliability and safety. The existing fuzzy image matching is prone to false detection / missing detection due to feature distortion. The present invention guarantees the result reliability through the full-process optimization: For local anomalies (such as the camera jitter artifacts caused by vibration), reconstruct the abnormal pixel values through the bilinear interpolation algorithm, improve the recognition accuracy of fine features such as track cracks, and avoid missing detection caused by fuzziness.

[0022] Reduce potential safety hazards, reduce processing latency, and gain critical time for train braking or manual intervention. Brief Description of the Drawings

[0023] Figure 1 is the flowchart of the method for quickly identifying the target object of the present invention; Figure 2 is the schematic diagram of the on-vehicle camera of the present invention; Figure 3 is the structural diagram of the image calibration model of the present invention; Figure 4 is the flowchart of the error elimination step of the present invention. Detailed Embodiment

[0024] The present invention will be further described in detail below with reference to the drawings and embodiments. The same reference numerals are used for the same components. It should be noted that the terms "front", "rear", "left", "right", "upper" and "lower" used in the following description refer to the directions in the drawings, and the terms "bottom surface" and "top surface", "inner" and "outer" refer to the directions towards or away from the geometric center of a specific component respectively.

[0025] The purpose of this embodiment is to solve the problems of image blur, processing latency, insufficient matching accuracy, and low efficiency of multi-camera collaborative calibration in the existing rail vision recognition method, and provide a method that can quickly and accurately identify rail target objects. As Figures 1 to 4 shown, it includes: A method for quickly identifying target objects in a rail area based on vision, including the following steps: The target object image sampling step, where an image of the rail area is collected by an on-vehicle camera to obtain a blurred sampling image of the target object and a spatial calibration image; The blurred database construction step, where the database is trained with the collected comparison blurred image data, the blurred image features are extracted and mapped to specific target objects to obtain a blurred database; The sampling image calibration step, where an image calibration model is constructed based on the spatial calibration image and a pre-collected standard calibration image, and the blurred sampling image of the target object is calibrated by the image calibration model to obtain the blurred image data of the target object; The target object quick identification step, where the blurred image data is compared with the blurred image features of the comparison blurred image data in the blurred database to obtain the target object with a mapping relationship and output it as the identified target object.

[0026] In practical applications, on-vehicle cameras are installed at the bottom or side of a rail train. When the train is running, the sampling camera and the calibration camera synchronously collect images of the track area. The sampling camera collects blurred sampling images of the target object, and the calibration camera collects spatial calibration images. The collected image data is transmitted to the on-vehicle computer system in real time. The computer system first processes the collected comparison blurred image data, extracts blurred image features such as edge features, texture features, color features, etc., and establishes a mapping relationship between these features and specific target objects such as foreign objects and track cracks to construct a blurred database. Then, based on the spatial calibration image and the pre-collected standard calibration image (the standard calibration image is a clear image of known track features collected by the calibration camera from a directly overhead perspective in a laboratory environment), an image calibration model is constructed. The blurred sampling image of the target object is calibrated to obtain clear blurred image data. Finally, the calibrated blurred image data is compared with the features in the blurred database, and the corresponding target object is found through a feature matching recognition algorithm, and the recognition result is output.

[0027] Specifically, the on-vehicle cameras include a sampling camera and a calibration camera, which are set as four cameras arranged in parallel, including a left sampling camera, a right sampling camera, a left calibration camera, and a right calibration camera. The left sampling camera and the left calibration camera are arranged on the left front side in the advancing direction of the rail train, and the right sampling camera and the right calibration camera are arranged on the right front side in the advancing direction of the rail train. The left sampling camera and the left calibration camera are respectively mirror-symmetric with the right sampling camera and the right calibration camera along the track center line. The camera sampling areas of the left calibration camera and the right calibration camera are known and standard calibration images from a directly overhead perspective are collected.

[0028] The four on-vehicle cameras are installed in such a way that the left sampling camera and the left calibration camera are arranged on the left front side in the advancing direction of the rail train, the right sampling camera and the right calibration camera are arranged on the right front side, and they are mirror-symmetric along the track center line, as shown in Figure 2 the reference numerals ①, ②, ③, ④ in the figure, where the inner ② and ③ are calibration cameras, and ① and ④ are sampling cameras. The sampling areas of the left calibration camera and the right calibration camera are determined in advance by measurement, for example, set as a rectangular area with a length of 50 cm and a width of 30 cm. Before the train leaves the factory or during regular maintenance, standard calibration images are collected by the calibration camera from a directly overhead perspective and stored in the on-vehicle computer system as a reference for subsequent image calibration.

[0029] Specifically, as shown in Figures 1 to 4 the figure, it also includes a calibration camera detection step. When the sampling areas and the sizes of the collected images of the left calibration camera and the right calibration camera are different, an abnormal calibration camera and a normal calibration camera are determined through a calibration layer correction strategy, and the normal calibration camera or the corrected calibration camera is selected as the calibration camera for collecting spatial calibration images.

[0030] When it is detected that the sampling areas and the sizes of the acquired images of the left calibration camera and the right calibration camera are different, for example, the sampling area of the left calibration camera is 50 cm in length and 30 cm in width, while the sampling area of the right calibration camera is 45 cm in length and 30 cm in width, it indicates that the right calibration camera may have an installation position deviation or a fault. At this time, through the calibration layer correction strategy, the differences between the images acquired by the two calibration cameras and the standard calibration image are compared to determine that the right calibration camera is the abnormal calibration camera and the left calibration camera is the normal calibration camera. The calibration layer correction strategy includes an anomaly detection step, which compares the resolution of the images acquired by the left / right calibration cameras and the geometric distortion rate calculated by computing the uniformity of the distribution of feature points through Harris corner detection (indicating the geometric non-uniformity of the feature points, the lower the uniformity of the feature point distribution, the higher the geometric distortion rate). If the difference exceeds the preset threshold, the resolution deviation > 5% or the distortion rate difference > 10%, it is determined as an abnormal camera; for the abnormal camera, the polynomial geometric correction algorithm is used to establish a coordinate transformation model through more than 20 control points to correct the sampling area deviation. In the subsequent image acquisition process, the left calibration camera is preferentially used to acquire the spatial calibration image. If the left calibration camera also shows abnormalities, correction processing is performed on the right calibration camera, for example, by adjusting its sampling area parameters through a software algorithm to make it consistent with the standard sampling area, and the corrected right calibration camera is used as the calibration camera for acquiring the spatial calibration image.

[0031] Specifically, such as Figures 1 to 4As shown, the sampling image calibration step includes constructing an image calibration model, performing pixel simplification processing on the spatial calibration images collected by the calibration camera, pixelating the spatial calibration images to obtain spatial calibration image pixel values and comparing them with the pixel thresholds preset according to the standard image features. The pixel threshold is determined by calculating the optimal segmentation point through the Otsu algorithm for the pixel value distribution of 1000 standard calibration images; the RGB gradient range value is statistically obtained based on the typical color distribution of the track scene. For example, the RGB range of the track steel is (150±30, 100±20, 50±10), and it is obtained by fitting with the Gaussian mixture model. The pixel offset threshold is 5 pixels (corresponding to an actual track distance of 50 mm, set according to the train detection accuracy requirements); the MSE threshold is 50 (corresponding to the acceptable range of image quality, determined through subjective evaluation experiments). When the pixel value of the spatial calibration image is greater than the pixel threshold, the spatial calibration image is reduced proportionally according to the standard image feature ratio, the reduced spatial calibration image is multi-valued, and the pixel points are divided into several preset RGB thresholds according to the preset RGB gradient range value. The multi-valued spatial calibration image is used to confirm the image boundary of the area corresponding to the standard image features in the calibration camera sampling area according to the threshold classification strategy, and the spatial calibration image is masked according to the image boundary to obtain the masked processing result. The calibration layer of the image calibration model is constructed based on the masked part and the background part of the masked processing result through the fuzzy image calibration rule. The combination method of the masked result and the calibration rule is to extract the contour centroid (x m , y m ) of the masked area and the centroid (x s , y s ) of the corresponding area of the standard image, calculate the angular deviation (rotation correction is only required when the camera optical axis is not vertical), determine the scaling factor s in combination with the aspect ratio difference, and finally complete the initialization of the calibration layer parameters through the affine transformation matrix. The fuzzy image calibration rule calibrates the angle and size of the image input to the image calibration model by comparing the masked processing result with the size of the standard image features; after processing the target object fuzzy sampling image through the calibration layer, it is input to the fuzzy image comparison layer to obtain the fuzzy image data of the target object through image comparison processing. The rule is implemented based on the geometric transformation matrix. The specific steps are as follows: calculate the aspect ratio difference between the minimum circumscribed rectangle of the masked area and the corresponding area of the standard image to determine the scaling factor; calculate the coordinate difference between the centroid of the masked area and the centroid of the standard image to determine the translation amount; complete the angle and size calibration through the affine transformation matrix.

[0032] When constructing an image calibration model, first, pixel simplification processing is performed on the spatial calibration images collected by the calibration camera. Using the clustering downsampling algorithm, the spatial calibration images are divided into 16×16 pixel blocks, and the pixel values within each block are averaged, reducing the image resolution from 1920×1080 to 120×67 pixels, and increasing the processing speed by 30 times. Assuming that the pixel threshold of the standard image features is 100, after pixelating the spatial calibration images, the pixel value of each pixel point is compared with 100. If the pixel value of a certain pixel point is greater than 100, according to the standard image feature ratio, such as the ratio of the pixel width of the track edge in the standard image to the actual width being 1:10, the pixel value of this pixel point is reduced proportionally, for example, reduced to 80% of the original. Then, the downsampled spatial calibration images are multi-valued processed, and the pixel points are divided into 5 RGB thresholds (such as track metal color, fastener gray, ballast brown, foreign object miscellaneous color, background black), and the division basis is the color distribution clustering result of 1000 standard images by the K-means clustering algorithm. Each threshold corresponds to the typical color range of a type of substance. The preset RGB gradient range values are R (100 - 200), G (50 - 150), B (0 - 100), and the pixel points are divided into several preset RGB threshold intervals, R1 (100 - 150), G1 (50 - 100), B1 (0 - 50), etc. Then, according to the threshold classification strategy, the color features and shape features of the spatial calibration images are extracted and compared with the color feature data of the standard image features. The color of the standard track is R = 150, G = 100, B = 50, to narrow the RGB value difference of the pixel points and obtain the RGB thresholds (the R threshold is 140 - 160, the G threshold is 90 - 110, and the B threshold is 40 - 60). The pixel points in the spatial calibration images with RGB values greater than the thresholds are extracted to generate an image boundary, such as the boundary of the track edge. According to the image boundary, the spatial calibration images are masked, with the masked part being the track area and the background part being the non-track area. According to the masked part and the background part of the masking processing result, by comparing the size of the masked part with the size of the standard image features through the fuzzy image calibration rule, the angle and scaling ratio of the image are adjusted to complete the construction of the calibration layer of the image calibration model. Finally, the blurred sampling image of the target object is processed through the calibration layer to correct the angle and size of the image, and then input into the blurred image comparison layer, and the similarity of the image is calculated through the mean square error and peak signal-to-noise ratio for image comparison processing to obtain the blurred image data of the target object.

[0033] Specifically, such as Figures 1 to 4As shown in the figure, the construction of the image calibration model includes the synchronization step of the calibration camera and the sampling camera. An anti-convolution layer is constructed from the spatial calibration images of the calibration camera. The U-Net network structure is adopted, with the input being the low-resolution spatial calibration images (512×256 pixels) of the calibration camera. Through three anti-convolution layers (convolution kernel size 4×4, stride 2), the resolution is restored to 1024×512 pixels. At the same time, in combination with the camera internal parameter matrix (focal length, principal point coordinates), a mapping between pixel coordinates and the actual coordinates of the orbit is established. Based on the known fixed orbit features collected by the calibration camera, a geometric mapping relationship between the image pixel values and the actual orbit features is established. The geometric mapping is realized based on the homography matrix. By setting at least four checkerboard calibration plates on the orbit plane, after collecting the calibration camera images, the homography matrix H is calculated using the findHomography function in OpenCV, and the mapping between pixel coordinates (u, v) and world coordinates (X, Y) is established: The step of correcting deformation through anti-convolution processing includes performing perspective transformation correction on the sampling image through the homography matrix to eliminate the trapezoidal distortion caused by the installation angle, and then inputting it into the anti-convolution layer for resolution restoration. The processing flow is: perspective-corrected image → anti-convolution layer (three layers, with the number of output channels halved for each layer) → high-resolution calibration image.

[0034] Based on the geometric mapping relationship, anti-convolution processing is performed on the blurred sampling images of the target object collected by the sampling camera to correct the image deformation caused by the installation angle difference between the sampling camera and the calibration camera, so that the feature mapping relationship of the calibrated blurred image is consistent with that in the blurred database, completing camera synchronization.

[0035] In the synchronization step of the calibration camera and the sampling camera, an anti-convolution layer is constructed from the spatial calibration images collected by the calibration camera. The anti-convolution layer is a neural network layer that can restore a low-resolution image to a high-resolution image. Based on the known fixed orbit features collected by the calibration camera, such as the seams and bolts of the orbit, a geometric mapping relationship between the image pixel values and the actual orbit features is established. For example, each pixel point corresponds to a 1-mm distance on the actual orbit. Based on this geometric mapping relationship, anti-convolution processing is performed on the blurred sampling images of the target object collected by the sampling camera. Through anti-convolution operations, the image deformation caused by the installation angle difference between the sampling camera and the calibration camera, such as perspective distortion, is corrected. Make the features of the calibrated blurred image, such as the position and shape of the orbit seam, consistent with the feature mapping relationship in the blurred database, completing camera synchronization.

[0036] Specifically, such as Figures 1 to 4As shown, the threshold classification strategy includes extracting the color features and shape features of the spatially calibrated image, setting thresholds for the pixel points of the pixelated spatially calibrated image according to the color feature data of the standard image features, and setting the pixel points of the spatially calibrated image to several fixed RGB values according to the threshold partition. By comparing the color features of the spatially calibrated image and the standard image features, the RGB value difference is reduced to obtain the RGB threshold, and the pixel points in the spatially calibrated image with RGB values greater than the threshold are extracted to generate the image boundary. In the threshold classification strategy, first, an image segmentation algorithm is used to extract the color features and shape features of the spatially calibrated image. The color feature extraction uses the mean-variance statistics in the HSV color space to calculate the color deviation of each pixel point. , where (H s , S s , V s ) is the color mean of the standard track feature. The shape feature extraction extracts the contour through Canny edge detection and calculates the ratio of the contour perimeter to the area, that is, the compactness. When partitioning the threshold, it is divided into 3 RGB threshold intervals in ascending order of D, corresponding to the track main body (D < 10), suspected features (10 ≤ D < 20), and background (D ≥ 20). The interval boundaries are automatically optimized by the Otsu algorithm. The color feature is represented by statistically analyzing the RGB value distribution of each pixel point in the image, and the shape feature is represented by extracting parameters such as the edge contour, perimeter, and area of the image. According to the color feature data of the standard image features, the color distribution of the standard track is that R is between 140 - 160, G is between 90 - 110, and B is between 40 - 60. Thresholds are set for the pixel points of the pixelated spatially calibrated image. For example, the pixel points with R less than 140 or greater than 160 are set as the thresholds for non-track areas. Then, the pixel points of the spatially calibrated image are set to several fixed RGB values according to the threshold partition. For example, the pixel points in the track area are set to R = 150, G = 100, B = 50, and the pixel points in the non-track area are set to R = 255, G = 255, B = 255. By comparing the color features of the spatially calibrated image and the standard image features, the threshold is adjusted to reduce the RGB value difference and obtain the RGB threshold. The R threshold is 145 - 155, the G threshold is 95 - 105, and the B threshold is 45 - 55. The pixel points in the spatially calibrated image with RGB values greater than the threshold are extracted to generate the image boundary, such as the edge boundary of the track.

[0037] Specifically, such as Figures 1 to 4As shown, it also includes a calibration layer correction step, and the calibration layer correction step includes a common-mode error correction step and a differential-mode error correction step. When the error amounts of the left calibration camera and the right calibration camera are the same, the output is a common-mode error; otherwise, the output is a differential-mode error. The quantization index is as follows: for the common-mode error, the difference in pixel offset between the left / right camera image and the standard image ≤ 2 pixels, and the mean difference in color error ≤ 5 (RGB values); for the differential-mode error, if it exceeds the above range, it is determined as a differential-mode error. The common-mode error correction step calculates the error through the overall translation in the spatial calibration images collected by the two calibration cameras; the differential-mode error calculates the error through the differential translation caused by abnormal local orbital features of the calibration cameras.

[0038] In the calibration layer correction step, when the error amounts of the left calibration camera and the right calibration camera are the same, for example, the images collected by both cameras are shifted 5 pixels to the left as a whole, it is judged as a common-mode error; when the image collected by the left calibration camera is shifted 5 pixels to the left while the image collected by the right calibration camera is shifted 3 pixels to the right, it is judged as a differential-mode error. The common-mode error correction step calculates the common-mode error coefficient by analyzing the overall translation amount in the spatial calibration images collected by the two calibration cameras, and generates a reverse compensation matrix translation matrix according to the common-mode error coefficient. The discrimination between the common-mode error and the differential-mode error of the image is realized by the error analysis module of the vehicle-mounted computer. This module monitors the image data of the left and right calibration cameras in real time and calculates the difference in pixel offset between the images collected by the two cameras and the standard calibration image. If the difference is within the preset tolerance range, such as ±2 pixels, it is determined as a common-mode error; if it exceeds the tolerance range, it is determined as a differential-mode error.

[0039] Specifically, as Figures 1 to 4 shown, the common-mode error correction step includes obtaining the spatial calibration images collected by the left calibration camera and the right calibration camera, comparing the overall offset of the spatial calibration images with the pre-stored standard calibration image, calculating the common-mode error coefficient according to the overall offset, and generating a reverse compensation matrix according to the common-mode error coefficient. The reverse compensation matrix is applied to the blurred sampling image of the sampling camera to perform translation or scaling compensation on the blurred sampling image to eliminate the common-mode error.

[0040] Specifically, as Figures 1 to 4As shown, the differential-mode error correction step includes extracting the pixel anomaly features in the local regions of the spatial calibration images of the left and right calibration cameras, and comparing the differences with the corresponding regions of the standard calibration image; determining the differential-mode error regions according to the difference comparison results, performing pixel-level segmentation on the differential-mode error regions, extracting the RGB values and spatial coordinates of the abnormal pixel points, and repairing the abnormal regions by reconstructing the target values of the abnormal pixel points based on the RGB gradient distribution of the neighboring pixels using the bilinear interpolation algorithm. Taking the repaired calibration image as the input, the calibration layer parameters are updated. The neighborhood range of the bilinear interpolation algorithm uses a 3×3 neighborhood, and the weight calculation method is , where d i is the Manhattan distance between the neighboring pixel and the abnormal pixel.

[0041] The common-mode error correction step includes: Error calculation: Through the image registration algorithm, based on the SIFT algorithm of feature points, globally match the spatial calibration images of the left and right calibration cameras with the standard calibration image, and calculate the overall translation amount (Δx, Δy) and the scaling factor k. For example, if the image is translated 10 pixels to the left and 5 pixels down as a whole, and the scaling ratio is 0.9, then the common-mode error coefficient is (Δx = -10, Δy = -5, k = 0.9).

[0042] Compensation matrix generation: Generate a reverse compensation matrix according to the common-mode error coefficient. The matrix contains translation, scaling, and rotation parameters. Among them, the translation compensation is realized through the affine transformation matrix, and the scaling compensation is realized through the scaling matrix. For example, the reverse compensation matrix is: ; Image compensation: Apply the reverse compensation matrix to the blurred sampled image of the sampling camera, and resample the pixel values through the bilinear interpolation algorithm to eliminate the overall offset and scaling error.

[0043] Specifically, the specific implementation of the common-mode error correction is based on the image transformation algorithm library. Taking the example of the overall right offset of the track: Offset calculation: Through the template matching algorithm, match the track centerlines in the left and right calibration images with the standard image, and calculate that the overall right offset is 15 pixels.

[0044] Reverse compensation matrix generation: Generate a translation matrix, , and compensate the sampled image by shifting it 15 pixels to the left.

[0045] Scaling compensation example: If the overall scaling ratio of the image is 0.8 (i.e., the image size is reduced by 20%), then generate a scaling matrix , and enlarge the image to the standard size.

[0046] The differential-mode error correction step includes: Local feature extraction: Use a convolutional neural network (CNN) to extract local features in the left and right calibrated camera images, such as track fasteners, weld seams, etc., compare them with the features in the corresponding areas of the standard calibrated image, and calculate the pixel-level differences, such as RGB value deviation and position offset.

[0047] Error area segmentation: Mark the areas with significant differences as differential mode error areas through the threshold segmentation algorithm. For example, if the RGB value deviation of a certain area from the standard value exceeds 30% and the area is larger than 50×50 pixels, it is determined as an abnormal area.

[0048] Pixel repair: Perform pixel-level segmentation on the abnormal area, extract the coordinates (x, y) and RGB values (r, g, b) of each abnormal pixel point. Based on the RGB gradient distribution of the 8 neighboring pixels, use the bilinear interpolation algorithm to calculate The target value of this pixel:

[0049] where the weight is inversely proportional to the distance from the neighboring pixels to the abnormal pixel. The repaired image is used as new calibration data and input into the calibration layer to update the parameters.

[0050] Specifically, taking the pixel abnormality in the local area of the right calibrated camera as an example for the implementation of differential mode error correction: Abnormality detection: By comparing the right calibrated image with the standard image, it is found that there is a color block abnormality of 50×30 pixels in the track fastener area, with the RGB value being (200, 180, 150), while the standard value is (150, 100, 50).

[0051] Pixel-level segmentation: Use the flood fill algorithm to mark this abnormal area and extract the coordinates and RGB values of all pixels within the area.

[0052] Neighborhood interpolation repair: For each abnormal pixel point, take the normal pixels (non-abnormal area pixels) within its 3×3 neighborhood, and calculate the RGB mean of the neighborhood pixels as the repair value. For example, if the neighborhood mean of a certain abnormal pixel is (145, 98, 52), then this pixel is repaired to this value.

[0053] Parameter update: Input the repaired calibrated image into the calibration layer, and adjust the weight parameters of the deconvolution layer through the backpropagation algorithm to improve the subsequent calibration accuracy.

[0054] Specifically, such as Figures 1 to 4As shown, it also includes an error elimination strategy. The error elimination strategy includes dynamically updating the preset pixel threshold and RGB gradient range threshold in pixel simplification processing according to the historical error data of the calibration camera, setting a feedback loop in the calibration layer, comparing the image after each correction with the standard calibration image for a second time, and adjusting the deconvolution layer parameters or compensation matrix according to the residual error until the error is lower than the preset threshold. When the error of a certain side calibration camera continuously exceeds the tolerance range, it is marked as an abnormal camera, and the system switches to the standby calibration camera or uses the calibration data of the mirror symmetric side to complete the calibration. The trigger mechanism is that the MSE of three consecutive calibrations of a certain side camera > 80 and the difference in MSE from the opposite side camera > 30, that is, it is determined as "continuously exceeding the tolerance", and automatically switches to the horizontally flipped data of the opposite side image. The residual error is calculated using the mean square error (MSE), and the formula is , where I i is the pixel value of the corrected image, and I std,i is the corresponding pixel value of the standard image. The threshold is set to 30 (this threshold is determined by statistics of 1000 calibration data). Its implementation includes: Dynamic threshold update: The vehicle system maintains an error history database that records the residual error data of the past 100 calibrations. When the calibration error in a certain area is detected to exceed the preset threshold continuously for 5 times, such as pixel offset > 3 pixels or RGB deviation > 20%, the preset pixel threshold in pixel simplification processing is automatically adjusted, such as from 100 to 120, and the RGB gradient range threshold, such as the R range from 100 - 200 to 80 - 220, to adapt to the image differences caused by light changes or camera aging.

[0055] Feedback loop correction: Set a closed-loop feedback mechanism in the calibration layer: The image after each calibration is compared with the standard calibration image for a second time, and the mean square error (MSE) is calculated. If MSE > 50, then trigger parameter adjustment: If it is a geometric deformation error, such as tilt, adjust the convolution kernel parameters of the deconvolution layer; If it is a brightness / color error, adjust the brightness scaling factor or color balance coefficient of the compensation matrix.

[0056] Repeat the comparison - adjustment process until MSE < 30 or the maximum number of iterations reaches 10 times.

[0057] Abnormal camera switching: When the error of a certain side calibration camera (such as the left calibration camera) continuously exceeds the tolerance range for 10 times, the system marks this camera as "abnormal", and automatically switches to the standby calibration camera or enables the mirror symmetry strategy: Use the image data of the right calibration camera to generate the left - view calibration image through the horizontal flipping algorithm to replace the abnormal camera data to complete the calibration.

[0058] The basic features, principles and advantages of the present invention have been shown and described above. It should be noted that the present invention is not limited by the above embodiments, which are only partial embodiments. Without departing from the spirit and scope of the present invention, several improvements and supplements made are regarded as the protection scope of the present invention.

Claims

1. A method for rapid identification of objects within an orbital region based on vision, characterized in that, It includes the following steps: Target object image sampling step: Collect images of the track area through an on-vehicle camera to obtain a blurred sampling image of the target object and a spatial calibration image; Blurred database construction step: Train the database with the collected comparison blurred image data, extract the blurred image features and establish a mapping relationship with specific target objects to obtain a blurred database; Sampling image calibration step: Construct an image calibration model based on the spatial calibration image and a pre-collected standard calibration image, and calibrate the blurred sampling image of the target object through the image calibration model to obtain the blurred image data of the target object; Target object rapid recognition step: Compare the blurred image data with the blurred image features of the comparison blurred image data in the blurred database to obtain the target object with a mapping relationship and output it as the recognized target object.

2. The rapid recognition method of an object within an orbital region based on vision according to claim 1, characterized in that The on-vehicle camera includes a sampling camera and a calibration camera, which are set as four cameras arranged in parallel, including a left sampling camera, a right sampling camera, a left calibration camera and a right calibration camera. The left sampling camera and the left calibration camera are arranged on the left front side of the advancing direction of the rail train, and the right sampling camera and the right calibration camera are arranged on the right front side of the advancing direction of the rail train. The left sampling camera and the left calibration camera are symmetrically arranged along the track center line with the right sampling camera and the right calibration camera respectively. The camera sampling areas of the left calibration camera and the right calibration camera are of known size and have collected standard calibration images with a directly overhead view.

3. The fast recognition method for the target object within the track area based on vision according to claim 2, wherein It also includes a calibration camera detection step. When the sampling areas and the collected image sizes of the left calibration camera and the right calibration camera are different, determine the abnormal calibration camera and the normal calibration camera through the calibration layer correction strategy, and select the normal calibration camera or the corrected calibration camera as the calibration camera for collecting the spatial calibration image.

4. A method for quickly identifying a target object within an orbital region based on vision according to claim 1, characterized in that The sampling image calibration step includes: constructing an image calibration model, performing pixel simplification processing on the spatial calibration image collected by the calibration camera, pixelizing the spatial calibration image to obtain the pixel values of the spatial calibration image and comparing them with the pixel threshold preset according to the standard image features. When the pixel value of the spatial calibration image is greater than the pixel threshold, perform an equi-ratio pixel value reduction on the spatial calibration image according to the standard image feature ratio, perform multi-valued processing on the reduced spatial calibration image, divide the pixel points into several preset RGB thresholds according to the preset RGB gradient range value, confirm the image boundary of the area corresponding to the standard image features of the calibration camera sampling area in the image according to the threshold classification strategy for the multi-valued processed spatial calibration image, and perform masking processing on the spatial calibration image according to the image boundary to obtain the masking processing result. Construct the calibration layer of the image calibration model according to the masked part and the background part of the masking processing result through the blurred image calibration rule. The blurred image calibration rule calibrates the angle and size of the image input into the image calibration model by comparing the masking processing result with the size of the standard image features; after processing the blurred sampling image of the target object through the calibration layer, input it into the blurred image comparison layer to obtain the blurred image data of the target object through image comparison processing.

5. A method for rapid recognition of an object within an orbital region based on vision according to claim 4, characterized in that, The construction of the image calibration model includes the steps of synchronizing the calibration camera and the sampling camera. An anti-convolution layer is constructed through the spatial calibration image of the calibration camera. According to the known fixed track features collected by the calibration camera, a geometric mapping relationship between the image pixel values and the actual track features is established; Based on the geometric mapping relationship, deconvolution processing is performed on the blurred sampling image of the target object collected by the sampling camera to correct the image deformation caused by the installation angle difference between the sampling camera and the calibration camera, so that the calibrated blurred image features are consistent with the feature mapping relationship in the blurred database, and the camera synchronization is completed.

6. The fast recognition method for an object within an orbital region based on vision according to claim 4, wherein The threshold classification strategy includes extracting the color features and shape features of the spatial calibration image, setting thresholds for the pixel points of the pixelated spatial calibration image according to the color feature data of the standard image features, and setting the pixel points of the spatial calibration image to several fixed RGB values according to the threshold partition. By comparing the color features of the spatial calibration image and the standard image features, the RGB threshold is obtained by reducing the RGB value difference. The pixel points in the spatial calibration image with RGB values greater than the threshold are extracted to generate an image boundary.

7. A method for quickly identifying an object within an orbital region based on vision according to claim 2, wherein It also includes a calibration layer correction step. The calibration layer correction step includes a common-mode error correction step and a differential-mode error correction step. When the error amounts of the left calibration camera and the right calibration camera are the same, the output is a common-mode error, otherwise the output is a differential-mode error. The common-mode error correction step calculates the error through the overall translation in the spatial calibration images collected by the two calibration cameras; the differential-mode error is calculated through the differential translation caused by the abnormal local track features of the calibration camera.

8. A method for quickly identifying an object within an orbital region based on vision according to claim 7, characterized in that, The common-mode error correction step includes obtaining the spatial calibration images collected by the left calibration camera and the right calibration camera, comparing the overall offset of the spatial calibration images with the pre-stored standard calibration image, calculating the common-mode error coefficient according to the overall offset, and generating a reverse compensation matrix according to the common-mode error coefficient. The reverse compensation matrix is applied to the blurred sampling image of the sampling camera to perform translation or scaling compensation on the blurred sampling image to eliminate the common-mode error.

9. A method for quickly identifying an object within an orbital region based on vision according to claim 7, characterized in that The differential-mode error correction step includes extracting the pixel abnormal features in the local areas of the spatial calibration images of the left calibration camera and the right calibration camera, and comparing them with the corresponding areas of the standard calibration image; determining the differential-mode error area according to the comparison result of the difference. The differential-mode error area is segmented at the pixel level, the RGB values and spatial coordinates of the abnormal pixel points are extracted, and the target values of the abnormal pixel points are reconstructed based on the RGB gradient distribution of the neighboring pixels by using the bilinear interpolation algorithm to repair the abnormal area. The repaired calibration image is used as the input to update the calibration layer parameters.

10. A method for quickly identifying an object within an orbital region based on vision according to claim 7, characterized in that It also includes an error elimination strategy, which includes dynamically updating the preset pixel threshold and the RGB gradient range threshold in pixel simplification processing according to the historical error data of the calibration camera, setting a feedback loop in the calibration layer, performing a secondary comparison between the image after each correction and the standard calibration image, and adjusting the deconvolution layer parameters or the compensation matrix according to the residual error until the error is lower than the preset threshold. When the error of a certain side calibration camera continuously exceeds the tolerance range, it is marked as an abnormal camera, and the standby calibration camera is switched to or the calibration data of the mirror symmetry side is used to complete the correction.

Citation Information

Patent Citations

  • Subway track foreign matter intelligent identification method

    CN112989931A

  • Rail transit structure crack dynamic detection method, device and equipment and storage medium

    CN118333980A

  • Track foreign matter identification method and system, electronic equipment and medium

    CN118736527A

  • Remote rail identification and obstacle detection method and system

    CN119625679A

  • Track foreign body detection method and system

    CN119763073A

Cited By

  • Image feature calibration method based on motion offset

    CN120655549A

  • Detection model generation method, monocular 3D target detection method and monocular 3D target detection device

    CN120877273A