Track alignment monitoring method and device based on binocular vision measurement

By using binocular visual measurement technology on the track, combined with encoded image markers and deep learning algorithms, high-precision three-dimensional linear monitoring of orbits is achieved, solving the problems of low equipment reliability and measurement accuracy in traditional methods, and has broad application prospects.

CN120088395AActive Publication Date: 2025-06-03SOUTHEAST UNIV

Patent Information

Application Number
CN202510003349.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-02
Publication Date
2025-06-03
Estimated Expiration
2045-01-02

AI Technical Summary

Technical Problem

The traditional orbital three-dimensional linear measurement method has reduced equipment reliability and measurement accuracy due to vibration and severe weather during long-term use, and the accuracy is not high.

Method used

Using a binocular visual measurement method, by dividing the measurement track into multiple segments and fixing markers marked with encoded images on each segment, a binocular camera is used to take photos on multiple measurement points, combining the adversarial neural network and YOLOv5 for image preprocessing and object detection, the encoded images are extracted and three-dimensional coordinate reconstruction is carried out, and the three-dimensional linear curve of the track is optimized through beam adjustment method.

Benefits of technology

Significantly improves the accuracy of track line monitoring, reduces equipment complexity, avoids additional installation and maintenance costs, and maintains high accuracy and reliability under a variety of environmental conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120088395A_ABST
    Figure CN120088395A_ABST
Patent Text Reader

Abstract

The invention discloses a track alignment monitoring method and device based on binocular vision measurement, and the method comprises the steps: fixing a marker marked with a coded image at each track segment, enabling a binocular camera to move along a measurement track, and carrying out the shooting at different measurement points; carrying out image preprocessing on the shot photo sequence, carrying out target detection, and extracting a coded image; calibrating internal parameters of the binocular camera and external parameters of each measuring point; obtaining a three-dimensional coordinate of each marker in a local binocular coordinate system according to the internal parameters of the binocular camera, and converting the three-dimensional coordinate into a three-dimensional coordinate of each marker in a global coordinate system by adopting the external parameters of the binocular camera; and performing beam adjustment method optimization on all three-dimensional coordinates of the markers with consistent unique mark numbers under the global coordinate system to obtain optimized three-dimensional coordinates of all the markers under the global coordinate system, and gathering according to a sequence to obtain a three-dimensional linear curve of the measurement orbit. The method is high in precision and free of environmental interference.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to computer vision technology, and particularly to a method and device for monitoring track alignment based on binocular vision measurement. Background Art

[0002] Under the action of factors such as high-speed train rolling, starting and braking, and external temperature changes, large longitudinal stresses will be generated inside the track, causing the rail to move longitudinally along the sleeper or the track frame along the top surface of the ballast bed. This phenomenon is called rail creep or longitudinal displacement of the rail. The longitudinal displacement of the rail causes the stress redistribution of the seamless line rail, and even induces buckling and running off the track, thus affecting the driving safety of high-speed trains. Therefore, the accurate monitoring of the three-dimensional track alignment is of great significance for ensuring the safe operation of high-speed railways.

[0003] Traditional methods for measuring the three-dimensional track alignment mainly use electrical displacement sensors and acceleration sensors, and usually require the installation of displacement sensing devices on railway facilities such as tracks or sleepers, such as magnetostrictive displacement sensors, laser displacement sensors, fiber optic sensors, etc. The change of the track alignment is obtained through the change of the displacement at the measurement points. Such methods are easy to operate and have a fast calculation speed, but due to the long-term influence of track vibration and bad weather on the sensing devices, the reliability and measurement accuracy of the devices decrease, and the accuracy is not high. Summary of the Invention

[0004] Aiming at the problems existing in the prior art, the purpose of the present invention is to provide a method and device for monitoring track alignment based on binocular vision measurement with higher accuracy.

[0005] In order to achieve the above invention purpose, the present invention provides the following technical solutions:

[0006] A method for monitoring track alignment based on binocular vision measurement, comprising the following steps:

[0007] S1. Divide the measurement track into several track segments, and fix markers with coded images on the web of each track segment. Among them, the coded images on the markers of each track segment are different, and each coded image is a code of the unique label of the track segment;

[0008] S2. Move the binocular camera with a mobile platform base along the measurement track and take pictures at different measurement points to obtain a sequence of photos containing markers;

[0009] S3. Perform image preprocessing on the sequence of photos using an adversarial neural network to obtain a sequence of high-quality images;

[0010] S4. For each photo in the sequence of photos, perform object detection using YOLOv5 to extract the coded image area therein;

[0011] S5. Preprocess each extracted encoded image region, and extract the encoded image from each preprocessed encoded image region through an edge extraction operator, and parse the unique label represented by each encoded image;

[0012] S6. Calibrate the internal parameters of the binocular camera through a calibration board;

[0013] S7. Solve the relative pose of the binocular camera through epipolar geometry constraints to obtain the external parameters of the binocular camera at each measurement point;

[0014] S8. Locate the center point of each encoded image in each photo, and based on the located center point, reconstruct the coordinates of each encoded image according to the internal parameters of the binocular camera to obtain the three-dimensional coordinates of the corresponding marker of each encoded image in the local binocular coordinate system;

[0015] S9. Convert the three-dimensional coordinates of each marker in the local binocular coordinate system to the three-dimensional coordinates of each marker in the global coordinate system by using the external parameters of the binocular camera;

[0016] S10. Optimize all the three-dimensional coordinates of the markers with the same unique label in the global coordinate system by using the bundle adjustment method to obtain the optimized three-dimensional coordinates of all the markers in the global coordinate system;

[0017] S11. Assemble the optimized three-dimensional coordinates of all the markers in the global coordinate system according to the order of the track segments where they are located to obtain the three-dimensional linear curve of the measurement track.

[0018] Further, the encoded image consists of three concentric circles, and the ratio relationship of their radii is: 5:12:20. The three concentric circles are divided into 15 equal parts to form a 15-bit encoding band. Each encoding band is an encoding area. If the encoding area is black, it represents the binary number 0, and if it is white, it represents the binary number 1. Finally, a 15-bit binary number is formed as the unique label for each track segment.

[0019] Further, step S3 specifically includes:

[0020] S301. Filter and denoise each photo in the photo sequence;

[0021] S302. Enhance the brightness and contrast of the photos with low light or overexposure, and for the photos with strong light and reflection problems, use image dehazing or dereflection algorithms for repair;

[0022] S303. Use the trained adversarial neural network to preprocess each photo to obtain high-quality photos, where the adversarial neural network is trained with a number of low-quality images and corresponding high-quality image labels.

[0023] Further, step S4 specifically includes:

[0024] S401. Denoise, adjust illumination, and perform geometric correction on each photo in the photo sequence;

[0025] S402. Load the trained YOLOv5 model, perform object detection on each photo in the photo sequence, and extract the encoded image regions and their confidence levels therein; wherein, when training the YOLOv5 model, pictures collected by the track under different scenarios, different angles, and different distances are used;

[0026] S403. Retain the encoded image regions with confidence levels greater than the preset threshold, and delete the others;

[0027] S404. If the position of the detected encoded image region deviates from the expected position of the track region, it is determined as a false detection; if the number of detected encoded image regions is less than the preset number, it is determined as a missed detection;

[0028] S405. If a missed detection is found, adjust the position of the binocular camera or re-adjust the exposure parameters of the binocular camera.

[0029] Further, step S5 specifically includes:

[0030] S501. Preprocess the extracted encoded image regions, including two-dimensional Gaussian filtering and binarization;

[0031] S502. Use the Canny operator to extract edges from the preprocessed encoded image regions to obtain the encoded images in the encoded image regions;

[0032] S503. Decode the encoded images to parse out the unique labels represented by the encoded images.

[0033] Further, step S6 specifically includes: calibrating the internal parameters of the binocular camera using a checkerboard calibration plate.

[0034] Further, step S8 specifically includes:

[0035] S801. Use the Gaussian curve fitting method to perform sub-pixel localization of the center points of each encoded image to obtain the coordinates of the center points of each encoded image on the two-dimensional image plane;

[0036] S802. According to the internal parameters of the binocular camera, reconstruct the three-dimensional coordinates of the center points of each encoded image on the two-dimensional image plane as the three-dimensional coordinates of each encoded image in the local binocular coordinate system.

[0037] A computer device includes a memory, a processor, and a computer program stored on the memory and executable on the processor. The processor executes the computer program to implement the above method.

[0038] A computer-readable storage medium stores a computer program / instructions, and the computer program / instructions implement the above method when executed by a processor.

[0039] A computer program product includes a computer program / instructions, and the computer program / instructions implement the above method when executed by a processor.

[0040] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0041] (1) The present invention adopts a three-dimensional reconstruction method based on coded markers and close-range photogrammetry to accurately measure the three-dimensional spatial coordinates of the markers. The present invention also uses the bundle adjustment method to optimize the coordinate reconstruction results of multiple measurement points, significantly improving the accuracy of the three-dimensional reconstruction world coordinate system.

[0042] (2) By moving a set of binocular camera devices at multiple measurement points, the present invention simulates multi-camera group measurement, significantly reducing the complexity of the equipment required for track measurement.

[0043] (3) The present invention adopts a non-contact measurement method, eliminating the need to install complex vibration sensors and avoiding additional equipment installation and maintenance costs.

[0044] (4) The present invention has low requirements for the background environment and has a wide range of applications. It can accurately identify under complex environmental conditions such as low light, strong light, rain, and snow, thereby enhancing the reliability and accuracy of the system in various environments and having broad application prospects. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] Figure 1 is a schematic flow chart of the track alignment monitoring method based on binocular vision measurement provided by the present invention;

[0046] Figure 2 is a schematic diagram of the pasting of track surface markers provided by the present invention;

[0047] Figure 3 is a schematic diagram of a binocular camera provided by the present invention;

[0048] Figure 4 is a schematic structural diagram of the computer device provided by the present invention;

[0049] In the figure, the measurement track is 1, the sleeper is 2, the marker is 3, the industrial camera is 4, the binocular camera bracket is 5, and the mobile platform is 6. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0050] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention.

[0051] Embodiment 1

[0052] The embodiment of the present invention provides a method for monitoring the track alignment based on binocular vision measurement, as Figure 1 shown, including the following steps:

[0053] S1. Divide the measurement track into several track segments, and fix the markers with coded images at the web of each track segment.

[0054] Among them, as Figure 2 shown, the measurement track 1 is divided into several segments. The measurement track 1 is arranged on the sleeper 2. The coded images on the markers 3 of each track segment are different, and each coded image is the code of the unique label of the track segment. Specifically, the coded image consists of three concentric circles, and the ratio relationship of their radii is: 5:12:20. The three concentric circles are divided into 15 equal parts to form a 15-bit coded band. Each coded band is a coding area. If the coding area is black, it represents the binary number 0. If it is white, it represents the binary number 1. Finally, a 15-bit binary number is formed as the unique label of each track segment. Since there are 15 coded bands in total, at most 15 binary numbers can be read. Due to the arrangement combination method and the uniqueness of the code, and removing the two cases where the coded band is all 0 or 1 (for the recognition of the marker points and the robustness of decoding), the capacity of this coded image is 2190, which can meet the needs of most camera group external parameter calibration and key point measurement.

[0055] S2. Move the binocular camera with a mobile platform base along the measurement track and take pictures at different measurement points to obtain a sequence of photos containing markers.

[0056] Set several measurement points along the track, and take pictures of the markers at each specified measurement point. A set of binocular cameras used is as Figure 3 shown. Two industrial cameras 4 are supported on the mobile platform 6 through the binocular camera bracket 5. The left and right view images of the track are simultaneously collected by the binocular camera to ensure that all the markers on the track are covered.

[0057] S3. Use an adversarial neural network to preprocess the sequence of photos to obtain a high-quality image sequence.

[0058] This step specifically includes:

[0059] S301. Filter and denoise each photo in the sequence of photos. Specifically, denoising algorithms such as Gaussian filtering and median filtering can be used to remove the noise generated under rainy, snowy or low-light conditions;

[0060] S302. Enhance the brightness and contrast of photos with low light or overexposure. For photos with strong light and reflection problems, use image dehazing or dereflection algorithms for repair;

[0061] S303. Use a trained adversarial neural network to perform image preprocessing on each photo to obtain high-quality photos. Among them, the adversarial neural network is trained with a number of low-quality images and corresponding high-quality image labels. The adversarial neural network can specifically be CycleGAN (suitable for unpaired data environments when paired low-quality and high-quality images cannot be obtained during training) or Pix2Pix (if corresponding high-quality images can be obtained, use Pix2Pix for end-to-end conversion of images), and process different low-quality images (such as low light, strong light, rain and snow, etc.). The generator is a generator based on the U-Net or ResNet structure, adding skip connections to retain details in the encoded point area, and the discriminator uses PatchGAN to evaluate the local quality of the generated image to ensure the realism and details of the enhanced image.

[0062] S4. For each photo in the photo sequence, use YOLOv5 for object detection and extract the encoded image area therein.

[0063] This step specifically includes:

[0064] S401. Denoise, adjust the lighting, and perform geometric correction on each photo in the photo sequence to ensure that the image quality meets the requirements when input into the YOLOv5 model;

[0065] S402. Load the trained YOLOv5 model, perform object detection on each photo in the photo sequence, and extract the encoded image area and its confidence level therein; among them, when training the YOLOv5 model, pictures collected at different scenes, different angles, and different distances are used, including pictures collected in multiple scenes such as day, night, strong light, shadow, rain and snow, and at different angles and distances (such as changes in camera installation height and tilt angle). Then use the test set to evaluate the performance of the trained model, and the indicators include mAP (mean Average Precision), Recall, Precision, etc. According to the evaluation results, further adjust the data augmentation strategy and model hyperparameters (learning rate, batch size, etc.) to ensure that the model can detect and accurately locate the track encoding points. The size of the photo needs to be adjusted according to the default input size of YOLO v5 (640×640), and output the category, bounding box coordinates, and detection confidence level of each detected encoded image area.

[0066] S403. Select and retain the encoded image areas with a confidence level greater than a preset threshold (such as 0.7), and delete the others;

[0067] S404. If the detected position of the encoded image region deviates from the expected position of the track region, it is determined as a false detection. If the number of detected encoded image regions is less than the preset number, it is judged as a missed detection.

[0068] S405. If a missed detection is found, adjust the position of the binocular camera or re-adjust the exposure parameters of the binocular camera. For example, if some regions are missed due to occlusion or perspective problems, adjust the pitch angle, roll angle, etc. of the camera position.

[0069] S5. Preprocess each extracted encoded image region, and extract the encoded image from each preprocessed encoded image region through an edge extraction operator, and parse the unique label represented by each encoded image.

[0070] This step specifically includes:

[0071] S501. Preprocess the extracted encoded image region, including two-dimensional Gaussian filtering and binarization; among them, edge extraction directly on the binarized image is easily affected by the binarization threshold, so binarization is only used in positioning, and integer-pixel edge extraction is performed on the image after two-dimensional Gaussian filtering.

[0072]

[0073] In the formula, G(u, v) represents the Gaussian filtering function, (u, v) represents the coordinates of the image, and σ is the standard deviation.

[0074] S502. Use the Canny operator to perform edge extraction on the preprocessed encoded image region to obtain the encoded image in the encoded image region.

[0075] The edge in integer-pixel edge extraction refers to the position where the gray level in the image is discontinuous, with a large gray level gradient. Commonly used edge detection operators include Canny, Robert, Sobel, and Prewitt, etc. The present invention uses the Canny operator that is insensitive to noise and can generate single-pixel edges, performs convolution calculation on the image to enhance the edge, and obtains the magnitude and direction of the gray level gradient:

[0076]

[0077] Among them, f(u, v) is the pixel gray level of the coordinate (u, v), is the convolution operation symbol, G x and G y are the magnitudes of the gray level gradient, H x and H y are respectively:

[0078]

[0079] Gradient magnitude and direction θ are as follows:

[0080]

[0081] Usually, multiple pixels' edges are obtained through the previous calculation. By non-maximum suppression of the gradient direction, the local maximum gradient is retained while other gradients are set to zero, and finally the position with the sharpest change in gray-scale gradient is obtained. Finally, a double-threshold method is adopted. The edge points with gray-scale gradient greater than the high threshold are regarded as strong edge points, and the edge points with gray-scale gradient less than the high threshold but greater than the low threshold are regarded as weak edge points, completing edge extraction. The encoded image in the encoded image region can be obtained through edge extraction.

[0082] S503. Decode the encoded image and parse out the unique label represented by the encoded image.

[0083] Specifically, decoding can be performed according to the correspondence between the encoded image and the unique label.

[0084] S6. Calibrate the internal parameters of the binocular camera through a calibration board.

[0085] Specifically, a checkerboard calibration board can be used to calibrate the internal parameters of the binocular camera.

[0086] S7. Solve the relative pose of the binocular camera through epipolar geometry constraints to obtain the external parameters of the binocular camera at each measurement point.

[0087] At each measurement point, let S 1 and S 2 represent any two consecutive imaging planes. P is a landmark point in space, and p 1 (u 1 , v 1 ) and p 2 (u 2 , v 2 ) are the imaging points of point P. O c1 and O c2 represent the optical center positions of the two cameras. The plane determined by P, O c1 and O c2 is the epipolar plane. The line connecting O c1 and O c2 is the baseline. The intersection points e 1 and e 2 of the baseline with planes S 1 and S 2 are the poles. The intersection lines l 1 and l 2 of the epipolar plane with planes S 1 and S 2is the epipolar line. According to the pinhole camera model, the conversion relationship between the pixel coordinate system and the camera coordinate system can be expressed as:

[0088]

[0089] where K is the internal parameter matrix of the camera; c x and c y are the principal points of the image (the intersection of the optical axis and the imaging plane); f x and f y are the equivalent focal lengths, that is, the ratio of the focal length to the horizontal and vertical dimensions of a single pixel; f s is the axis tilt parameter, which is related to the distance between the imaging planes, the horizontal pixel size, and the true angle of the pixel arrangement on the sensor plane. Combining the above formula, we have:

[0090] s 1 p 1 = KP s 2 p 2 = K(RP + t)

[0091] where s 1 and s 2 are the depths of point P in the two images. After transforming the above formula, we can get the following equation:

[0092] s 2 x 2 = s 1 Rx 1 + t

[0093] where x 1 = K -1 p 1 , x 2 = K -1 p c .

[0094] Define the skew-symmetric symbol ^ and let:

[0095]

[0096] Left-multiply the transformed equation by to get:

[0097]

[0098] According to the perpendicular relationship, the left side of the above formula is equal to zero. Therefore, the equation of the epipolar constraint can be obtained:

[0099]

[0100] Substitute p 2 and p 2 into the above formula to get another form of the epipolar constraint:

[0101]

[0102] Among them, t ∧ R is the essential matrix E, and F is the fundamental matrix. Thus, the solution of the relative pose of the cameras can be transformed into the solution of the essential matrix E or the fundamental matrix F. The pose of the cameras is solved by using 6-8 pairs of well-matched fiducial points in two images. That is, the external parameters R and t are obtained.

[0103] S8. Locate the center point of each encoded image in each photo, and based on the located center point, reconstruct the coordinates of each encoded image according to the internal parameters of the binocular cameras, so as to obtain the three-dimensional coordinates of the corresponding markers of each encoded image in the local binocular coordinate system.

[0104] S801. Use the Gaussian curve fitting method to perform sub-pixel localization of the center point of each encoded image, and obtain the coordinates of the center point of each encoded image on the two-dimensional image plane.

[0105] During sub-pixel localization, each square represents a pixel point. Let P 4 (u 0 , v 0 ) be the gray gradient direction of the integer pixel edge. P 1 , P 7 , P 2 , P 6 , P 3 , P 5 are the points at distances of 3, 2, and 1 pixel from the P 4 point in the gradient direction and the opposite direction of the P point respectively. Their gray values are obtained by bilinear interpolation of the gray values of the neighboring pixels and are denoted as I i . Use the difference formula to calculate the gray differences of P 2 , P 3 , P 4 , P 5 , P 6 :

[0106]

[0107] Use one-dimensional Gaussian curve fitting to calculate the obtained gray differences. The objective function is as follows:

[0108]

[0109] Take the logarithm on both sides, and denote g = lny, then there is:

[0110] g = ax 2 + bx + c

[0111] According to the square aperture sampling theorem, the pixel gray difference is:

[0112]

[0113] where \(g(n)=\ln D\) n (\(n = 1, 2, 3, 4, 5\)), and the corresponding \(x\) values are -2, 1, 0, 1, 2 respectively. Using the least squares method to solve for the three unknowns \(a\), \(b\), and \(c\), taking the extreme value of the previous equation gives:

[0114]

[0115] where

[0116]

[0117] Denote that on the gradient direction \(\theta\), the offset \(\delta = x\) of the sub-pixel boundary relative to the integer-pixel boundary, then the sub-pixel boundary coordinates can be expressed as:

[0118]

[0119] u ′ 0 , v ′ 0 represent the sub-pixel boundary coordinates, and \(u\) 0 , v 0 represent the integer-pixel boundary coordinates.

[0120] Using the calculated sub-pixel boundaries for fitting, since the image of a circle after perspective projection is usually an ellipse, the objective function is generally selected as the general equation of the ellipse:

[0121] u 2 + Auv + Bv 2 + Cu + Dv + E = 0

[0122] There are 5 unknown coefficients A, B, C, D, and E. Therefore, the number of sub-pixel boundaries must be more than 5. Generally, the number of boundaries will be much greater than 5. So, the least squares method can be used to solve the above coefficients, and finally the center coordinates \((u\) c , v c ) are as follows:

[0123]

[0124] S802. According to the internal parameters of the binocular camera, reconstruct the three-dimensional coordinates of the center point of each encoded image on the two-dimensional image plane as the three-dimensional coordinates of each encoded image in the local binocular coordinate system.

[0125] Specifically, the three-dimensional coordinates \((X\) C , Y C , ZC ) is:

[0126]

[0127] In the formula, K is the internal parameter matrix of the camera, and Z C is the depth matrix.

[0128] S9. Convert the three-dimensional coordinates of each marker in the local binocular coordinate system into the three-dimensional coordinates of each marker in the global coordinate system by using the external parameters of the binocular camera.

[0129] Specifically, the three-dimensional coordinates (X W , Y W , Z W ) of the marker in the global coordinate system are:

[0130]

[0131] S10. Optimize all the three-dimensional coordinates of the markers with the same unique label in the global coordinate system by using the bundle adjustment method to obtain the optimized three-dimensional coordinates of all the markers in the global coordinate system.

[0132] List the optimization objective of the bundle adjustment method according to the collinearity condition equation as follows:

[0133]

[0134] In the formula, is the image coordinate of the center point of the encoded image of the i-th marker at the j-th measurement point, is the projected image coordinate of the center point of the corresponding encoded marker image, and K, D, R j , T j are the internal parameters of the camera, the lens distortion, and the external parameters at the j-th measurement point respectively, represents the three-dimensional space coordinate of the i-th marker at the j-th measurement point.

[0135] Perform adjustment processing uniformly in the entire area, and optimize and solve the internal parameters of the camera, the lens distortion, the external parameters of the camera coordinate system relative to the world coordinate system at different measurement points, and the spatial point coordinates.

[0136] S11. Assemble the optimized three-dimensional coordinates of all the markers in the global coordinate system according to the order of the segments of the track where they are located to obtain the three-dimensional linear curve of the measurement track.

[0137] Embodiment 2

[0138] Figure 4 is a schematic structural diagram of a computer device provided by an embodiment of the present invention. The embodiment of the present invention provides services for implementing the method of the above Embodiment 1 of the present invention. As Figure 4As shown, the device may include: a memory 301 storing computer-executable programs; a processor 302 coupled to the memory 301; the processor 302 invoking the computer-executable programs stored in the memory 301 for performing the steps in the method described in Embodiment 1.

[0139] The memory 301 may include a computer system-readable medium in the form of volatile memory, such as random access memory (RAM) and / or cache memory. The device may further include other removable / non-removable, volatile / non-volatile computer system storage media. By way of example only, the memory 301 may be used for reading and writing non-removable, non-volatile magnetic media (commonly referred to as a "hard disk drive"). Programs / utilities having a set (at least one) of program modules may be stored, for example, in the memory 301, and such program modules include but are not limited to an operating system, one or more application programs, other program modules, and program data, and the implementation of a network environment may be included in each or some combination of these examples. The computer-executable programs of the program modules generally perform the functions and / or methods in the embodiments described in the present invention.

[0140] The processor 302 executes various functional applications and data processing by running the programs stored in the memory 301, such as implementing the method provided in Embodiment 1 of the present invention.

[0141] The code of the computer-executable programs may be written in one or more programming languages or combinations thereof, including object-oriented programming languages such as Java, Smalltalk, C++, and also including conventional procedural programming languages such as the "C" language or similar programming languages.

[0142] Embodiment 3

[0143] The embodiment of the present invention provides a storage medium containing computer-executable programs, and the computer-executable programs are used for performing the method of Embodiment 1 when executed by a computer processor.

[0144] The storage medium according to an embodiment of the present invention may adopt any combination of one or more computer-readable media. The computer-readable media may be a computer-readable signal medium or a computer-readable storage medium. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (non-exhaustive list) of the computer-readable storage medium include: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this document, the computer-readable storage medium may be any tangible medium that contains or stores a program, and this program may be used by or in combination with an instruction execution system, apparatus, or device.

[0145] Of course, for a storage medium provided by an embodiment of the present invention that contains a computer-executable program, the computer-executable program is not limited to the above method operations, and may also execute related operations in the methods provided by any embodiment of the present invention.

[0146] Embodiment 4

[0147] The embodiment of the present invention further provides a computer program product, such as an app on a mobile phone or a tablet, an installation program on a computer, etc. The product includes a computer program / instructions, and when the computer program / instructions are executed by a processor, the method described in Embodiment 1 is implemented. The code of the computer-executable program for performing the operations of the present invention may be written in one or more programming languages or a combination thereof. The programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, executed as an independent software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (for example, by using an Internet service provider to connect through the Internet).

[0148] It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations.

[0149] It should be understood that the above embodiments and the descriptions in the specification are only the principles, main features and advantages of the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and these changes and improvements all fall within the protection scope of the present invention.

Claims

1. A track line shape monitoring method based on binocular vision measurement, characterized in that: The steps include: S1. Divide the measurement track into several track segments, and fix a marker marked with a coded image on the rail waist of each track segment, wherein the coded image on the marker of each track segment is different, and each coded image is a code of a unique label of the track segment; S2, moving the binocular camera with a mobile platform base along the measurement track and taking photos at different measurement points to obtain a sequence of photos containing markers; S3, using an adversarial neural network to preprocess the photo sequence to obtain a high-quality image sequence; S4. For each photo in the photo sequence, use YOLOv5 to perform target detection and extract the encoded image area; S5, preprocessing each extracted coded image region, extracting a coded image from each preprocessed coded image region by an edge extraction operator, and parsing a unique label represented by each coded image; S6. Calibrate the internal parameters of the binocular camera using a calibration plate; S7, solving the relative position and posture of the binocular camera through the epipolar geometry constraint to obtain the external parameters of the binocular camera at each measuring point; S8, locating the center point of each coded image in each photo, and reconstructing the coordinates of each coded image based on the located center point according to the internal parameters of the binocular camera to obtain the three-dimensional coordinates of the marker corresponding to each coded image in the local binocular coordinate system; S9, converting the three-dimensional coordinates of each marker in the local binocular coordinate system into the three-dimensional coordinates of each marker in the global coordinate system using the external parameters of the binocular camera; S10, optimizing all three-dimensional coordinates of markers with the same unique number in the global coordinate system by bundle adjustment method, and obtaining optimized three-dimensional coordinates of all markers in the global coordinate system; S11. The optimized three-dimensional coordinates of all markers in the global coordinate system are assembled in the order of the track segments in which they are located, to obtain a three-dimensional linear curve of the measurement track.

2. The track line shape monitoring method based on binocular vision measurement according to claim 1 is characterized in that: The coded image is composed of three concentric rings, and the ratio of their radii is 5:12:

20. The three concentric rings are divided into 15 equal parts to form a 15-bit coding band. Each coding band is a coding area. If the coding area is black, it represents the binary number 0, and if it is white, it represents the binary number 1. Finally, a 15-bit binary number is formed as the unique label of each track segment.

3. The track line shape monitoring method based on binocular vision measurement according to claim 1 is characterized in that: Step S3 specifically includes: S301, filtering and denoising each photo in the photo sequence; S302, enhancing the brightness and contrast of the photos with low light or overexposure, and repairing the photos with strong light and reflection problems using image defogging or dereflection algorithms; S303, using a trained adversarial neural network to perform image preprocessing on each photo to obtain high-quality photos, wherein the adversarial neural network is trained using a number of low-quality images and corresponding high-quality image labels.

4. The track line shape monitoring method based on binocular vision measurement according to claim 1 is characterized in that: Step S4 specifically includes: S401, performing denoising, lighting adjustment and geometric correction on each photo in the photo sequence; S402, loading the trained YOLOv5 model, performing target detection on each photo in the photo sequence, and extracting the coded image area and its confidence; wherein the YOLOv5 model is trained using images collected by the track in different scenes, different angles, and different distances; S403, selecting and retaining the coded image regions whose confidence is greater than a preset threshold, and deleting the others; S404: if the position of the detected coded image area deviates from the expected position of the track area, it is determined to be a false detection; if the number of the detected coded image areas is less than a preset number, it is determined to be a missed detection; S405: If missed detection is found, adjust the binocular camera position or readjust the binocular camera exposure parameters.

5. The track line shape monitoring method based on binocular vision measurement according to claim 1 is characterized in that: Step S5 specifically includes: S501, preprocessing the extracted coded image area, including two-dimensional Gaussian filtering and binarization; S502, using the Canny operator to perform edge extraction on the preprocessed coded image region to obtain a coded image in the coded image region; S503: Decode the encoded image and parse out a unique label represented by the encoded image.

6. The track line shape monitoring method based on binocular vision measurement according to claim 1 is characterized in that: Step S6 specifically includes: calibrating the internal parameters of the binocular camera using a checkerboard calibration plate.

7. The track line shape monitoring method based on binocular vision measurement according to claim 1 is characterized in that: Step S8 specifically includes: S801, using a Gaussian curve fitting method to perform sub-pixel positioning of the center point of each encoded image, and obtain the coordinates of the center point of each encoded image on a two-dimensional image plane; S802: Reconstruct the three-dimensional coordinates of the coordinates of the center point of each encoded image on the two-dimensional image plane according to the internal parameters of the binocular camera, as the three-dimensional coordinates of each encoded image in the local binocular coordinate system.

8. A computer device comprising a memory, a processor and a computer program stored in the memory and executable on the processor, characterized in that: The processor executes the computer program to implement the method according to any one of claims 1 to 7.

9. A computer-readable storage medium having a computer program / instruction stored thereon, characterized in that: The computer program / instructions, when executed by a processor, implement the method of any one of claims 1-7.

10. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instructions are executed by a processor, the method according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Steel coil binocular-vision-positioning method and device

    CN108335331A

  • Method for measuring rainfall-induced landslide three-dimensional deformation by using binocular stereoscopic vision technology

    CN115112035A

Cited By

  • Visual displacement monitoring method for railway engineering

    CN120702351A