Train vibration test machine vision assisted positioning method based on smart phone

By generating high-definition videos on trains and using image recognition algorithms to detect the contact wire support numbers, the problem of marking the location of contact wire supports in railway vibration detection on trains has been solved, achieving precise positioning of vibration signals and improving the accuracy of anomaly diagnosis.

CN121728237APending Publication Date: 2026-03-24SHANGHAI INST OF TECH +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-31
Publication Date
2026-03-24

AI Technical Summary

Technical Problem

When conducting railway vibration testing on trains, it is difficult to accurately mark the specific locations of the overhead contact line supports, which affects the accuracy and traceability of vibration anomaly analysis.

Method used

By arranging lenses and stabilizer phone holders, high-definition videos are generated, and image recognition algorithms are used to detect the contact wire support numbers. Combined with time-amplitude signal curves, the precise location of vibration occurrence and number position is achieved.

Benefits of technology

It integrates vibration detection with geolocation, enabling real-time acquisition of railway vibration signals and automatic identification of catenary support numbers. This improves the accuracy and traceability of anomaly diagnosis, and enhances the intelligence level and fault location accuracy of railway vibration monitoring.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121728237A_ABST
    Figure CN121728237A_ABST
Patent Text Reader

Abstract

The invention discloses a train vibration testing machine vision auxiliary positioning method based on a smart phone, and the method comprises the steps: installing shooting equipment at the position of a train head or a train window through arranging a stable support with four glass suckers and a mobile phone fixing device of an anti-reflection cover, and achieving the high-definition video recording of a contact network pillar. The system uses YOLO and OCR algorithms to carry out automatic identification and time synchronization marking on the contact network pillar number in the video, and generates a picture file according to a time sequence. And then, the identified strut number is associated with a time axis of the train vibration waveform, and specific position marking and image information synchronous display are realized on the vibration waveform. According to the method, real-time detection and accurate positioning of abnormal vibration can be achieved in the train running process, and an efficient visual means is provided for track safety monitoring and maintenance.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of image recognition, in particular to a train vibration test machine vision assisted positioning method based on a smart phone. BACKGROUND

[0002] With the development of information technology, especially in recent years, the rapid development of artificial intelligence, artificial intelligence technology is also more and more widely used in mobile phones, such as smart chip technology, intelligent processing system technology, intelligent camera technology and intelligent image technology, etc., constantly improve and develop, so that the current mobile phone function is also more and more strong. The technology of mobile phone on the train for railway vibration detection has become a research point in the railway vibration detection. Detecting railway vibration can obtain vibration anomaly to judge whether the railway is damaged, so as to make more safe and efficient maintenance and operation. When the mobile phone on the train for railway vibration detection, the effect depends on the detection system on the mobile phone, how to obtain the specific location and vibration waveform anomaly point marking method of the actual railway.

[0003] With the research of people on the train for railway vibration detection by mobile phone in recent years, the mobile phone on the train for railway vibration detection has achieved certain results, however, the mobile phone on the train for railway vibration detection to touch net support number for specific location marking is a blank. If the railway vibration detection is carried out on the train, the specific location marking can greatly improve the analysis of the results, so as to more accurately judge or analyze the results in the actual railway line. SUMMARY

[0004] In view of the above problems, the present application is proposed.

[0005] To solve the above technical problems, the present application provides the following technical scheme: a train vibration test machine vision assisted positioning method based on a smart phone, comprising: By arranging the lens, the catenary support is captured, high-definition video is generated and stored; The high-definition video is disassembled into pictures, and an image recognition algorithm is used for target detection to obtain a number recognition result; The number recognition result is embedded into the time-amplitude signal curve collected and displayed by the sensor during the train running, the time relationship between vibration occurrence and number position is realized, so that the spatial position of the vibration signal can be accurately positioned, thereby providing a basis for safety monitoring and fault analysis of the railway line.

[0006] As a preferred scheme of the train vibration test machine vision assisted positioning method based on a smart phone, wherein: the lens arrangement comprises a stabilizer mobile phone support provided with four glass suction cups and a straight rod support; Four suction cups are fixed to the left, right, top, and bottom of the phone clip, and the reflector is placed at the back of the phone clip. Mount your phone with a clip on the train's front or side, such as near the window, to ensure a clear shot of the overhead contact line support.

[0007] As a preferred embodiment of the machine vision-assisted positioning method for train vibration testing based on smartphones described in this invention, the step of capturing the contact wire support includes setting the lens to automatically focus and real-time tracking and capturing objects outside the train, wherein the focused object is the railway contact wire support.

[0008] As a preferred embodiment of the machine vision-assisted positioning method for train vibration testing based on smartphones described in this invention, the following is provided: during autofocus, two focusing modules alternately focus; when the focusing modules are in use, only one focusing module is allowed to operate at a time, while the other focusing module is in standby mode. Once one focusing module completes identification, it enters standby mode, while another focusing module proceeds to focus on the next contact wire support.

[0009] As a preferred embodiment of the machine vision-assisted positioning method for train vibration testing based on smartphones described in this invention, the method of decomposing the high-definition video into images includes extracting keyframes from the frame sequence based on the clarity and content variation characteristics of the frames.

[0010] As a preferred embodiment of the machine vision-assisted positioning method for train vibration testing based on smartphones described in this invention, the image recognition algorithm includes performing image enhancement processing on the extracted frames to improve the image clarity and contrast. The location of the overhead contact line support pillars in the image was detected using a target detection model built with a neural network; and the number on the pillar was identified using OCR technology.

[0011] As a preferred embodiment of the machine vision-assisted positioning method for train vibration testing based on smartphones described in this invention, the embedding process of the number recognition result includes: extracting all images from the database in sequence; marking the images on the vibration waveform diagram in chronological order, and marking real-time time points on the time axis of the vibration waveform diagram; marking the contact wire support number at the corresponding time point position, and displaying the image information of each time point in association.

[0012] A machine vision-assisted positioning system for train vibration testing based on a smartphone, employing the method described in this invention, is characterized in that: the acquisition unit captures the contact wire support by arranging lenses, generates high-definition video, and stores it; The identification unit decomposes the high-definition video into images and uses an image recognition algorithm to perform target detection, thereby obtaining the identification result by number. The embedding unit embeds the identification result into the time-amplitude signal curve collected and displayed by sensors during train operation, realizing the time relationship between vibration occurrence and number position, so that the spatial position of vibration signal can be accurately located, thereby providing a basis for railway line safety monitoring and fault analysis.

[0013] A computer device includes: a memory and a processor; the memory stores a computer program, wherein: when the processor executes the computer program, it implements the steps of the method described in any one of the present invention.

[0014] A computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the steps of the method described in any one of the present invention.

[0015] The beneficial effects of this invention are as follows: This invention integrates vibration detection and geolocation, enabling real-time acquisition of railway vibration signals during train operation. AI algorithms automatically identify the contact wire support column numbers and accurately correlate them with the time points of the vibration waveforms. Compared to traditional vibration detection methods that only provide time information, this invention, through image recognition and time synchronization, allows each vibration anomaly to be located at a specific railway support column, greatly improving the accuracy and traceability of anomaly diagnosis. The system supports real-time dual-target tracking, autofocus, and dynamic tracking, ensuring image clarity and data continuity. Lightweight deployment is achieved through mobile deep learning frameworks such as TensorFlowLite or CoreML, improving operational efficiency and field adaptability. Furthermore, this method offers advantages such as low cost, simple structure, and convenient installation, making it suitable for on-site detection and intelligent operation and maintenance of conventional and high-speed railway lines, significantly improving the intelligence level and fault location accuracy of railway vibration monitoring. Attached Figure Description

[0016] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0017] Figure 1 The present invention provides an overall flowchart of a machine vision-assisted positioning method for train vibration testing based on a smartphone.

[0018] Figure 2The present invention provides a flowchart of a machine vision-assisted positioning method for train vibration testing based on a smartphone, during the recording stage.

[0019] Figure 3 This invention provides a machine vision-assisted positioning method for train vibration testing based on a smartphone, including a flowchart of the specific process of video-to-image processing.

[0020] Figure 4 This invention provides a machine vision-assisted positioning method for train vibration testing based on a smartphone, including a flowchart of the specific process of video-to-image processing.

[0021] Figure 5 The present invention provides an overall workflow and data association flowchart for a machine vision-assisted positioning method for train vibration testing based on a smartphone.

[0022] Figure 6 This invention provides a machine vision-assisted positioning method for train vibration testing based on a smartphone, and a workflow for a smartphone-based intelligent video recording system. Detailed Implementation

[0023] To make the above-mentioned objects, features, and advantages of the present invention more apparent and understandable, specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of them. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the protection scope of the present invention.

[0024] Reference Figures 1-6 As one embodiment of the present invention, a machine vision-assisted positioning method for train vibration testing based on a smartphone is provided, comprising: Step 1, camera setup.

[0025] This device includes a stabilizer phone holder with four glass suction cups and a straight rod bracket. The four suction cups are fixed to the left, right, top, and bottom of the phone clamp, acting like reflectors to protect against internal and external light, similar to umbrellas. The reflectors are located behind the phone clamp. The stabilizer phone holder with four glass suction cups and the straight rod bracket are as follows: The four suction cups are manually adjustable for suction and release, with appropriate suction force. The suction cups are connected together via a central support rod that can rotate left and up / down, forming a "plus" sign and securing the stabilizer device. The directional joint has a manual loosening device. The straight rod bracket is also connected to the "plus" sign from below. The stabilizer device includes a phone clamp and an anti-reflective shield. The anti-reflective shield is connected to the bracket via a central rod. The shield is made of elastic black velvet, is retractable, and has ribs. Ropes are attached to the connections between the four suction cups and the support rod.

[0026] Choose a phone with high resolution, high frame rate, low distortion, wide dynamic range, a wide-angle lens to cover a wider field of view, and a zoom lens to flexibly adjust the shooting distance. Mount the phone on the train's front or side, such as a window, to ensure the shooting angle clearly captures the overhead contact line support.

[0027] Step 2: Record video. For example... Figure 2 As shown.

[0028] The system automatically focuses and tracks objects outside the train in real time, based on the train's speed. The focused object is the railway catenary support pole, and the video is recorded in high definition using a similar recording method. The video is automatically saved to a designated folder.

[0029] The recording process includes: a dedicated recording program or recording system.

[0030] The video recording system utilizes deep learning image recognition to achieve the following: Recording angles can be from the train window, door, and locomotive. From the window and door, only a single overhead contact line support is visible; from the locomotive, multiple supports are visible. Therefore, during recording, the system prioritizes tracking the two closest supports (H1, H2) in the frame. It supports real-time dual-target tracking (M1, M2 focus frames), which can track synchronously or in relay. M1 and M2 are independently assigned to different supports, achieving precise tracking of each target. The sequence is: when the first overhead contact line support H1 enters the frame, the first autofocus and... The system uses real-time tracking or M1 for autofocus and real-time tracking. If the first contact wire support H1 hasn't left the frame yet, but the second contact wire support H2 has already entered the frame, the second system will use autofocus and real-time tracking or M2 will use autofocus and real-time tracking. When the first one leaves the frame, the first system will wait or track the third contact wire support into the frame. Similarly, when the second one leaves the frame, the second system will wait or track the fourth contact wire support into the frame, and so on until the end. During recording, the system automatically ends recording when the train enters a tunnel for too long (black screen), stops at a station, or the train stops (still screen), and saves the video as Video 1, Video 2, etc. Then, when the train leaves the tunnel (lit screen) or starts moving (moving screen), the system automatically starts recording and saves all videos to a designated folder.

[0031] The video recording, using a program or system with real-time tracking and autofocus modules, is achieved based on the following technologies: (1) Data collection: Collect image or video data, including the numbering of railway catenary supports, as well as relevant specifications. The height, shape, standard position of the catenary supports and their numbering are mainly collected from existing books, which is the height from the rail surface to the numbering position (for example, for conventional railways ≤160km / h, the numbering is usually located 2.0~2.5 meters above the rail surface). Video and image collection mainly involves taking photos and recording videos from train windows, doors, and locomotives and then collecting them (collecting data for conventional railways and high-speed railways respectively).

[0032] (2) Target Detection: YOLO models for image analysis and recognition, such as YOLOv5s, YOLOv5n, YOLOv8n, or MobileNetSSD, are used to achieve real-time detection of catenary support pole numbers. The process is as follows: 1) Data preparation stage: First, collect image or video data containing railway catenary supports and their numbers as the basis for subsequent training. Then, use annotation tools such as LabelImg to accurately label the bounding boxes and categories of the supports and numbers in each image. Next, improve the diversity of the dataset and the robustness of the model to different lighting and angles through data augmentation methods such as flipping, scaling, and brightness adjustment. Finally, divide the dataset into a training set (70%), a validation set (20%), and a test set (10%) to provide a scientific data basis for model training and evaluation. 2) Model selection and training stage: Lightweight models such as YOLOv5n, YOLOv8n, or MobileNetSSD are preferred. These models have efficient real-time inference capabilities, are suitable for mobile deployment, and can achieve frame rates of over 30 FPS. The training process includes: inputting labeled image data and performing preprocessing such as image resizing and pixel normalization; using the PyTorch or TensorFlow framework, setting hyperparameters such as learning rate (e.g., 0.001), batch size (e.g., 16), and number of training epochs (e.g., 100 epochs), and using the Adam optimizer to monitor the loss function decrease to optimize the model. After training, the model is converted to TensorFlowLite or CoreML format to ensure efficient operation on mobile devices such as smartphones, achieving real-time contact network support pillar number detection. 3) In the real-time detection stage, input video frames are first acquired from the camera or video stream, and each frame undergoes preprocessing such as resizing and normalization. The preprocessed frames are then input into the target detection model, which outputs the bounding box coordinates, confidence score, and category of the support pillar number. Redundant detection boxes are removed using non-maximum suppression (NMS), and OCR (e.g., Tesseract) is further applied to the detected support pillar regions to identify the numbers. Finally, the recognition results are overlaid and displayed on the video frames in real time. To achieve stable detection across consecutive frames, an autofocus and tracking module (such as DeepSORT) can be integrated to improve the consistency and accuracy of detection.

[0033] The corresponding model algorithms or formulas are based on commonly used real-time object detection algorithms such as YOLO and MobileNetSSD. YOLO (e.g., YOLOv5 / YOLOv8) divides the image into a grid, with each grid predicting multiple bounding boxes and class probabilities. Its core process includes: a backbone network (e.g., CSPDarknet, EfficientNet) extracts features; a neck network (PANet / FPN) fuses multi-scale features; and a head network outputs bounding boxes (x, y, w, h), confidence scores, and class probabilities. The prediction formula is: Center coordinate calculation: =σ( )+ =σ( )+ in It predicts the coordinates of the bounding box center (absolute coordinates relative to the entire image). , These are the two values ​​directly output by the model (which can be understood as the offset predicted by the network); σ is the sigmoid function, and its formula is: σ(x)= Its function is to map any real number to the interval (0,1); , The coordinates (integer) of the top-left corner of the current grid.

[0034] Width and height calculation: = = in It predicts the width and height of the bounding box; , These are the two values ​​directly output by the model (which can be understood as the logarithm of the scaling factor of the network prediction). , This is the preset anchor frame size (reference size). This is the scaling factor (the exponent ensures the result is positive). Where σ is the sigmoid function. , , , For model output, , For grid offset, , This refers to the anchor frame dimensions.

[0035] The loss function is: Loss= ∑CIOU+ ∑ + ∑ in , , These are the weighting coefficients; ∑ represents the summation over all predicted boxes; CIOU is the location loss. It is the object detection loss; It is a classification loss. MobileNetSSD uses MobileNet as its backbone network and employs multi-scale feature layers to predict bounding boxes, making it suitable for lightweight deployment.

[0036] Its loss function is: L(x,c,l,g)= ( (x,c)+α (x,l,g)) in This represents dividing by the number of positive samples, N; (x,c) is the confidence loss; (x,l,g) is the localization loss; α is the balance factor.

[0037] (3) Target Tracking: DeepSORT, such as MobileNetV2 or KCF (Kernelized Correlation Filter), is used as a feature extractor to track the catenary support column numbers. The process is as follows: First, target detection and feature extraction are performed. For each support column number region identified in the video frame, an image block I_crop is cropped. If the DeepSORT framework is used, I_crop is input into the MobileNetV2 network to extract a 128-dimensional or 256-dimensional feature vector f. This network uses depthwise separable convolution to reduce computation. If the KCF tracker is used, HOG features are calculated for the image block. First, the gradient values ​​Gx and Gy in the x and y directions of each pixel are calculated to obtain the gradient magnitude. The direction is calculated as θ = arctan(Gy / Gx), and finally, the gradient direction histogram within the cell is statistically analyzed and normalized. In the association measurement stage, DeepSORT simultaneously calculates motion and appearance similarity. Motion similarity is calculated using Mahalanobis distance. Where dᵢⱼ represents the position and velocity difference vector between the detection box i and the predicted trajectory j. Let S be the mean of historical differences, and S be the covariance matrix. Appearance similarity is calculated as follows: ;in and Let be two feature vectors, and <·,·> be the cosine similarity.

[0038] The final cost is the weighted sum. λ is the weighting parameter. For trajectory prediction and management, DeepSORT uses a Kalman filter to recursively estimate the target state. Prediction steps, =F· , =F· · +Q. (Among them) The prior state estimate (including position and velocity) at time k. Let F be the state transition matrix and Q be the process noise covariance, used for prior estimation of the covariance. Update steps: , , .in Here, H represents the observed value (the location of the detection box), H is the observation matrix, and R is the observation noise covariance. This is the Kalman gain. The trajectory is initialized, confirmed (for consecutive successful matches), put into hibernation (for brief loss), or deleted (for exceeding the loss limit) based on the matching results.

[0039] The KCF tracker employs frequency domain correlation filtering. First, training samples are generated from the base samples x using a circulant matrix, and the Gaussian kernel function k(x,z) = exp(-1 / ( )· ) Measure sample similarity. Solve the filter ŵ= in the frequency domain. , where F represents the Fourier transform and y is the desired output (Gaussian response plot). The kernel is the autocorrelation kernel, and λ is the regularization coefficient. During tracking, the response plot of the test sample z is calculated as f(z) = (F( The model is updated using linear interpolation, where F(w) is used to find the position with the maximum response (x,y)=argmaxf(z) as the new target position. =(1-η)· +η·ŵ, where η is the learning rate. In actual deployment, the algorithm is dynamically selected according to the scenario. For example, when multiple pillars appear simultaneously (such as from the perspective of a train engine), DeepSORT is used for multi-target tracking and identity maintenance; for single-pillar scenarios (such as from the perspective of a train window), a KCF tracker is instantiated independently for each target to improve computational efficiency.

[0040] (4) Autofocus: Contrast detection (variance or Laplacian operator) is used to focus on the contact wire support column numbers. The process is as follows: In the system, the autofocus module and the target tracking module work together. The system first locates the column number area through target detection, then uses contrast detection technology to achieve precise focus, and finally combines the tracking algorithm to maintain continuous lock on the target. The contact wire support column number area is located by processing the input image using detection models such as YOLO. Let the detection box coordinates be (x, y, w, h), and the image area containing the number is cropped from it as the region of interest (ROI) for subsequent focus calculation. Two contrast detection methods can be used for focus evaluation: 1) Variance-based focusing: Calculate the variance of pixel values ​​within the ROI region. = × .in σ² represents the grayscale value of each pixel, μ is the average grayscale value of all pixels in the region, and N is the total number of pixels. The system controls the lens movement via the camera API to find the lens position that maximizes σ².

[0041] 2) Laplacian operator focusing: Apply a Laplacian convolution kernel to the ROI region.

[0042] Commonly used core 1 Common Core 2 Calculate the focus measurement value: F= Or F= Where ∇²I(x,y) represents the Laplacian operation result at position (x,y). The system adjusts the camera position to find the maximum value of F. When multiple pillars appear simultaneously, the DeepSORT algorithm is used for tracking: MobileNetV2 is used to extract the feature vector f of the cropped region and calculate the motion similarity. .in Let i be the state difference vector between the detection box i and the trajectory j. Let S be the historical variance mean, and S be the covariance matrix.

[0043] The system processes the video stream in real time. First, it detects the area with the support column number, then evaluates the focus of that area, maximizing the focus metric by adjusting the lens position. During tracking, the system uses the current target location as the focus area and adjusts the focal length in real time to ensure image clarity, forming a closed-loop system of detection-focusing-tracking, achieving stable identification and tracking of railway catenary support column numbers.

[0044] (5) Video recording: The video recording stage is performed using Android CameraX API or iOS AV Foundation.

[0045] (6) Mobile deep learning frameworks: TensorFlowLite and CoreML; According to the following, the system will deploy the deep learning model trained on the server side to Android or iOS mobile devices after specific transformation and optimization, so that the mobile phone can complete the detection, recognition and tracking of pillar numbers locally without relying on network connection.

[0046] The core workflow of mobile deep learning frameworks begins with model conversion. Object detection models (YOLO series or MobileNetSSD) and feature extraction models (MobileNetV2), trained in server environments (such as using TensorFlow or PyTorch frameworks), are converted to mobile-specific formats. For Android devices, the TensorFlowLite conversion tool is used to convert the model to .tflite format; for iOS devices, it is converted to the CoreML-supported .mlmodel format. This conversion process not only changes the model format but also lays the foundation for subsequent optimization operations. Model optimization is performed immediately after conversion, which is a core step to ensure real-time performance on mobile devices. The main optimization techniques include quantization, pruning, and knowledge distillation. Quantization converts model weights from 32-bit floating-point numbers (float32) to 8-bit integers (int8) or 16-bit floating-point numbers (float16), reducing the model size by 75% and significantly improving inference speed; pruning further compresses the model by removing unimportant weights and connections from the network; knowledge distillation allows smaller student models to learn the behavior of the original large teacher model, reducing computational complexity while maintaining accuracy. After optimization, the size of detection models such as YOLOv8n can be controlled to below 5MB, and the inference speed can reach 30+ FPS on mobile device GPUs. During deployment, the optimized model files are integrated into the mobile application. Android applications load the .tflite model files in the assets directory via the TensorFlowLite interpreter, while iOS applications load the .mlmodel files via the CoreML framework. The model is preloaded during system initialization. During video stream processing, the model receives preprocessed image frames and outputs detection results, which are then accelerated by the device's GPU to achieve real-time identification and tracking of the overhead contact line support numbers.

[0047] Step 3: Video and image processing. For example... Figure 3 As shown.

[0048] The video-to-image generation system is a module that generates images from videos and names them according to a formula. An AI system (image recognition or deep learning technology) crops the most clearly visible contact wire support number from the video footage to generate an image, identifies the number, and marks it using real-time data. In the video-to-image processing step, real-time marking is as follows: Video 1 is read from a file, the most clearly visible contact wire number is cropped into an image, and the number within the image is identified using OCR technology. The marking method is naming.

[0049] The naming formula is: rename = dateh + h1 hours min + min1 minutes s + s1 seconds (contact wire support number), where date is the date, h is the time when video recording started (in hours), h1 is the time after video playback started (in hours), min is the time when video recording started (in minutes), min1 is the time after video playback started (in minutes), s is the time when video recording started (in seconds), and s1 is the time after video playback started (in seconds). For example, assuming video 1 has a duration of 25 minutes and starts recording at 9:23:08 AM on September 8, 2024, the video is imported into the processing system... The first image (p1) was cut out 7 seconds after the video started playing, and the contact wire support number it identified was 003. The second image (p2) was cut out 11 seconds after the video started playing, and the contact wire support number it identified was 004, ..., and the nth image (pn) was cut out 24 minutes and 10 seconds after the video started playing. So the first image was named 9:23:15 on September 8, 2024 (003), the second image was named 9:23:19 on September 8, 2024 (004), ..., and the nth image (pn) was named 9:47:18 on September 8, 2024 (contact wire support number).

[0050] The program or system for generating images from videos and naming them according to formulas is implemented based on the following technologies: (1) Video decoding: First, a video decoding library (such as FFmpeg) is needed to decode the video file into a frame sequence. The process is as follows: The video decoding process begins with the demuxing stage. The avformat_open_input(...) function is used to open the input video file and read its header information, identify the container format (such as MP4, FLV, etc.), and then separate the encapsulated video stream and audio stream to complete the preparation for reading the encoded data packet (AVPacket). In the core stage of video decoding, the system first uses the avcodec_find_decoder() function to find a matching decoder according to the video stream encoding format, and uses avcodec_open2() to initialize the decoder context. Then, av_read_frame() is used to read AVPacket data packets from the video stream in a loop, and the avcodec_send_packet() and avcodec_receive_frame() functions (or the traditional version avcodec_decode_video2()) are used to decode the encoded data into the original video frame (AVFrame). To improve system efficiency, a keyframe filtering strategy is adopted in the frame extraction and processing stage. I-frames are selectively extracted using FFmpeg filters, and only keyframes containing complete image information are processed. The decoded frame data is usually in YUV format, which needs to be initialized with sws_getContext() and converted to RGB format by calling the sws_scale() function for subsequent image processing and object detection.

[0051] The decoded frame data is passed to the image enhancement module for sharpness and contrast improvement. The enhanced frames are then passed to the target detection module for catenary support identification. All detection results are strictly correlated with the corresponding timestamp and stored in strict accordance with the naming formula dateh+h1hourmin+min1minutess+s1second (catenary support number) to ensure data consistency and traceability.

[0052] (2) Frame extraction: Keyframes are extracted from the frame sequence and can be filtered based on indicators such as frame sharpness and content changes. The process is as follows: 1) Sharpness assessment is the key step in the filtering. The edge sharpness can be effectively assessed by calculating the gradient information of the image. For example, the Tenengrad function uses the Sobel operator to calculate the gradient in the x and y directions, defined as S=Σ(G +G In the image, Gx and Gy represent the gradient magnitudes in the x and y directions, respectively, and S is the sum of gradient energy; a larger value indicates sharper image edges. The Laplace operator evaluates edge sharpness by calculating the second derivative of the image, highlighting areas of rapid intensity change through a second-order differential operator. The Brenner gradient function judges edge sharpness by calculating the sum of squared differences between adjacent pixel grayscale values, directly reflecting the intensity of local grayscale changes in the image. Analysis of variance calculates the variance of image pixel values. To evaluate contrast, I(x,y) represents the gray value of the pixel at coordinates (x,y), μ represents the average gray value of all pixels in the image, and N is the total number of pixels. A larger variance generally indicates higher image contrast and better clarity. Frequency domain analysis transforms the image to the frequency domain using Fourier transform, judging clarity by analyzing the distribution and quantity of high-frequency components. More high-frequency components indicate richer image details and more obvious edge information. 2) Content change detection is another important criterion for selecting keyframes, such as train stopping or entering a tunnel. Significant changes in the scene can be identified by calculating the differences between consecutive frames: pixel-level differences are calculated by directly calculating the sum of the absolute differences between pixel values ​​in two frames: Diff = Σ| (x,y)- (x,y)| is used to quantify the change, where (x,y) and (x, y) represent the pixel values ​​at position (x, y) in the two consecutive frames, respectively; the mean square error (MSE) is calculated using MSE = Σ The calculation of differences, where N is the total number of pixels, is more sensitive to larger differences and can highlight significant areas of change. The Structural Similarity (SSIM) metric comprehensively evaluates the similarity between two frames based on brightness, contrast, and structure, calculating a similarity score by comparing the features of local image patches. Histogram comparison detects content changes by statistically analyzing changes in color distribution: it calculates the color histograms of adjacent frames and uses metrics such as Bach's distance or chi-square distance to assess differences, where Bach's distance measures the overlap between two probability distributions; the histogram cross-validation method uses H(A,B)= The similarity is measured by the sum of the minimum values ​​of the two histograms in each bin, where A(i) and B(i) represent the values ​​of the two histograms in the i-th bin, respectively.

[0053] (3) Image Enhancement: Image enhancement processing is performed on the extracted frames to improve the image's clarity and contrast. This can be achieved using image processing libraries such as OpenCV. The image enhancement stage aims to systematically improve the quality of the input image and provide optimal visual data for the core recognition task. The process is as follows: 1) In terms of clarity improvement. Gradient information is obtained by calculating the gradient magnitudes of the image in the x and y directions and summing them up: S=Σ(G +G The edge strength is quantified by S, where Gx and Gy represent the rate of change of grayscale at a pixel in the horizontal and vertical directions, respectively. A larger S value indicates sharper edges and a clearer overall image. The Laplacian operator utilizes the second derivative property, calculating the second derivative of the image to highlight areas of rapid pixel intensity change. A larger response value indicates a higher probability that the point is at an edge or corner, resulting in a sharper image. Analysis of variance calculates the variance of grayscale values ​​in local regions. To evaluate texture richness, I(x,y) is the pixel value at coordinates (x,y), μ is the average pixel value in that region, and N is the number of pixels. A larger variance indicates stronger local contrast and richer details. 2) Contrast Enhancement. Linear transformation uses the formula g(x)=α×f(x)+β for global adjustment, where f(x) is the input pixel value, g(x) is the output pixel value, α is the contrast control coefficient (contrast increases when α>1, decreases when α<1), and β is the brightness offset (brightens when β>0, darkens when β<0). Histogram equalization expands the dynamic range by redistributing pixel grayscale values. Standard histogram equalization (HE) globally adjusts contrast, adaptive histogram equalization (AHE) individually processes the local neighborhood of each pixel to enhance local contrast, and contrast-limited adaptive histogram equalization (CLAHE) avoids excessive noise enhancement by limiting local contrast amplification, making it particularly suitable for scenarios requiring enhanced fine textures, such as contact wire support numbering. There is also sharpening specifically for enhancing edge and detail information. In the spatial domain, the Laplacian operator uses a specific convolution kernel, such as [0,-1,0;-1,5,-1;0,-1,0], to convolve with the image, highlighting edges by enhancing the difference between the center pixel and its neighboring pixels. Unsharpened masking, on the other hand, first blurs the original image to obtain low-frequency information, then subtracts the blurred image from the original to obtain high-frequency details, and finally weights and superimposes the high-frequency details back into the original image to enhance edges. In the frequency domain, high-pass filtering directly attenuates the low-frequency components after the Fourier transform while preserving the high-frequency components, thus enhancing edges; homomorphic filtering decomposes the image into illumination and reflectance components through logarithmic transformation, and filters these two components separately in the frequency domain, simultaneously compressing dynamic range and enhancing contrast. Image restoration focuses on addressing degradation issues such as noise and blurring. In denoising, mean filtering effectively smooths Gaussian noise using neighborhood averaging; median filtering eliminates salt-and-pepper noise by taking the median of neighboring pixels; Gaussian filtering uses a Gaussian kernel for weighted averaging, preserving edges well while denoising; wavelet transform denoising achieves adaptive denoising by thresholding different frequency coefficients after multi-scale decomposition. In deblurring, Wiener filtering searches for the optimal restoration filter in the frequency domain based on the statistical characteristics of the image and noise; the blind restoration principle, when the point spread function is unknown, iteratively estimates and simultaneously restores the clear image and the blur kernel. 3) Brightness optimization. The overall brightness level is evaluated by calculating the average gray value of the entire image, and the proportion of overexposed / underexposed areas is statistically analyzed to detect areas requiring local correction.

[0054] (4) Object Detection: Use object detection models (such as YOLO, SSD, Faster R-CNN) to detect the location of catenary poles in the image. The process is as follows: 1) In the model training stage, build a dedicated detector that can identify the target "catenary pole". First, a large-scale dataset needs to be built to collect railway scene images covering different weather, lighting and angle conditions. Then, use annotation tools (such as LabelImg) to draw accurate bounding boxes around each catenary pole in all images and assign the label "catenary_pole". This process provides the model with the "standard answer" needed for learning. The model usually chooses an architecture that balances speed and accuracy, such as YOLOv8n or MobileNetSSD, and uses pre-trained weights on a large dataset (such as COCO) for transfer learning to accelerate convergence. During training, the model extracts multi-level features of the image through its backbone network (such as CSPDarknet) and fuses these features through the neck network (such as FPN+PAN) to capture information from global contours to local details. The essence of training is to adjust the model parameters by iteratively optimizing the loss function, so that the predicted bounding boxes continuously approximate the true annotations. The loss function of the YOLO model is usually expressed as: Loss= ∑ (CIoU)+ ∑ + ∑ Where ∑ CIoU (CompleteIntersectionoverUnion) is the positional loss, responsible for calculating the overlap between the predicted bounding box and the ground truth bounding box. CIoU takes into account the overlap area, center point distance, and aspect ratio. It is the confidence loss, which measures the accuracy of the model in determining whether an object exists within the bounding box; ∑ This is a classification loss, and since only the single category "pillar" is detected here, the weight of this item can be reduced. , , To balance the weighting coefficients of different loss terms, after model training, its performance needs to be evaluated on a reserved test set using mAP (mean accuracy) as the main metric. Once the target is met, it is exported in deployment format. 2) The inference application stage involves putting the trained model into practical use. The system first loads the model into memory, and then receives a single-frame image from the image enhancement module. The image is scaled and normalized before being fed into the model. The model performs forward propagation and outputs a series of raw detection boxes. Each box contains initial center coordinates, width and height, confidence score, and class probability. The core algorithm used by its detection model is the same as that of the formula and numbered detection models (such as the formula for calculating the center coordinates and width and height of the bounding boxes in the YOLO series algorithms), which will not be repeated here. The model's raw output undergoes Non-Maximum Suppression (NMS) post-processing: First, unreliable predictions are filtered out based on a confidence threshold (e.g., 0.5). Then, the remaining predicted boxes are sorted by confidence level, and other boxes with an Intersection over Union (IoU) exceeding a set threshold (e.g., 0.45) are removed one by one to eliminate duplicate detections and ensure that each pillar corresponds to only one most accurate bounding box. The final output of this stage is structured data, generating an entry for each detected pillar containing bounding box coordinates (x, y, width, height), a confidence score, and a corresponding timestamp. This provides a precise operating area for the subsequent numbering and recognition module.

[0055] (5) Number recognition: Use OCR technology (such as Tesseract-OCR) to recognize the numbers on the pillars. The process mainly relies on optical character recognition (OCR) technology, and its implementation can be divided into: 1) Text region detection: The aim is to locate the specific position of the number text in the pillar image. Since the numbers are usually sprayed on the surface of the pillar, there may be complex situations such as verticality, tilt, perspective distortion, and uneven lighting. Therefore, a robust detection algorithm is required. A deep learning-based text detection model (such as the detection part of EAST or CRNN) can be used. Its backbone network (such as PVANet) first extracts the image feature map, and then generates a rectangular or quadrilateral region containing the text line through the algorithm. For numbers with regular shapes, traditional methods such as the Maximum Stable Extreme Region (MSER) ​​algorithm can also be used. It binarizes the image through a series of thresholds and extracts the connected regions that remain stable during the threshold change process. Its mathematical expression is: For an image I and a series of thresholds δ, MSER can be defined as a connected region that satisfies the condition. : in This represents the binary region below threshold i. The threshold step size, The stability threshold parameter controls the sensitivity to changes in region area. The output of this process is the bounding box coordinates of the numbered text lines. 2) Character recognition is responsible for converting the detected text image regions into machine-encoded character sequences. Currently, the mainstream approach is to use an end-to-end recognition model based on convolutional recurrent neural networks (CRNNs). This model first extracts deep features of the image region through a CNN (such as VGGNet) and converts them into feature sequences. Let the input image region be X. After multiple convolutions and pooling, the output feature map is reconstructed into a feature sequence. Where T is the sequence length, and each feature vector This corresponds to the information of a receptive field in the horizontal direction of the image. Subsequently, this sequence is fed into a bidirectional LSTM (Bi-LSTM) layer, and its forward and backward computation processes can be represented as follows: =LSTM( , ) =LSTM( , ) in and These represent the forward and backward hidden states at time t, respectively. Bi-LSTM, by fusing contextual information, outputs a more robust feature sequence H=( , ,..., ),in =[ ; Finally, the Connectionist Temporal Classification (CTC) layer is responsible for solving the problem of predefined segmentation in sequence labeling. CTC calculates conditional probabilities through dynamic programming. ,in It is a length of The path (allowing whitespace "-" and repeating characters) is summed to find all paths that can be mapped to actual tags. Path probability: = Its loss function is the negative log-likelihood: =- During decoding, beam search or greedy decoding (taking the most likely output at each time step, then merging duplicate characters and removing whitespace) is typically used to obtain the final most likely character sequence. = .

[0056] (6) Image cropping: Based on the detected pillar positions, crop out the area containing the number.

[0057] (7) Time stamp: Get the recording time and playback time of the video and generate a file name according to the specified format.

[0058] (8) File naming: Name the cropped image files according to the naming formula.

[0059] Step 4: Mark the railway catenary support post number on the vibration waveform.

[0060] The system extracts all image names and numbers from the database sequentially, importing or labeling them into the railway vibration system based on the labeling time or image name. These images are then labeled on the vibration waveform diagram in chronological order, with real-time time points marked on the time axis. The contact wire support number is labeled at the corresponding time point, and the image information for that time point is displayed accordingly. The system, using image names for labeling, is implemented as a program or system using the following techniques: First, a lightweight model (such as YOLOv5s, YOLOv8n, or MobileNetSSD) is selected to train and optimize the target detection model (e.g., processing image names), and it is deployed to mobile devices to ensure rapid inference and efficient detection. The text recognition function of the integrated OCR module is embedded into the system to achieve accurate detection and recognition of text in the target area (e.g., recognizing image names). Simultaneously, an efficient data storage scheme is designed to save detection results, text information, and related data, ensuring efficient and real-time data management. Furthermore, data visualization development is required, using MPAndroidChart (for Android) or Charts (for iOS) to visually display the processed data in chart form and support user interaction. Furthermore, it is essential to fully consider the integration with cloud services, combining mobile development with deep learning model optimization, and deploying trained models (such as TensorFlow Lite or CoreML formats) to the system to ensure efficient operation and a smooth user experience. (Reference) Figure 4 This diagram illustrates the spatiotemporal correlation between vibration signals and visual positioning, combining vibration data and visual images to achieve more accurate railway condition analysis. The vibration signals in the diagram are measured in millimeters, showing the variation of vibration amplitude over time at different locations (e.g., K55+000, K55+060), fluctuating between -4mm and +3mm. Each time point (T0, T1, T2, etc.) is associated with its corresponding railway location. Through spatiotemporal correlation methods, vibration data is matched with visual images of the corresponding locations (e.g., G1037, G1038, etc.). The purpose of this diagram is to use this spatiotemporal synchronous analysis to help further understand the condition of railway equipment and tracks, promptly identify potential faults or safety hazards, and thus improve the safety and efficiency of railway transportation.

[0061] Figure 5The overall workflow and data association process are demonstrated, divided into two main parts: vibration signal and image data acquisition, and intelligent vision processing. First, image data is acquired via a smartphone, and vibration signals are collected via vibration sensors to obtain vibration curves and corresponding image data. Next, a data fusion and localization layer performs spatiotemporal correlation between the vibration signals and image data, ensuring data synchronization and accuracy. In the intelligent vision processing layer, video data is input and preprocessed through image enhancement and keyframe extraction to prepare for subsequent target recognition. A YOLO model is used for target detection, identifying targets in the images, such as railway catenary supports, and OCR technology is combined to identify identification numbers within the images. Finally, by combining timestamp information, precise spatiotemporal data correlations are generated, enabling real-time monitoring and analysis of the targets.

[0062] Figure 6 This demonstration showcases the workflow of a smartphone-based intelligent video recording system. Upon startup, the system first captures high-definition, high-frame-rate video by calling the phone's camera API, selecting appropriate shooting angles based on different scene settings, such as the angle of a car window, door, or train locomotive. During scene setting, the system can choose between single-target tracking or dual-target relay tracking modes, using the DeepSORT algorithm for dynamic target tracking. Next, the system uses an autofocus module to adjust the focus in real time to ensure the target remains clearly visible, while simultaneously performing target comparison and detection to ensure accurate target identification within the image. During recording, the system automatically encodes and stores the video, adjusting lighting and motion control based on environmental awareness technology to ensure image stabilization. Finally, when an environmental change is detected (such as entering a tunnel or station), the system automatically stops recording and saves the current video. Through this process, the intelligent video recording system can efficiently and accurately complete recording tasks in dynamic environments.

[0063] On the other hand, this embodiment also provides a smartphone-based machine vision-assisted positioning system for train vibration testing, which includes: The acquisition unit uses lenses to capture images of the overhead contact line supports, generating and storing high-definition video.

[0064] The identification unit decomposes the high-definition video into images and uses an image recognition algorithm to perform target detection, thereby obtaining the identification result by number.

[0065] The embedding unit embeds the identification result into the time-amplitude signal curve collected and displayed by sensors during train operation, realizing the time relationship between vibration occurrence and number position, so that the spatial position of vibration signal can be accurately located, thereby providing a basis for railway line safety monitoring and fault analysis.

[0066] Hardware and angle requirements allow for the installation of mobile phones with high resolution, high frame rate, low distortion, wide dynamic range, wide-angle lenses to cover a wider field of view, and zoom lenses to flexibly adjust the shooting distance. These phones can be mounted on the front or side of the train, such as windows, to ensure that the shooting angle can clearly capture the contact wire support.

[0067] If the above functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0068] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (such as a computer-based system, a processor-including system, or other system that can fetch and execute instructions from, an instruction execution system, apparatus, or device). For the purposes of this specification, "computer-readable medium" can be any means that can contain, store, communicate, propagate, or transmit programs for use by, or in conjunction with, an instruction execution system, apparatus, or device.

[0069] More specific examples of computer-readable media (a non-exhaustive list) include: electrical connections (electronic devices) having one or more wires, portable computer disk drives (magnetic devices), random access memory (RAM), read-only memory (ROM), erasable and editable read-only memory (EPROM or flash memory), fiber optic devices, and portable optical disc read-only memory (CDROM). Furthermore, computer-readable media can even be paper or other suitable media on which the program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, interpreting, or otherwise processing as necessary, and then stored in computer memory.

[0070] It should be understood that various parts of the present invention can be implemented in hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented in software or firmware stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, it can be implemented using any or a combination of the following techniques known in the art: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.

[0071] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.

Claims

1. A machine vision-assisted positioning method for train vibration testing based on a smartphone, characterized in that, include: By deploying cameras, the overhead contact line supports are captured, generating high-definition video and storing it; The high-definition video is decomposed into images, and target detection is performed using image recognition algorithms to obtain the identification results by number; The identification result of the number is embedded into the time-amplitude signal curve collected and displayed by the sensors during the operation of the train, so as to realize the time relationship between the vibration occurrence and the number position, and make the spatial position of the vibration signal can be accurately located, thereby providing a basis for the safety monitoring and fault analysis of railway lines.

2. The machine vision-assisted positioning method for train vibration testing based on a smartphone as described in claim 1, characterized in that: The lens arrangement includes a stabilizer phone holder with four glass suction cups and a straight rod holder. Four suction cups are fixed to the left, right, top, and bottom of the phone clip, and the reflector is placed at the back of the phone clip. Mount your phone with a clip on the train's front or side, such as near the window, to ensure a clear shot of the overhead contact line support.

3. The machine vision-assisted positioning method for train vibration testing based on a smartphone as described in claim 2, characterized in that: The capture of the overhead contact line support includes setting the camera to automatically focus as the camera follows the train's speed and to track and capture objects outside the train in real time; wherein the object to be focused on is the railway overhead contact line support.

4. The machine vision-assisted positioning method for train vibration testing based on a smartphone as described in claim 3, characterized in that: During autofocus, two focusing modules alternately focus; when in use, only one focusing module is allowed to operate at a time, while the other focusing module is in standby mode. Once one focusing module completes identification, it enters standby mode, while another focusing module proceeds to focus on the next contact wire support.

5. The machine vision-assisted positioning method for train vibration testing based on a smartphone as described in claim 4, characterized in that: Decomposing the high-definition video into images involves extracting keyframes from the frame sequence based on the frame's clarity and content variation characteristics.

6. The machine vision-assisted positioning method for train vibration testing based on a smartphone as described in claim 5, characterized in that: The image recognition algorithm includes performing image enhancement processing on the extracted frames to improve the image clarity and contrast; The location of the overhead contact line support pillars in the image was detected using a target detection model built with a neural network; and the number on the pillar was identified using OCR technology.

7. The machine vision-assisted positioning method for train vibration testing based on a smartphone as described in claim 6, characterized in that: The embedding process of the identification results includes: extracting all images from the database in sequence; marking the images on the vibration waveform diagram in chronological order and marking real-time time points on the time axis of the vibration waveform diagram; marking the contact wire support number at the corresponding time point and displaying the image information of each time point in association.

8. A smartphone-based machine vision-assisted positioning system for train vibration testing, employing the method described in any one of claims 1-7, characterized in that: The acquisition unit captures images of the overhead contact line supports by arranging lenses, generates high-definition video, and stores it. The identification unit decomposes the high-definition video into images and uses an image recognition algorithm to perform target detection, thereby obtaining the identification result by number. The embedding unit embeds the identification result into the time-amplitude signal curve collected and displayed by sensors during train operation, realizing the time relationship between vibration occurrence and number position, so that the spatial position of vibration signal can be accurately located, thereby providing a basis for railway line safety monitoring and fault analysis.

9. A computer device, comprising: A memory and a processor; the memory stores a computer program, characterized in that: when the processor executes the computer program, it implements the steps of the method as described in any one of claims 1-7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, it implements the steps of the method as described in any one of claims 1-7.