An image-based method and system for monitoring and identifying the operating status of a train at a station

Through technical means such as two-dimensional wavelet decomposition, Gaussian pyramid and SGM algorithm, the problem of image data being disturbed by light and noise in the existing technology is solved, and efficient and stable feature extraction and train operating status recognition are achieved, improving the recognition efficiency and accuracy.

CN119693884BActive Publication Date: 2025-06-03LIAONING QIHUI ELECTRONIC SYST ENG CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510221176.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-27
Publication Date
2025-06-03
Estimated Expiration
2045-02-27

AI Technical Summary

Technical Problem

The existing image-based train operating status monitoring method is disturbed by light changes, background complexity and noise, making it difficult to achieve efficient and stable feature extraction in complex scenarios, resulting in a reduced recognition efficiency.

Method used

Two-dimensional wavelet decomposition is used for denoising, Gaussian pyramid is constructed for multi-scale feature extraction, and the disparity map is generated using the SGM algorithm, combined with the random sampling consistency algorithm to remove error matching points, form point cloud data, and identify key point cloud data through deep learning models, calculate comprehensive feature values ​​for abnormal judgment.

Benefits of technology

It significantly improves the local details and multi-scale features of the image, improves the noise resistance and adaptability of point cloud data, enhances the efficiency of identifying train operating states, and ensures high accuracy and robustness in complex scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119693884B_ABST
    Figure CN119693884B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and system for monitoring and identifying the running state of a train at a station based on images, which relates to the technical field of train state identification. The method includes collecting image data for two-dimensional wavelet decomposition, applying a soft threshold for denoising processing and regenerating the denoised aligned image data; constructing a Gaussian pyramid of the aligned image data, performing equalization mapping through histogram equalization, and merging each layer of the Gaussian pyramid into complete image data. The method of the present invention significantly enhances the local details of the image under different lighting conditions through CLAHE adaptive contrast adjustment, making the originally low-contrast areas clearer. Through multi-scale feature extraction, the Gaussian pyramid decomposition generates image levels with different resolutions, ensuring that the structural information of the image at different scales is retained. Through the combination of SGM and RANSAC, the point cloud still has strong noise resistance in complex scenes and can adapt to environmental conditions such as light changes and reflection interference.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of train status recognition, and particularly to a method and system for monitoring and recognizing the running status of a train in a station based on images. Background Art

[0002] With the rapid development of modern railway transportation, the safety and efficiency of train operation have become the core issues concerned in the field of transportation. In the scenario of train operation in a station, the dynamic status monitoring of the train and the accurate recognition of its key components are directly related to operation safety and maintenance efficiency. Traditional train operation status monitoring methods usually rely on physical sensors or manual inspections, such as physical devices like acceleration sensors and strain gauges. Although these methods have a certain degree of reliability;

[0003] However, image data is usually significantly affected by illumination changes, background complexity, and noise interference, which directly affects the accuracy of subsequent feature extraction and anomaly judgment. On the other hand, existing methods mostly adopt single-scale feature extraction or global equalization processing, ignoring the effective information of multi-scale details and local contrast in the image. This single processing method is difficult to achieve efficient and stable feature extraction in complex scenarios, thus reducing the recognition efficiency of the train operation status. Summary of the Invention

[0004] In view of the above problems existing in the existing method and system for monitoring and recognizing the running status of a train in a station based on images, the present invention is proposed.

[0005] Therefore, the problems to be solved by the present invention are that image data is usually significantly affected by illumination changes, background complexity, and noise interference, which directly affects the accuracy of subsequent feature extraction and anomaly judgment. On the other hand, existing methods mostly adopt single-scale feature extraction or global equalization processing, ignoring the effective information of multi-scale details and local contrast in the image. This single processing method is difficult to achieve efficient and stable feature extraction in complex scenarios, thus reducing the recognition efficiency of the train operation status.

[0006] To solve the above technical problems, the present invention provides the following technical solution: A method for monitoring and recognizing the running status of a train in a station based on images, which includes,

[0007] Collecting image data for two-dimensional wavelet decomposition, applying soft thresholding for denoising processing and regenerating the aligned image data after denoising;

[0008] Constructing a Gaussian pyramid of the aligned image data, performing equalization mapping through histogram equalization, and merging the Gaussian pyramid of each layer into complete image data;

[0009] Based on the complete image data, use the SGM algorithm to generate a disparity map, use the random sample consensus algorithm to remove the matching points with large errors, and calculate the spatial coordinates of each pixel to form point cloud data, and construct a deep learning model to identify the key point cloud data;

[0010] Determine the gradient feature and texture feature to calculate the comprehensive eigenvalue based on the image data decomposed by the Gaussian pyramid, perform an anomaly judgment on the comprehensive eigenvalue corresponding to the key point cloud data, and issue a warning and generate an anomaly point report based on the anomaly judgment.

[0011] As a preferred solution of the method for monitoring and identifying the running state of a train at a station based on images according to the present invention, wherein: the collected image data is subjected to two-dimensional wavelet decomposition, and soft thresholding is applied for denoising processing and the denoised aligned image data is regenerated, including,

[0012] Collect image data of multiple perspectives of the train through a graphic acquisition camera;

[0013] Select the Daubechies wavelet as the wavelet basis function, perform two-dimensional wavelet decomposition on the image data, and the obtained subbands include low-frequency components and high-frequency components, representing different frequency details of the image data;

[0014] Apply soft thresholding to the high-frequency component subband for denoising processing to eliminate image noise, which is expressed as:

[0015]

[0016] where λ represents the soft threshold, σ represents the standard deviation of the high-frequency subband, and N represents the total number of pixels in the subband;

[0017] Apply the soft threshold to each high-frequency coefficient of the high-frequency component subband, which is expressed as:

[0018] T(x) = sign(x)·max(|x| - λ, 0);

[0019] where T(x) represents the high-frequency coefficient value after denoising by soft thresholding, λ represents the soft threshold, and sign(x) represents the sign function, representing the positive or negative sign of x;

[0020] Based on the denoised high-frequency coefficient values, perform inverse wavelet transform on the denoised high-frequency component subband, and merge the decomposed band data into the denoised aligned image data.

[0021] As a preferred solution of the method for monitoring and identifying the running state of a train at a station based on images according to the present invention, wherein: constructing a Gaussian pyramid of the aligned image data, performing equalization mapping through histogram equalization, and merging each layer of the Gaussian pyramid into the complete image data, including,

[0022] Construct a Gaussian pyramid of the aligned image data, decompose the aligned image data into different levels of multiple scales, where each level represents the scale of the image at the k-th level, expressed as:

[0023]

[0024] where I k+1 (x, y) represents the image of the (k + 1)-th layer of the Gaussian pyramid, and ω(m, n) represents the Gaussian weight coefficient;

[0025] Use the Contrast Limited Adaptive Histogram Equalization (CLAHE) method for the images of each layer of the Gaussian pyramid, divide each layer of the image into blocks, count the gray-level distribution of each sub-block, and generate a gray histogram H g , where g represents the gray level;

[0026] Based on the gray histogram H g Calculate the cumulative distribution function CDF, expressed as:

[0027]

[0028] where CDF g represents the cumulative pixel frequency from the minimum gray level 0 to the current gray level g, and H g' represents the pixel frequency of the gray level g', where g' is the gray-level variable used for summation starting from 0 and gradually accumulating to the current gray level g;

[0029] Determine the contrast limit threshold T based on the product of the clipping threshold ratio of the historical histogram and the total number of pixels in the sub-block;

[0030] Based on the contrast limit threshold T, if the pixel frequency H of the gray level g' in the gray histogram g' is greater than the contrast limit threshold T, then reduce H g' to T, and calculate the pixel frequency H of the gray-level pixels exceeding the contrast limit threshold T g' and evenly distribute it to all gray levels to obtain the adjusted gray-level pixel frequency, and recalculate the cumulative distribution function CDF;

[0031] Calculate the equalization mapping according to the recalculated cumulative distribution function, expressed as:

[0032]

[0033] where P n (i, j) represents the equalized pixel value of the image block pixel (i, j), and CDF gr represents the cumulative distribution function of the adjusted gray-level pixel frequency, and CDF minrepresents the minimum value of the non - zero cumulative distribution, L' represents the total number of gray levels, and v represents the total number of pixels in the sub - block;

[0034] Map the equalized values back to the corresponding pixels in the sub - block;

[0035] For the overlapping regions of adjacent sub - blocks, use bilinear interpolation for smooth transition, and combine all sub - blocks to generate the image data processed by CLAHE;

[0036] Based on the images of each layer of the Gaussian pyramid, starting from the image with the lowest resolution, upsample it layer by layer to a higher resolution through bilinear interpolation, and finally merge it into the complete image data.

[0037] As a preferred scheme of the method for monitoring and identifying the running state of a train at a station based on images according to the present invention, wherein: based on the complete image data, use the SGM algorithm to generate a disparity map, use the random sample consensus algorithm to remove the matching points with large errors, and calculate the spatial coordinates of each pixel to form point cloud data, including,

[0038] According to the complete image data processed by CLAHE collected from the left and right viewpoints, use the SGM algorithm to generate a disparity map, and calculate the matching cost of each pixel through the sum of absolute differences (SAD), expressed as:

[0039]

[0040] where D(x, y, d) represents the matching cost of the pixel located at (x, y) when the disparity is d, k represents the window size for controlling the matching range of the local area, I L (x + i, y + j) represents the pixel value of the left view, I R (x + i - d, y + j) represents the pixel value of the right view shifted by the disparity d;

[0041] Optimize the matching cost through the cumulative cost;

[0042] Assign a disparity value to each pixel point based on the disparity d of the minimum optimized matching cost to generate a preliminary disparity map,

[0043] Use the random sample consensus algorithm RANSAC, based on the pixel point positions of the left and right view images collected from the preliminary disparity map, detect and remove the matching points with large errors to generate an optimized disparity map;

[0044] Based on the disparity values at different positions of the optimized disparity map, calculate the spatial coordinates of each pixel, expressed as:

[0045]

[0046] Where Z represents the depth of the pixel point, f represents the camera focal length, B represents the baseline distance, and d(x, y) represents the disparity value of the pixel point (x, y); X and Y respectively represent the abscissa and ordinate of the spatial coordinates, and c x and c y represent the abscissa and ordinate of the camera principal point coordinates;

[0047] Traverse all pixels in the disparity map, calculate their corresponding spatial coordinates, and form point cloud data.

[0048] As a preferred solution of the method for monitoring and identifying the running state of a train at a station based on images according to the present invention, wherein: the construction of the deep learning model to identify key point cloud data includes,

[0049] Construct a PointNet++ deep learning model, including an input layer, a local feature extraction layer, a feature propagation layer, a hierarchical feature aggregation layer, and an output layer;

[0050] Among them, the input layer inputs point cloud data, the local feature extraction layer uses a distance-based grouping mechanism to select points within a local area, the hierarchical feature aggregation layer aggregates the extracted local features layer by layer, the feature propagation layer uses a feature propagation mechanism to transfer high-level feature information back to the low-level point cloud data, and the output layer outputs a prediction result based on the segmentation label;

[0051] Use the point cloud samples calibrated with labels as training data, select the cross-entropy loss function to calculate the difference between the class probability predicted by the model and the actual label, use the Adam optimizer for optimization, and stop iterating and output the model when the loss calculated during continuous iteration no longer decreases significantly;

[0052] Input the newly calculated point cloud data into the model, and output the label probabilities of being recognized as different key components of the train and the background. Select the maximum probability to determine the label of the point cloud data, and select the point cloud data of the key components of the train as the key point cloud data.

[0053] As a preferred solution of the method for monitoring and identifying the running state of a train at a station based on images according to the present invention, wherein: the determination of the gradient feature and texture feature of the image data based on Gaussian pyramid decomposition to calculate the comprehensive feature value, and the abnormal judgment of the comprehensive feature value corresponding to the key point cloud data includes,

[0054] Determine the gradient feature and texture feature of the image data based on Gaussian pyramid decomposition, which is expressed as:

[0055]

[0056] Where G k (x, y) represents the gradient feature of the pixel point (x, y) on the k-th layer of the Gaussian pyramid, and I k(x, y) represents the image of the pixel point (x, y) in the k-th layer of the Gaussian pyramid, represents the horizontal gradient, represents the vertical gradient, T k (x, y) represents the texture feature of the pixel point (x, y) in the k-th layer of the Gaussian pyramid, p k (i, j) represents the frequencies of the gray values i and j appearing in the neighborhood window centered on the pixel point (x, y) in the k-th layer of the Gaussian pyramid, (i - j) 2 represents the square of the gray value difference, which is used to measure the contrast degree of the gray values;

[0057] Non-linearly combine the gradient features and texture features of each layer, which is expressed as:

[0058] F k (x, y) = ln(1 + G k (x, y) · T k (x, y));

[0059] where F k (x, y) represents the comprehensive feature value of the k-th layer image. The comprehensive feature values of all layers are weighted and fused to generate the full-resolution comprehensive feature value;

[0060] Based on the coordinates of the key point cloud data, recalculate its pixel coordinates on the plane, and extract the feature values of the pixel coordinates on the plane according to the full-resolution comprehensive feature value;

[0061] Statistically analyze the historical distribution of the feature values of the key point cloud data, and set the abnormal threshold range according to the difference between the feature value mean and twice the standard deviation and the sum of the feature value mean and twice the standard deviation;

[0062] If the feature value of the extracted pixel coordinates on the plane is greater than or equal to the maximum value of the abnormal threshold range, record it as the abnormal point coordinates. If the feature value of the extracted pixel coordinates on the plane is less than the minimum value of the abnormal threshold range, record it as the abnormal point coordinates.

[0063] As a preferred solution of the method for monitoring and identifying the running state of a train at a station based on an image according to the present invention, wherein: the early warning is performed based on the abnormal judgment and an abnormal point report is generated, including,

[0064] Form an abnormal point list with the point cloud coordinates and corresponding feature values of the abnormal point coordinates, and color-mark the plane image according to the abnormal point coordinates;

[0065] And give an early warning to the maintenance personnel according to the sound and light alarm, and specifically judge and record the abnormal type according to the maintenance personnel;

[0066] Use a table generation tool to generate an abnormal point report according to the abnormal point list, the plane image marked with colors, the maintenance records of the maintenance personnel, and the abnormal type records.

[0067] Another object of the present invention is to provide a system for monitoring and identifying the running state of a train at a station based on images, which includes:

[0068] An image preprocessing module, which collects multi-view image data and performs preprocessing, performs two-dimensional wavelet decomposition, separates low-frequency and high-frequency components, applies soft thresholding for denoising, and generates denoised and aligned image data;

[0069] A multi-scale feature enhancement module, which constructs a Gaussian pyramid of the aligned image data, generates multi-level image representations, performs histogram equalization on each layer of the image to enhance local contrast, and combines the Gaussian pyramid images of each layer into complete image data;

[0070] A point cloud construction module, which based on the complete image data, uses the SGM algorithm to calculate the disparity map, calculates the spatial coordinates of each pixel, and generates point cloud data;

[0071] A key point recognition module, which constructs a deep learning model to classify and segment the point cloud data;

[0072] A comprehensive eigenvalue calculation module, which based on the image data decomposed by the Gaussian pyramid, calculates gradient features and texture features, performs non-linear fusion on the eigenvalue of each layer, calculates the comprehensive eigenvalue, and extracts the comprehensive eigenvalue corresponding to the key point cloud data;

[0073] An abnormal warning module, which determines the abnormal threshold range according to the historical distribution of the comprehensive eigenvalue, marks the abnormal points and triggers a warning;

[0074] A report generation module, which generates a report according to the abnormal point list and the corresponding point cloud coordinates, marks the color of the abnormal points in the corresponding planar image, and integrates the records of the maintenance personnel on the abnormal types.

[0075] A computer device, including: a memory and a processor; the memory stores a computer program, and when the processor executes the computer program, the steps of the above-mentioned method for monitoring and identifying the running state of a train at a station based on images are implemented.

[0076] A computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the above-mentioned method for monitoring and identifying the running state of a train at a station based on images are implemented.

[0077] The beneficial effects of the present invention are as follows: Through CLAHE adaptive contrast adjustment, the local details of the image are significantly improved under different lighting conditions, making the areas with relatively low contrast clearer. Through multi-scale feature extraction, Gaussian pyramid decomposition generates image levels with different resolutions, ensuring that the structural information of the image is retained at different scales. Through the combination of SGM and RANSAC, the point cloud still has strong noise resistance in complex scenes, can adapt to environmental conditions such as lighting changes and reflection interference, and improves the adaptability of the point cloud in various scenarios. Based on the mapping of the comprehensive eigenvalue and the point cloud data, the point cloud data of the key components of the train can be efficiently extracted, thereby improving the recognition efficiency of the train operation state. BRIEF DESCRIPTION OF THE DRAWINGS

[0078] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0079] Figure 1 It is a schematic flow chart of the method for monitoring and recognizing the running state of a train at a station based on an image.

[0080] Figure 2 It is a schematic structural diagram of the system for monitoring and recognizing the running state of a train at a station based on an image. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0081] In order to make the above objects, features, and advantages of the present invention more obvious and understandable, the following will describe the specific embodiments of the present invention in detail with reference to the drawings of the specification.

[0082] In the following description, many specific details are set forth in order to fully understand the present invention. However, the present invention can also be implemented in other ways different from those described herein. Those skilled in the art can make similar generalizations without departing from the connotation of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed below.

[0083] Secondly, the so-called "one embodiment" or "embodiment" herein refers to a specific feature, structure, or characteristic that can be included in at least one implementation manner of the present invention. The phrase "in one embodiment" appearing in different places in this specification does not necessarily refer to the same embodiment, nor is it an independent or selectively exclusive embodiment from other embodiments.

[0084] Example 1, refer to Figure 1, which is the first embodiment of the present invention. This embodiment provides a method for monitoring and identifying the running state of a train at a station based on images. The method for monitoring and identifying the running state of a train at a station based on images includes,

[0085] S1, collect image data for two-dimensional wavelet decomposition, apply a soft threshold for denoising processing, and regenerate the aligned image data after denoising;

[0086] Preferably, collecting image data for two-dimensional wavelet decomposition, applying a soft threshold for denoising processing, and regenerating the aligned image data after denoising includes,

[0087] Collect image data of multiple perspectives of the train through a graphic acquisition camera;

[0088] Select the Daubechies wavelet as the wavelet basis function, perform two-dimensional wavelet decomposition on the image data, and the obtained sub-bands include low-frequency components and high-frequency components, representing different frequency details of the image data, expressed as:

[0089]

[0090] where φ and respectively represent the low-frequency and high-frequency filters, φ(x - m)·φ(y - n) represents calculating the distance between the filter center and the pixel coordinates (m, n) to perform low-frequency weighting on different positions of the image, represents calculating the distance between the filter center and the pixel coordinates (m, n) to perform high-frequency weighting on different positions of the image, I(m, n) represents the pixel value of the image at the coordinates (m, n), (x, y) represents the image coordinates, L represents the low-frequency component sub-band, containing the main structure information in the image, usually used to retain the general outline of the image in wavelet decomposition, H represents the high-frequency component sub-band, containing the detail information and edges in the image, usually used to enhance the detail part of the image, represents the summation operation for all pixel coordinates (m, n), traversing each pixel of the image, and is used to calculate the low-frequency or high-frequency components at all positions in the image;

[0091] Apply a soft threshold to the high-frequency component sub-band for denoising processing to eliminate image noise, expressed as:

[0092]

[0093] where λ represents the soft threshold, σ represents the standard deviation of the high-frequency sub-band, and N represents the total number of pixels in the sub-band;

[0094] Apply the soft threshold to each high-frequency coefficient of the high-frequency component sub-band. The high-frequency coefficient is the image signal value obtained from the high-frequency sub-band, expressed as:

[0095] T(x) = sign(x)·max(|x| - λ, 0);

[0096] Where T(x) represents the high-frequency coefficient value after soft-threshold denoising processing, λ represents the soft threshold, and sign(x) represents the sign function, indicating the positive or negative sign of x;

[0097] Based on the high-frequency coefficient value after denoising processing, perform inverse wavelet transform on the high-frequency component sub-band after denoising processing, and merge the decomposed data of each frequency band into the denoised aligned image data, which is expressed as:

[0098]

[0099] Where I d (x, y) represents the pixel value of the denoised image at the position (x, y); H d represents the high-frequency component sub-band after denoising, which contains the detailed information of the image.

[0100] By using Daubechies wavelet decomposition, reduce the random noise in the image. Without destroying the main structural information, enhance the edge and detail quality of the image. Through the soft-threshold denoising of high-frequency coefficients, it is possible to retain the detailed information with larger amplitudes and suppress the noise signals with smaller amplitudes, avoiding the edge blurring caused by traditional filtering methods (such as mean filtering). Based on the denoised high-frequency component sub-band and low-frequency component sub-band, re-synthesize the denoised image through inverse wavelet transform, providing high-quality input for the projection and comprehensive analysis of point cloud data.

[0101] S2. Construct a Gaussian pyramid of the aligned image data, perform equalization mapping through histogram equalization, and merge it into the complete image data based on each layer of the Gaussian pyramid;

[0102] Preferably, construct a Gaussian pyramid of the aligned image data, perform equalization mapping through histogram equalization, and merge it into the complete image data based on each layer of the Gaussian pyramid, including,

[0103] Construct a Gaussian pyramid of the aligned image data, decompose the aligned image data into different levels of multiple scales, where each layer represents the scale of the image at the k-th layer, which is expressed as:

[0104]

[0105] Where I k+1 (x, y) represents the image of the (k + 1)-th layer of the Gaussian pyramid, ω(m, n) represents the Gaussian weight coefficient, which is determined by the Gaussian function based on the Gaussian filter, and I k (2x + m, 2y + n) represents the pixel position offset of the image at the k-th layer. By offsetting 2 units in the x and y directions respectively with the new image coordinates, the downsampled image will be reduced to half of the original image;

[0106] For each layer of the Gaussian pyramid, the Contrast Limited Adaptive Histogram Equalization (CLAHE) method is used on the image. The image of each layer is divided into blocks, and the gray level distribution is statistically analyzed for each sub-block to generate a gray histogram H g , where g represents the gray level;

[0107] Based on the gray histogram H g Calculate the cumulative distribution function CDF, expressed as:

[0108]

[0109] where CDF g represents the cumulative pixel frequency from the minimum gray level 0 to the current gray level g, H g' represents the pixel frequency of the gray level g', where g' is the gray level variable used for summation starting from 0 and incrementing step by step to the current gray level g;

[0110] Determine the contrast limit threshold T based on the product of the clipping threshold ratio of the historical histogram and the total number of pixels in the sub-block;

[0111] Based on the contrast limit threshold T, if the pixel frequency H of the gray level g' in the gray histogram g' is greater than the contrast limit threshold T, then reduce H g' to T, and calculate the pixel frequency H g' of the gray level pixels exceeding the contrast limit threshold T, and evenly distribute them to all gray levels to obtain the adjusted gray level pixel frequency, and recalculate the cumulative distribution function CDF;

[0112] Calculate the equalization mapping according to the recalculated cumulative distribution function, expressed as:

[0113]

[0114] where P n (i, j) represents the equalized pixel value of the image block pixel (i, j), CDF gr represents the cumulative distribution function of the adjusted gray level pixel frequency, CDF min represents the minimum value of the non-zero cumulative distribution, L' represents the total number of gray levels, and v represents the total number of pixels in the sub-block;

[0115] Map the equalized values back to the corresponding pixels in the sub-block;

[0116] For the overlapping regions of adjacent sub-blocks, use bilinear interpolation for smooth transition, combine all sub-blocks, and generate the image data processed by CLAHE;

[0117] Based on the images of each layer of the Gaussian pyramid, starting from the image with the lowest resolution, it is upsampled layer by layer to a higher resolution through bilinear interpolation and finally merged into complete image data.

[0118] Through CLAHE adaptive contrast adjustment, the local details of the image are significantly improved under different lighting conditions, making the originally low-contrast areas clearer. Through multi-scale feature extraction, the Gaussian pyramid decomposition generates image levels with different resolutions, ensuring that the structural information of the image at different scales is retained. The contrast and feature information are evenly processed in the high-frequency and low-frequency regions. The multi-scale and CLAHE processing enhance the edge and detail information in the image, making the features of corresponding pixels in different views more consistent, thus making the calculation of the SGM matching cost more accurate and facilitating the subsequent use of RANSAC, which can more efficiently identify and eliminate incorrect matching points, ultimately optimizing the quality of the disparity map and improving the accuracy and robustness of the depth map.

[0119] Through the combination of CLAHE and the Gaussian pyramid, the image information of the left and right views becomes more uniform and detailed, significantly improving the efficiency and accuracy of the SGM matching cost calculation, thus supporting more accurate 3D point cloud analysis and state recognition. The local adaptive adjustment mechanism of CLAHE can improve the image quality under complex conditions such as uneven lighting and reflection interference. Combining the multi-scale representation of the Gaussian pyramid ensures the information balance of the image in the global and local structures, adapting to the image characteristics under different shooting conditions in the multi-view acquisition of trains and providing stable and high-quality input data.

[0120] S3. Based on the complete image data, use the SGM algorithm to generate a disparity map, use the random sample consensus algorithm to remove the matching points with large errors, and calculate the spatial coordinates of each pixel to form point cloud data, and construct a deep learning model to identify key point cloud data.

[0121] Preferably, based on the complete image data, use the SGM algorithm to generate a disparity map, use the random sample consensus algorithm to remove the matching points with large errors, and calculate the spatial coordinates of each pixel to form point cloud data, including

[0122] According to the complete image data processed by CLAHE collected from the left and right views, use the SGM algorithm to generate a disparity map, and calculate the matching cost of each pixel through the sum of absolute differences (SAD), expressed as:

[0123]

[0124] where D(x, y, d) represents the matching cost of the pixel at (x, y) when the disparity is d, and k represents the window size, which is used to control the matching range of the local area, and I L(x + i, y + j) represents the pixel value of the left view, I R (x + i - d, y + j) represents that the pixel value of the right view is offset by the disparity d;

[0125] The matching cost is optimized by accumulating the cost, expressed as:

[0126] C(x, y, d) = D(x, y, d) + min(C(x - 1, y, d), C(x, y - 1, d), C(x + 1, y, d));

[0127] Where C(x, y, d) represents the optimized matching cost of the disparity d at the position (x, y), and C(x - 1, y, d), C(x, y - 1, d) and C(x + 1, y, d) represent the matching costs of the disparity d at the adjacent positions (x - 1, y), (x, y - 1) and (x + 1, y) respectively;

[0128] Based on the minimum optimized matching cost, the disparity d assigns a disparity value to each pixel point to generate a preliminary disparity map.

[0129] Using the Random Sample Consensus algorithm RANSAC, based on the preliminary disparity map, the pixel point positions of the left and right view images are collected, and the matching points with large errors are detected and removed to generate an optimized disparity map;

[0130] Based on the disparity values at different positions of the optimized disparity map, the spatial coordinates of each pixel are calculated, expressed as:

[0131]

[0132]

[0133] Where Z represents the depth of the pixel point, f represents the camera focal length, B represents the baseline distance, and d(x, y) represents the disparity value of the pixel point (x, y); X and Y represent the abscissa and ordinate of the spatial coordinates respectively, and c x and c y represent the abscissa and ordinate of the camera principal point coordinates;

[0134] Traverse all pixels in the disparity map, calculate their corresponding spatial coordinates, and form point cloud data.

[0135] By optimizing the matching cost through the SGM algorithm, combining the cumulative cost calculation of adjacent pixels, making full use of the local and global information of the image, reducing the matching error, the generated disparity map is more accurate, can better reflect the depth information of the objects in the image, removes the incorrect matching points in the disparity map, improves the quality and robustness of the depth map, introduces the Random Sample Consensus (RANSAC) algorithm, effectively removes the incorrect matching points in the preliminary disparity map, and the generated optimized disparity map has higher accuracy and robustness, significantly improving the reliability of subsequent depth map and point cloud generation, reducing the matching error caused by noise and local feature inconsistency. The details of the generated point cloud data are rich and the accuracy is high, which can truly restore the three-dimensional shape of the key components of the train (such as pantographs and catenary wires), providing high-quality three-dimensional data input to support the precise analysis in subsequent train condition monitoring tasks. Through the combination of SGM and RANSAC, the point cloud still has strong anti-noise ability in complex scenes, can adapt to environmental conditions such as light changes and reflection interference, improves the adaptability of the point cloud in various scenarios, provides stable data support for multi-view monitoring of the train, reduces the interference of incorrect points on subsequent analysis (such as anomaly detection), and improves the credibility of the detection results. The SGM algorithm utilizes the local correlation and global constraints of image pixels in the cost calculation and optimization process, reduces the computational redundancy, and significantly improves the overall system efficiency, meeting the real-time requirements of train operation condition monitoring.

[0136] Furthermore, a deep learning model is constructed to identify key point cloud data, including

[0137] constructing a PointNet++ deep learning model, including an input layer, a local feature extraction layer, a feature propagation layer, a hierarchical feature aggregation layer, and an output layer;

[0138] wherein the input layer inputs the point cloud data, the local feature extraction layer selects the points within the local area using a distance-based grouping mechanism, the hierarchical feature aggregation layer aggregates the extracted local features layer by layer, the feature propagation layer uses a feature propagation mechanism to transfer the high-level feature information back to the low-level point cloud data, and the output layer outputs the prediction results based on the segmentation labels;

[0139] using the point cloud samples calibrated with labels as training data, selecting the cross-entropy loss function to calculate the difference between the class probabilities predicted by the model and the actual labels, using the Adam optimizer for optimization, and stopping the iteration to output the model when the loss calculated in the continuous iteration process no longer decreases significantly;

[0140] Input the newly calculated point cloud data into the model, and output the label probabilities identified as different key components of the train and the background. Select the maximum probability to clarify the label of the point cloud data, and select the point cloud data of the key components of the train as the key point cloud data.

[0141] The model can accurately identify key components of the train (such as pantographs, catenary conductors, etc.) and background points, thereby generating point-by-point segmentation labels. Through the segmentation results, redundant calculations of non-key points can be significantly reduced, providing high-quality input data for subsequent condition monitoring and anomaly detection. Through the distance-based grouping mechanism of the local feature extraction layer, points within the local area are dynamically selected, which can adapt to the sparsity and irregularity of point cloud data. Through the feature propagation layer, high-level features are propagated to low-level features, ensuring information flow and feature fusion between different levels of the point cloud. The PointNet++ model considers multi-scale characteristics during the hierarchical feature aggregation process, can adapt to point cloud data with different resolutions, and can significantly improve the accuracy, efficiency, and applicability of the system, providing comprehensive data support and technical guarantee for train operation safety.

[0142] S4. Determine the gradient feature and texture feature to calculate the comprehensive eigenvalue based on the image data decomposed by the Gaussian pyramid, perform anomaly judgment on the comprehensive eigenvalue corresponding to the key point cloud data, give an early warning based on the anomaly judgment, and generate an anomaly point report;

[0143] Preferably, determining the gradient feature and texture feature to calculate the comprehensive eigenvalue based on the image data decomposed by the Gaussian pyramid, and performing anomaly judgment on the comprehensive eigenvalue corresponding to the key point cloud data includes:

[0144] Determining the gradient feature and texture feature based on the image data decomposed by the Gaussian pyramid, which is expressed as:

[0145]

[0146] where G k (x, y) represents the gradient feature of the pixel point (x, y) on the k-th layer of the Gaussian pyramid, and I k (x, y) represents the image of the pixel point (x, y) on the k-th layer of the Gaussian pyramid, represents the horizontal gradient, represents the vertical gradient, and T k (x, y) represents the texture feature of the pixel point (x, y) on the k-th layer of the Gaussian pyramid, and p k (i, j) represents the frequencies of the gray values i and j appearing in the neighborhood window centered on the pixel point (x, y) on the k-th layer of the Gaussian pyramid, and (i - j) 2 represents the square of the difference in gray values, which is used to measure the degree of contrast of gray values;

[0147] Perform non-linear combination on the gradient feature and texture feature of each layer, which is expressed as:

[0148] F k (x, y) = ln(1 + G k (x, y) · T k(x, y));

[0149] where F k (x, y) represents the comprehensive eigenvalue of the k-th layer image. The comprehensive eigenvalues of all layers are weighted and fused to generate the full-resolution comprehensive eigenvalue;

[0150] Based on the coordinates of the key point cloud data, recalculate its pixel coordinates on the plane, and extract the eigenvalues of the pixel coordinates on the plane according to the full-resolution comprehensive eigenvalue;

[0151] Statistically analyze the historical distribution of the eigenvalues of the key point cloud data, and set the abnormal threshold range according to the difference between the eigenvalue mean and twice the standard deviation and the sum of the eigenvalue mean and twice the standard deviation;

[0152] If the eigenvalue of the extracted pixel coordinates on the plane is greater than or equal to the maximum value of the abnormal threshold range, record it as the abnormal point coordinates. If the eigenvalue of the extracted pixel coordinates on the plane is less than the minimum value of the abnormal threshold range, record it as the abnormal point coordinates.

[0153] By combining gradient features and texture features, it can accurately capture the edges, details, and texture information of the key components of the train, enhancing the multi-dimensional expressiveness of the image features. The extraction of gradient and texture features enhances the recognition ability of the details and edge features of the local areas of the image (such as pantographs and catenary wires), improving the accurate monitoring of the key components of the train. Through the calculation of the full-resolution comprehensive eigenvalue and the dynamic setting of the abnormal threshold range, it can accurately identify and locate the abnormal points in the image, reducing false alarms and missed alarms. Through the multi-level feature extraction of the Gaussian pyramid, it ensures that the image information at different scales can be fully retained, adapting to the changes of train images under different resolutions and lighting conditions, and improving the robustness of the system. Based on the mapping of the comprehensive eigenvalue and the point cloud data, it can efficiently extract the point cloud data of the key components of the train, providing accurate data support for subsequent status monitoring and anomaly detection. The CLAHE method and multi-scale processing improve the visibility of the image under different lighting conditions, can effectively enhance the local details of the image in a complex lighting environment, provide efficient and reliable status feedback of the key components, and contribute to improving the running safety of the train.

[0154] Furthermore, based on the anomaly judgment, give an early warning and generate an abnormal point report, including,

[0155] Form an abnormal point list with the point cloud coordinates and corresponding eigenvalues of the abnormal point coordinates, and color-mark the plane image according to the abnormal point coordinates;

[0156] And give an early warning to the maintenance personnel according to the sound and light alarm, and make specific judgments and records on the abnormal types by the maintenance personnel;

[0157] Use a table generation tool to generate an anomaly point report based on the anomaly point list, the color - marked planar image, and the maintenance records and anomaly type records of the maintenance personnel.

[0158] By forming an anomaly point list with the anomaly point coordinates and the corresponding eigenvalue, the position of the anomaly point can be accurately located, and the anomaly area can be visually displayed in the planar image through color marking. Through color marking, the maintenance personnel can clearly identify the anomaly points on the image, improving the intuitiveness and real - time nature of anomaly detection. Based on the marking of the anomaly point coordinates and eigenvalues, combined with the sound and light alarm system, an instant warning is provided for the maintenance personnel, enabling the system to react quickly when an anomaly is detected. The sound and light alarm helps to quickly attract the attention of the maintenance personnel, shorten the response time, and ensure the timely discovery and solution of potential problems in train operation. The anomaly point report not only records the position of the anomaly point but also includes detailed information on eigenvalues and anomaly types. The maintenance personnel can make more accurate judgments and processing based on these data. By automatically generating the anomaly point report using the table generation tool, combining the anomaly point information, planar image, and maintenance records provides a detailed and easily traceable report for management and record - keeping.

[0159] Example 2, refer to Figure 2 , which is the second embodiment of the present invention. This embodiment is different from the previous one and provides a system for monitoring and identifying the running state of a train at a station based on images, including,

[0160] An image pre - processing module that collects multi - perspective image data and performs pre - processing, conducts two - dimensional wavelet decomposition, separates low - frequency and high - frequency components, applies soft thresholding for denoising, and generates denoised and aligned image data;

[0161] A multi - scale feature enhancement module that constructs a Gaussian pyramid of the aligned image data, generates multi - level image representations, performs histogram equalization on each layer of the image to enhance local contrast, and combines the Gaussian pyramid images of each layer into complete image data;

[0162] A point cloud construction module that, based on the complete image data, uses the SGM algorithm to calculate the disparity map, calculates the spatial coordinates of each pixel, and generates point cloud data;

[0163] A key point recognition module that constructs a deep learning model to classify and segment the point cloud data;

[0164] A comprehensive eigenvalue calculation module that, based on the image data decomposed by the Gaussian pyramid, calculates gradient features and texture features, performs non - linear fusion on each layer of eigenvalues, calculates the comprehensive eigenvalue, and extracts the comprehensive eigenvalue corresponding to the key point cloud data;

[0165] An anomaly warning module that determines the anomaly threshold range according to the historical distribution of the comprehensive eigenvalue, marks the anomaly points, and triggers an alarm;

[0166] A report generation module generates a report according to the list of abnormal points and the corresponding point cloud coordinates, marks the colors of the abnormal points in the corresponding planar image, and integrates the records of the maintenance personnel on the abnormal types.

[0167] If the above-mentioned functions are implemented in the form of software function units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The foregoing storage medium includes: USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs and other various media that can store program codes.

[0168] The logic and / or steps represented in the flowchart or described in other ways herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device), or in conjunction with these instruction execution systems, apparatus, or devices. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device.

[0169] More specific examples (a non-exhaustive list) of computer-readable media include the following: an electrical connection part with one or more wirings (electronic device), a portable computer disk cartridge (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). Additionally, the computer-readable medium can even be paper or other suitable media on which the program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other media, followed by editing, interpretation, or, if necessary, other suitable processing, and then stored in a computer memory.

[0170] It should be understood that various parts of the present invention can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), etc.

[0171] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention, and they should all be covered within the scope of the claims of the present invention.

Claims

1. An image-based train operation status monitoring and identification method, characterized in that: include, Collect image data for two-dimensional wavelet decomposition, apply soft thresholding for denoising and regenerate aligned image data after denoising; Construct a Gaussian pyramid of aligned image data, perform equalization mapping through histogram equalization, and merge each layer of the Gaussian pyramid into complete image data; Based on the complete image data, the SGM algorithm is used to generate a disparity map, the random sampling consistency algorithm is used to remove matching points with large errors, and the spatial coordinates of each pixel are calculated to form point cloud data. A deep learning model is built to identify key point cloud data. Based on the image data decomposed by Gaussian pyramid, the gradient features and texture features are determined to calculate the comprehensive feature values, and the comprehensive feature values ​​corresponding to the key point cloud data are used for abnormal judgment. Based on the abnormal judgment, an early warning is issued and an abnormal point report is generated; The construction of a deep learning model to identify key point cloud data includes: Build a PointNet++ deep learning model, including input layer, local feature extraction layer, feature propagation layer, hierarchical feature aggregation layer, and output layer; The input layer inputs point cloud data, the local feature extraction layer uses a distance-based grouping mechanism to select points in the local area, the hierarchical feature aggregation layer aggregates the extracted local features layer by layer, the feature propagation layer uses a feature propagation mechanism to pass high-level feature information back to the low-level point cloud data, and the output layer outputs the prediction results based on the segmentation label; Use the labeled point cloud samples as training data, select the cross entropy loss function to calculate the difference between the category probability predicted by the model and the actual label, and use the Adam optimizer for optimization. If the calculated loss no longer decreases significantly during the continuous iteration process, stop iterating and output the model. Input the newly calculated point cloud data into the model, and output the label probabilities of being identified as different key train components and backgrounds. Select the label of the point cloud data with the largest probability, and select the point cloud data of the key train components as the key point cloud data. The method of determining the gradient features and texture features based on the image data decomposed by the Gaussian pyramid to calculate the comprehensive feature value, and making an abnormal judgment on the comprehensive feature value corresponding to the key point cloud data, includes: The gradient features and texture features are determined based on the image data decomposed by Gaussian pyramid, which can be expressed as: Among them G k (x, y) represents the gradient feature of the pixel (x, y) at the kth layer of the Gaussian pyramid, I k (x,y) represents the image of the pixel (x,y) at the kth level of the Gaussian pyramid. represents the horizontal gradient, represents the vertical gradient, T k (x, y) represents the texture feature of the pixel (x, y) at the kth layer of the Gaussian pyramid, p k (i,j) represents the frequency of occurrence of grayscale values ​​i and j in the neighborhood window centered at the pixel point (x,y) at the kth layer of the Gaussian pyramid, (ij) 2 Represents the square of the gray value difference, which is used to measure the contrast of gray values; The gradient features and texture features of each layer are nonlinearly combined, expressed as: F k (x,y)=ln(1+G k (x,y)·T k (x,y)); where F k (x, y) represents the comprehensive feature value of the k-th layer image. The comprehensive feature values ​​of all layers are weighted fused to generate the full-resolution comprehensive feature value; Recalculate the pixel coordinates on the plane based on the coordinates of the key point cloud data, and extract the eigenvalues ​​of the plane pixel coordinates based on the full-resolution comprehensive eigenvalues; The historical distribution of the eigenvalues ​​of key point cloud data is counted, and the abnormal threshold range is set according to the difference between the eigenvalue mean and twice the standard deviation and the sum of the eigenvalue mean and twice the standard deviation; If the feature value of the extracted plane pixel coordinates is greater than or equal to the maximum value of the abnormal threshold range, it is recorded as the abnormal point coordinates. If the feature value of the extracted plane pixel coordinates is less than the minimum value of the abnormal threshold range, it is recorded as the abnormal point coordinates.

2. The image-based train operation status monitoring and identification method according to claim 1, characterized in that: The collected image data is subjected to two-dimensional wavelet decomposition, a soft threshold is applied to perform denoising processing and the denoised aligned image data is regenerated, including: Collect image data from multiple perspectives of the train through a graphics acquisition camera; Select Daubechies wavelet as the wavelet basis function, perform two-dimensional wavelet decomposition on the image data, and the obtained sub-bands include low-frequency components and high-frequency components, representing different frequency details of the image data; Apply soft threshold to the high-frequency component sub-band for denoising to eliminate image noise, which is expressed as: Where λ represents the soft threshold, σ represents the standard deviation of the high-frequency subband, and N represents the total number of pixels in the subband; Apply a soft threshold to each high-frequency coefficient of the high-frequency component subband, expressed as: T(x)=sign(x)·max(|x|-λ,0); Where T(x) represents the high-frequency coefficient value after soft threshold denoising, λ represents the soft threshold, and sign(x) represents the sign function, indicating the positive or negative sign of x; Based on the high-frequency coefficient value after denoising, the high-frequency component subband after denoising is inversely transformed, and the decomposed frequency band data are merged into the aligned image data after denoising.

3. The image-based train operation status monitoring and identification method according to claim 2, characterized in that: The Gaussian pyramid of the aligned image data is constructed, and the equalization mapping is performed through histogram equalization, and each layer of the Gaussian pyramid is merged into complete image data. include, Construct a Gaussian pyramid of aligned image data and decompose the aligned image data into different levels of multiple scales, where each level represents the scale of the image at the kth level, expressed as: Among them I k+1 (x, y) represents the image of the k+1th layer of the Gaussian pyramid, ω(m, n) represents the Gaussian weight coefficient; The contrast limited adaptive histogram equalization (CLAHE) method is used for the image of each layer of the Gaussian pyramid. Each layer of the image is divided into blocks, and the gray level distribution of each sub-block is statistically analyzed to generate a gray level histogram Hg, where g represents the gray level. Based on the grayscale histogram H g Calculate the cumulative distribution function CDF, expressed as: Where CDF g It represents the cumulative pixel frequency from the minimum gray level 0 to the current gray level g, H g' represents the pixel frequency of gray level g', g' represents the gray level variable used for summation starting from 0 and gradually accumulating to the current gray level g; Determine a contrast limit threshold T based on the product of the clipping threshold ratio of the history histogram and the total number of pixels of the sub-block; Based on the contrast limit threshold T, if the pixel frequency H of gray level g' in the gray histogram g' If it is greater than the contrast limit threshold T, H g' Cut to T and calculate the pixel frequency H g' The grayscale pixel frequency that exceeds the contrast limit threshold T is evenly distributed to all gray levels, the adjusted grayscale pixel frequency is obtained, and the cumulative distribution function CDF is recalculated; The equalization mapping is calculated based on the recalculated cumulative distribution function, expressed as: Where P n (i, j) represents the equalized pixel value of the image block pixel (i, j), CDF gr Represents the cumulative distribution function of the adjusted grayscale pixel frequency, CDF min represents the minimum value of the non-zero cumulative distribution, L' represents the total number of gray levels, and v represents the total number of pixels in the sub-block; Map the equalized values ​​back to the corresponding pixels in the sub-block; For the overlapping areas of adjacent sub-blocks, bilinear interpolation is used to smoothly transition, and all sub-blocks are combined to generate image data processed by CLAHE; Based on the image of each layer of the Gaussian pyramid, starting from the lowest resolution image, it is upsampled layer by layer to a higher resolution through bilinear interpolation and finally merged into the complete image data.

4. The image-based train operation status monitoring and identification method according to claim 3, characterized in that: Based on the complete image data, the disparity map is generated using the SGM algorithm, the matching points with large errors are removed using the random sampling consistency algorithm, and the spatial coordinates of each pixel are calculated to form point cloud data, including: According to the complete image data collected from the left and right perspectives and processed by CLAHE, the SGM algorithm is used to generate a disparity map, and the matching cost of each pixel is calculated by the absolute difference cost SAD, which is expressed as: Where D(x,y,d) represents the matching cost of the pixel at (x,y) when the disparity is d, k represents the window size, which is used to control the matching range of the local area, and I L (x+i, y+j) represents the pixel value of the left view, I R (x+id, y+j) indicates that the pixel value of the right view is offset by the disparity d; Optimize matching cost by accumulating cost; Based on the disparity d with the minimum optimized matching cost, a disparity value is assigned to each pixel to generate a preliminary disparity map. The random sampling consensus algorithm RANSAC is used to collect the pixel positions of the left and right view images based on the preliminary disparity map, detect and remove matching points with large errors, and generate an optimized disparity map. Based on the disparity values ​​at different positions of the optimized disparity map, the spatial coordinates of each pixel are calculated, expressed as: Where Z represents the depth of the pixel, f represents the focal length of the camera, B represents the baseline distance, d(x,y) represents the disparity value of the pixel (x,y); X and Y represent the horizontal and vertical coordinates of the spatial coordinates, respectively, and c x and c y Represents the horizontal and vertical coordinates of the camera's principal point coordinates; Traverse all pixels in the disparity map, calculate their corresponding spatial coordinates, and form point cloud data.

5. The image-based train operation status monitoring and identification method according to claim 4, characterized in that: The method of providing an early warning based on abnormality judgment and generating an abnormality point report includes: The point cloud coordinates and corresponding eigenvalues ​​of the abnormal point coordinates are combined into an abnormal point list, and the plane image is color-marked according to the abnormal point coordinates; And according to the sound and light alarm, the maintenance personnel are warned, and the maintenance personnel make specific judgments and records on the abnormal types; Use the table generation tool to generate anomaly reports based on the anomaly list, color-coded plane images, and the maintenance personnel's maintenance records and anomaly type records.

6. A system based on the image-based train operation status monitoring and identification method according to any one of claims 1 to 5, characterized in that: include, The image preprocessing module collects and preprocesses multi-view image data, performs two-dimensional wavelet decomposition, separates low-frequency and high-frequency components, applies soft thresholding for denoising, and generates denoised aligned image data; The multi-scale feature enhancement module constructs a Gaussian pyramid of aligned image data, generates a multi-level image representation, performs histogram equalization on each layer of the image, improves the local contrast, and merges the Gaussian pyramid images of each layer into a complete image data; Point cloud construction module, based on the complete image data, uses the SGM algorithm to calculate the disparity map, calculate the spatial coordinates of each pixel, and generate point cloud data; Key point recognition module, builds a deep learning model to classify and segment point cloud data; The comprehensive eigenvalue calculation module calculates the gradient features and texture features based on the image data decomposed by the Gaussian pyramid, performs nonlinear fusion on the eigenvalues ​​of each layer, calculates the comprehensive eigenvalues, and extracts the comprehensive eigenvalues ​​corresponding to the key point cloud data; The abnormal warning module determines the abnormal threshold range based on the historical distribution of the comprehensive characteristic value, marks the abnormal point and triggers the warning; The report generation module generates a report based on the abnormal point list and the corresponding point cloud coordinates, marks the abnormal points with colors in the corresponding plane images, and integrates the maintenance personnel's records of abnormal types.

7. A computer device comprising: Memory and processor; The memory stores a computer program, characterized in that: when the processor executes the computer program, the steps of the image-based train operation status monitoring and identification method in any one of claims 1 to 5 are implemented.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the image-based train operation status monitoring and identification method according to any one of claims 1 to 5 are implemented.

Citation Information

Patent Citations

  • Method of identifying spinal sagittal image anomaly and computing equipment

    CN108830835A

  • Lung image segmentation method and system

    CN118115742A