Three-dimensional train modeling method and system based on binocular parallax prediction model

By combining a binocular parallax prediction model with a linear array camera, the frame rate limitation and environmental vibration impact of 3D measurement in rail transit were solved, achieving high-precision 3D reconstruction of trains, especially stable data acquisition under high-speed and complex lighting conditions.

CN121883704APending Publication Date: 2026-04-17CRRC QINGDAO SIFANG ROLLING STOCK RESEARCH INSTITUTE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CRRC QINGDAO SIFANG ROLLING STOCK RESEARCH INSTITUTE CO LTD
Filing Date
2025-12-18
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

Existing 3D measurement technologies in rail transit are limited in terms of acquisition and processing speed in high-speed scenarios. Laser triangulation is prone to edge breakage and distortion in occluded areas under a single viewpoint. The detection accuracy of low-texture areas is not ideal, and environmental vibration affects the measurement quality.

Method used

A binocular parallax prediction model is adopted, which uses two parallel linear array cameras for scanning. Pixel-level parallax estimation is performed by combining a spatiotemporal transformation network and a regression prediction network. A linear laser/LED light source is used to suppress ambient light interference, and vibration offset is compensated by longitudinal parallax component to achieve three-dimensional coordinate reconstruction.

Benefits of technology

It solves the frame rate limitation problem of traditional methods, improves the integrity of 3D reconstruction of edges and occluded areas, and enhances measurement accuracy and robustness in high-speed and vibration environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121883704A_ABST
    Figure CN121883704A_ABST
Patent Text Reader

Abstract

The embodiment of the invention relates to a method and system for carrying out train three-dimensional modeling based on a binocular parallax prediction model. The method comprises the steps that installation, collection and splicing rules of double parallel line-scan digital cameras are set; setting a binocular parallax prediction model; training data collection is carried out based on installation, collection and splicing rules to obtain a first data set; training a binocular parallax prediction model based on the first data set; after training is finished, double parallel arrangement linear array cameras are installed; when a train passes through the two installed cameras on the current train track, image collection and splicing are carried out based on collection and splicing rules to obtain images I1 and I2, the images I1 and I2 are input into the binocular parallax prediction model for prediction to obtain a parallax image D1-2, and three-dimensional modeling is carried out on the train passing this time based on I1, I2 and D1-2. The method can improve the prediction precision in a high-speed scene and a complex illumination environment, can effectively overcome the limitation of a single visual angle, and improves the three-dimensional reconstruction integrity of edges and shielding areas.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of rail transit technology, and in particular to a method and system for three-dimensional modeling of trains based on a binocular parallax prediction model. Background Technology

[0002] The most commonly used 3D measurement technology in trackside inspection systems is laser triangulation, typically implemented using a "laser line + area array camera + laser triangulation" scheme: one or more laser lines are placed along the trackside, and an area array camera is used to collect the stripes on the train surface illuminated by the laser at a certain angle. Then, laser triangulation is used to reconstruct the 3D information of the train's outline for subsequent dimensional inspection and fault analysis. This traditional method has significant drawbacks: 1) The acquisition and processing speed of the area array camera is limited, making stable acquisition and reconstruction difficult in high-speed scenarios; 2) Laser triangulation relies on laser stripe extraction from a single viewpoint, which is prone to breakage or distortion at the edges of the target object and in occluded areas, making it difficult to meet high-precision measurement requirements; 3) In low-texture areas or in scenarios with strong light interference, the output detection accuracy and robustness are not ideal; 4) Traditional schemes usually assume continuous camera stability, but in real-world scenarios, there are many vibration interferences, such as ground vibrations caused by passing trains, which degrades the output detection quality of traditional schemes in such situations. Summary of the Invention

[0003] The purpose of this invention is to address the shortcomings of existing technologies by providing a method and system for three-dimensional modeling of trains based on a binocular parallax prediction model. This invention employs dual parallel linear array cameras for scanning, utilizing the line scanning characteristics of the cameras to achieve continuous line-by-line acquisition of high-speed moving trains. Combined with a first acquisition rule, it achieves synchronous acquisition throughout the train's passage, fundamentally solving the frame rate limitation problem of traditional solutions and ensuring stable data acquisition in high-speed scenarios. A binocular disparity prediction model is used for pixel-level disparity estimation. A spatiotemporal transformation network learns the global correspondence of the train's surface structure, and multi-scale feature fusion is combined with a regression prediction network, effectively overcoming the limitations of a single viewpoint and significantly improving the completeness of 3D reconstruction of edges and occluded areas. The binocular disparity prediction model achieves robust feature matching for low-texture areas through spatiotemporal joint coding layers and correlation coding layers. The reliability of the model output can be used as a quality assessment indicator. Combined with a dedicated line laser / line LED light source illumination scheme, it can effectively suppress ambient light interference and ensure measurement accuracy under various lighting conditions. In the 3D modeling stage, 3D coordinate reconstruction is performed based on a binocular disparity depth calculation method, and the longitudinal disparity component is used to compensate for positional offset information caused by vibration, thereby improving robustness under vibration environments.

[0004] In view of this, a first aspect of the present invention provides a method for three-dimensional modeling of trains based on a binocular parallax prediction model, the method comprising: The installation rules, image acquisition rules, and image stitching rules of the dual parallel linear array cameras are set to obtain the corresponding first installation rules, first acquisition rules, and first stitching rules; A corresponding binocular disparity prediction model is set up for the stitched images of dual parallel linear array cameras; the binocular disparity prediction model is used to perform pixel-level binocular disparity prediction based on the input images I1 and I2 and output the corresponding disparity map D. 1-2 ; The first dataset is obtained by collecting training data based on the first installation rule, the first collection rule, and the first splicing rule. The binocular disparity prediction model is trained based on the first dataset; After the model training is completed, any train track is taken as the ground motion trajectory of the current scanning object, and any train passing on the current train track is taken as the current scanning object; and a dual parallel linear array camera is installed on the side of the current train track based on the first installation rule. When a train passes over two cameras on the current track, images are acquired and stitched together based on the first acquisition rule and the first stitching rule to obtain a current set of images I1 and I2; and the current images I1 and I2 are input into the binocular disparity prediction model to predict the corresponding disparity map D. 1-2 Based on the image I1, the image I2, and the disparity map D... 1-2 A 3D model of the train that passed through this time was created.

[0005] Preferably, the first installation rule is as follows: two identical line scan cameras are selected to form a dual parallel line scan camera; the two cameras are installed on the same side of the ground motion trajectory of the scanned object; the horizontal distance between the center points of the two cameras and the motion trajectory of the scanned object is equal; the vertical height between the center points of the two cameras and the ground is a preset installation height H; the horizontal distance between the center points of the two cameras is fixed as a preset camera spacing L; the vertical direction of the ground motion trajectory of the scanned object is the line scan direction of the two cameras; each of the two cameras is equipped with a corresponding line laser light source or line LED light source; the image acquisition operation of each camera is synchronized with the illumination operation of its corresponding light source; the image acquisition operations of the two cameras are synchronized. The first acquisition rule is as follows: when the scanned object moves towards the two cameras, the camera closer to the scanned object is designated as the first camera, and the other camera is designated as the second camera; when the distance between the scanned object and the first camera is lower than a preset first distance threshold, the image acquisition operation of both cameras is simultaneously started, and the illumination operation of the two corresponding light sources is simultaneously started; when the shortest distance between the scanned object and the second camera is not lower than the first distance threshold, the image acquisition operation of both cameras is simultaneously stopped, and the illumination operation of the two corresponding light sources is simultaneously stopped; and the two image sequences acquired by the first and second cameras during the current passage time of the scanned object are designated as the corresponding image sequences P1 and P2; wherein, each of the image sequences P1 and P2 is associated with T corresponding images p 1,t p 2,t Composition, 1 ≤ index t ≤ T, where T is the total number of samples collected; the graph p 1,i p 2,i The image format is the line scan format of the line scan camera. The image shape is H0×W0×D0, where H0, W0, and D0 are the corresponding line scan height, line scan width, and line scan pixel feature dimension, respectively. D0=3. The pixel features of the line scan are composed of the corresponding RGB three primary color features. The first stitching rule is: according to the image stitching principle of the line scan camera, stitch the T images p of the image sequence P1 along the image height direction. 1,t The corresponding image I1 is obtained by stitching the images line by line, and the T images p of the image sequence P2 are stitched together along the image height direction. 2,t Image I2 is obtained by stitching each row together. Both images I1 and I2 have a shape of H1×W1×D1, where H1, W1, and D1 are the first height, first width, and first pixel feature dimension, respectively. H1=T×H0, W1=W0, and D1=D0. Images I1 and I2 are each composed of H1×W1 pixels. , Composition; 1 ≤ horizontal coordinate x ≤ W1, 1 ≤ vertical coordinate y ≤ H1; each pixel , The pixel features are all composed of a corresponding set of three primary color features ( , , ), ( , , )composition; The disparity map D 1-2 The image shape is H out ×W out ×D out H out W out D outThese represent the corresponding disparity map height, disparity map width, and disparity map pixel feature dimension, H. out =H1, W out =W1,D out =3, from the corresponding H out ×W out 1 pixel Composition; the pixel It consists of a set of corresponding horizontal disparity Δx, vertical disparity Δy, and confidence level c; for any pixel point on the image I1 In other words, it is related to the pixels on the image I2. Corresponding to the same real point; The first dataset includes multiple first data records; each first data record consists of a set of images I1, I2 and their corresponding label disparity maps D. tag Composition; the label disparity map D tag The data format is consistent with that of the disparity map D1.

[0006] Preferably, the input terminal of the binocular disparity prediction model is used to receive the image I1 and the image I2, and the output terminal is used to output the corresponding disparity map D. 1-2 ; The binocular disparity prediction model is composed of a spatiotemporal transformation network and a regression prediction network connected sequentially. The spatiotemporal transformation network is composed of a feature extractor, a spatiotemporal coding embedding layer, a spatiotemporal joint coding layer, and a spatiotemporal feature decoding layer connected sequentially. The regression prediction network is composed of a correlation coding layer, a multi-scale feature fusion layer, a global modeling layer, and a disparity prediction layer connected sequentially. The spatiotemporal transformation network is used to extract features from images I1 and I2 to obtain corresponding feature vectors H1 and H2, respectively; to fuse the feature vectors H1 and H2, to embed the fused features using spatiotemporal encoding, and to perform sequence transformation on the embedded features to obtain the corresponding sequence S1; to encode the sequence S1 to obtain the corresponding sequence S2; and to decompose the sequence S2 into feature vectors and upsample the two decomposed vectors to obtain the corresponding feature vectors X1 and X2, which are then sent to the regression prediction network. The regression prediction network is used to identify the local correlation of feature vector X2 with feature vector X1 as the query, and obtain the corresponding correlation vector C. 1-2 And the feature vectors X1 and X2 and the correlation vector C are compared. 1-2 The corresponding feature vector X is obtained by fusion. C ; and for the feature vector X C Multi-scale feature extraction is performed, and feature fusion is applied to obtain the corresponding feature vector H. mul ; and based on the feature vector Hmul Global feature modeling is performed to obtain the corresponding feature vector H. glb ; and based on the feature vector H glb Disparity map D is obtained by performing disparity map prediction. 1-2 And output it.

[0007] Preferably, the feature extractor is implemented based on a CNN network or a residual network; the feature extractor is used to extract features from the image I1 and the image I2 respectively to obtain corresponding feature vectors H1 and H2, which are then sent to the spatiotemporal coding embedding layer; Wherein, the vector shapes of the feature vectors H1 and H2 are both H2×W2×D2, where H2, W2, and D2 are the corresponding second height, second width, and second pixel feature dimensions, respectively, and H2 < H1, W2 < W1, D2 > D0; The spatiotemporal coding embedding layer is used to perform feature fusion on the feature vectors H1 and H2 according to the feature channel concatenation method to obtain a feature vector H3 with shape H2×W2×2D2; and based on the feature data h of the feature vector H3... i,j,k The height coordinates i, the second height H2, and the total number of data collected T are used to confirm the index t corresponding to the current feature data, and based on each feature data h... i,j,k The corresponding camera number is set according to the correspondence between the feature vector H1 or the feature vector H2; and based on each feature data h i,j,k The corresponding index t, position coordinates (i,j,k), and camera number are used to initialize the corresponding time embedding code, position embedding code, and camera number embedding code. The corresponding code pe is obtained by adding the three types of embedding codes. i,j,k ; and composed of each of the aforementioned feature data h i,j,k and its corresponding encoding pe i,j,k Generate corresponding feature data α is a preset scaling factor; and the obtained H2×W2×2D2 feature data The corresponding embedded feature vector H4 is formed; and H2×W2 sub-vectors with a length of 2D2 are extracted from the embedded feature vector H4, and the H2×W2 sub-vectors are used to form the corresponding sequence S1 and sent to the spatiotemporal joint coding layer. Wherein, the vector shapes of both the feature vector H3 and the embedded feature vector H4 are H2×W2×2D2, 1≤indexi≤H2, 1≤indexj≤W2, 1≤indexk≤2D2; the feature vector H3 consists of H2×W2×2D2 feature data h i,j,k The embedded feature vector H4 is composed of H2×W2×2D2 feature data. Composition; the sequence S1 includes H2×W2 sub-vectors of length 2D2. , 1 ≤ index q ≤ H2 × W2; The spatiotemporal joint coding layer is implemented based on the encoder of the Transformer model; the spatiotemporal joint coding layer is used to perform multi-head self-attention coding on the sequence S1 and send the resulting feature vector sequence as the corresponding sequence S2 to the spatiotemporal feature decoding layer; The sequence S2 includes H2×W2 sub-vectors of length 2D2. ; The spatiotemporal feature decoding layer is used to convert the sequence S2 into a feature vector H5 with shape H2×W2×2D2; and extract the H2×W2×D2 feature data corresponding to the feature vectors H1 and H2 respectively from the feature vector H5 to form two decomposition vectors H6 and H7 with shape H2×W2×D2; and upsample the decomposition vectors H6 and H7 respectively to obtain the corresponding feature vectors X1 and X2, which are then sent to the correlation coding layer; The feature vectors X1 and X2 both have a shape of H3×W3×D3, where H3, W3, and D3 are the corresponding third height, third width, and third pixel feature dimension, respectively, and H3=H1, W3=W1, and D3=D2. Each feature vector X1 and X2 has three corresponding H3×W3 sub-vectors of length D3. , composition.

[0008] Preferably, the correlation encoding layer is used to encode each of the sub-vectors of the feature vector X1. h as the current query vector que ; and the current query vector h que The corresponding coordinates x and y are denoted as x. que y que Based on two preset local incremental parameters Δx * , △y * Set the local alignment range in feature vector X2 to the corresponding [x que -△x * ,x que +△x * ]、[y que -△y * ,y que +△y * ]; and for the current query vector h que Each sub-vector of the local comparison range The cosine similarity is calculated, and the resulting 2(Δx) is used to calculate the similarity. * +△y* +1) similarity values ​​form the corresponding relevance vector c. x,y ; and the resulting H3×W3 correlation vectors c x,y The corresponding correlation vector C is formed. 1-2 And according to the feature channel concatenation method, the feature vector X1, the feature vector X2, and the correlation vector C are concatenated. 1-2 Feature fusion is performed to obtain the corresponding feature vector X. C Send to the multi-scale feature fusion layer; Among them, the two local incremental parameters Δx * , △y * All are preset positive integers; the correlation vector C 1-2 The vector shape is H4×W4×D4, where H4, W4, and D4 are the corresponding fourth height, fourth width, and fourth pixel feature dimension, respectively. H4=H3, W4=W3, and D4=2(Δx). * +△y * +1); the feature vector X C The vector shape is H5×W5×D5, where H5, W5, and D5 are the corresponding fifth height, fifth width, and fifth pixel feature dimension, respectively. H5=H3, W5=W3, and D5=D3+D3+2(△x) * +△y * +1); The multi-scale feature fusion layer is implemented based on the lightweight U-Net model; the multi-scale feature fusion layer is used to apply the downsampling network of the U-Net model to the feature vector X. C Multi-scale feature extraction is performed, and the upsampling network of the U-Net model is used to fuse all the obtained multi-scale features to obtain the corresponding feature vector H. mul Send to the global modeling layer; Wherein, the feature vector H mul The vector shape is H6×W6×D6, where H6, W6, and D6 are the corresponding sixth height, sixth width, and sixth pixel feature dimension, respectively. H6=H5, W6=W5, and D6≤D5. The global modeling layer is used to process the feature vector H. mul Global pooling yields a pooling vector H of shape D6×1. pool ; and the pooling vector H pool The mapping is a weight matrix A of shape H6×W6. pool ; and using the weight matrix A pool For the feature vector H mul The corresponding feature vector H is obtained by weighting. glb Send to the disparity prediction layer; Wherein, the weight matrix A pool The mapping method is as follows: ; W g1 W g1 B represents the first and second weight matrices corresponding to the global modeling layer. g1 B g2 The first and second bias vectors corresponding to the global modeling layer; the first weight matrix W g1 The shape is H6×D6, and the first bias vector is B. g1 The shape is H6×1, and the vector... The shape is H6×1; ReLU() is the ReLU activation function; the second weight matrix W g2 The shape is W6×1, and the second bias vector is B. g2 The shape is H6×W6; the weight matrix A pool Its shape is H6×W6; The feature vector H glb The weighting method is as follows: ; Sigmoid() is the Sigmoid activation function; The disparity prediction layer is used to predict the disparity using three built-in independent convolutional layers based on the feature vector H. glb Predicting the three features of the disparity map—lateral disparity, vertical disparity, and confidence—results in three corresponding shapes, H. out ×W out The three predicted vectors are multiplied by 1, and the resulting disparity map D is composed of these three predicted vectors. 1-2 And output it.

[0009] Preferably, the step of collecting training data based on the first installation rule, the first acquisition rule, and the first splicing rule to obtain the corresponding first dataset specifically includes: Select one or more train track segments to form a track set; and use each train track segment of the track set as the current track; and install dual parallel linear array cameras on the side of the current track based on the first installation rule; and use any train passing on the current track as the current scanning object; and each time a train passes between the two installed cameras on the current track, perform an image acquisition and stitching based on the first acquisition rule and the first stitching rule to obtain a set of corresponding images I1 and I2 to form a corresponding first image group; and use all the obtained first image groups to form a corresponding first image set; And take each of the first image groups in the first image set as the current image group; and use preset 3D modeling tools and optical flow analysis tools, according to the disparity map D 1-2 The data format constructs a corresponding disparity map based on images I1 and I2 of the current image group; and performs binarization on the credibility of each pixel in the current disparity map based on a preset first credibility threshold. If the credibility is less than the first credibility threshold, it is reset to 0; if the credibility is greater than or equal to the first credibility threshold, it is reset to 1; and the current disparity map with completed credibility binarization is used as the corresponding label disparity map D. tag ; and the current image group and its corresponding label disparity map D tag The first data record is formed into a corresponding data record; and the first dataset is formed from all the first data records obtained; wherein the three-dimensional modeling tools include at least OpenCV, COLMAP, and Meshroom; and the optical flow analysis tools include at least DaVis, Halcon, VisionPro, RAFT model, and FlowNet2 model.

[0010] Preferably, training the binocular disparity prediction model based on the first dataset specifically includes: Step 71: Based on a preset first segmentation ratio, the first dataset is divided into two sub-datasets, denoted as the first training set and the first evaluation set. Wherein, both the first training set and the first evaluation set are composed of multiple first data records; the ratio of the total number of records in the first training set to the total number of records in the first evaluation set satisfies the first segmentation ratio; Step 72: Take each of the first data records in the first training set as the corresponding current training record; and input the image I1 and the image I2 of the current training record into the binocular disparity prediction model for processing to obtain the corresponding disparity map D. 1-2 ; and by the current disparity map D 1-2 and the label disparity map D of the current training record tag Form a corresponding first predicted label pair; Wherein, the disparity map D of each of the first predicted label pairs 1-2 The lateral parallax Δx, the longitudinal parallax Δy, and the confidence level c are denoted as the corresponding... , , The label disparity map D tag The lateral parallax Δx, the longitudinal parallax Δy, and the confidence level c are denoted as the corresponding... , , ; 1 ≤ index r ≤ N tr N tr The total number of records in the first dataset; Step 73: Substitute all the obtained first prediction-label pairs into the preset model loss function L. M The corresponding first loss value is obtained through calculation; Wherein, the model loss function L M From the lateral disparity loss function L △x Longitudinal disparity loss function L △y Credibility loss function L c The composition is as follows: ; ; ; ; ; ; ; , ; λ1, λ2, and λ3 are three preset balance coefficients; L smoothL1 () represents the smooth L1 loss function; w r β1 and β2 are the preset first and second threshold parameters, respectively; µ △y σ △y The mean and standard deviation of longitudinal parallax; Step 74: Identify whether the first loss value meets the preset first loss value range; if it does, proceed to step 75; if not, based on the preset first model optimizer, move towards making the model loss function L... M The direction that reaches the minimum value modulates the model parameters of the binocular disparity prediction model in one round, and returns to step 72 when the modulation ends; The first model optimizer includes the Adam optimizer and the SGD optimizer. Step 75: Take each of the first data records in the first evaluation set as the corresponding current evaluation record; and input the image I1 and the image I2 of the current evaluation record into the binocular disparity prediction model for processing to obtain the corresponding disparity map D. 1-2 ; and by the current disparity map D 1-2 and the label disparity map D of the current evaluation record tagA corresponding second prediction label pair is formed; and based on all the obtained second prediction labels, the RMSE error values ​​of the horizontal and vertical disparities are calculated to obtain the corresponding first and second error values; Step 76: Identify the first and second error values; if the first error value does not meet the preset first error value range or the second error value does not meet the preset second error value range, return to step 71; if the first error value meets the first error value range and the second error value meets the second error value range, stop training and confirm that the training of the binocular disparity prediction model has ended.

[0011] Preferably, the method based on the image I1, the image I2, and the disparity map D... 1-2 The 3D modeling of the train involved in this incident includes: Step 81: Take the point on the current track plane where the optical center of the first camera of the two installed cameras is projected onto the current track plane as the origin, take the forward direction of the train passing through this time as the positive X-axis, take the upward direction perpendicular to the track plane as the positive Y-axis, and take the straight line on the track plane that starts from the origin, is perpendicular to the track line on the camera side and points to the other side of the track as the positive Z-axis, thereby constructing the corresponding right-handed coordinate system. Step 82, the disparity map D 1-2 The pixels whose confidence level c is greater than the preset second confidence level threshold. Record these as valid points; Step 83: Take image I1 as the current image; perform a traversal of all valid points; during this traversal, take the currently traversed valid point as the current point; take the pixel in the current image corresponding to the current point as the current corresponding point; take the three primary color features of the current corresponding point as the corresponding point color feature; and calculate the corresponding depth based on the lateral disparity Δx of the current point. Based on the coordinates (x, y) of the current point, the depth Z, and the preset vertical reference coordinates y... ref Camera image acquisition frequency f cam and the image coordinates (x) of the image center point cen ,y cen ), for the current point in the right-handed coordinate system, the three-dimensional coordinates (x, y) of the point. rh ,y rh ,z rh ) to perform calculations, , , ; and based on the confidence level c corresponding to the current point, the three primary color features, and the three-dimensional coordinates (x, y) of the point. rh ,y rh ,z rh() form a corresponding first point; and at the end of this round of traversal, all the obtained first points form a first point cloud; Step 84: Take image I2 as the current image; perform a traversal of all valid points; during this traversal, take the currently traversed valid point as the current point; take the pixel in the current image corresponding to the current point as the current corresponding point; take the three primary color features of the current corresponding point as the corresponding point color feature; and calculate the corresponding depth based on the lateral disparity Δx of the current point. Based on the coordinates (x, y) of the current point, the depth Z, and the preset vertical reference coordinates y... ref Camera image acquisition frequency f cam and the image coordinates (x) of the image center point cen ,y cen ), for the current point in the right-handed coordinate system, the three-dimensional coordinates (x, y) of the point. rh ,y rh ,z rh ) to perform calculations, , , ; and based on the confidence level c corresponding to the current point, the three primary color features, and the three-dimensional coordinates (x, y) of the point. rh ,y rh ,z rh A corresponding second point is formed by combining all the second points obtained in this round of traversal; and at the end of this round of traversal, a cloud of second points is formed by combining all the second points obtained. Step 85, set the cloud coordinates z of the first and second points. rh The absolute value exceeds the preset depth coordinate threshold z hold Delete the first and second points; and obtain the six maximum and minimum boundary values ​​x of the X, Y, and Z axes from the latest cloud of the first and second points. max x min y max y min z max z min Based on the obtained six maximum and minimum boundary values ​​and the preset three-dimensional voxel mesh size, a corresponding voxel mesh network V is constructed in the right-hand coordinate system. The voxel grid network V is composed of multiple voxel grids v; Step 86: Take each voxel grid v containing at least one first or second point as the current grid; and assign the three-dimensional coordinates (x, y, y) of all points in the current grid to the voxel grid network V. rh ,y rh ,z rh A new three-dimensional coordinate (x, y) of the point is obtained by averaging. rh,y rh ,z rh The process involves calculating the mean of the three primary color features for all points in the current grid to obtain a new three primary color feature, and taking the maximum value within the confidence level c for all points in the current grid as a new confidence level c. This new confidence level c, the three primary color features, and the three-dimensional coordinates (x, y) of the point are then used to determine the new confidence level c. rh ,y rh ,z rh A corresponding third point is formed by combining all the first and second points in the current grid, and all the first and second points in the current grid are deleted; and a corresponding third point cloud is formed by combining all the obtained third points. Step 87: In the voxel grid network V, the third point cloud is completed by interpolation, and it is required that each voxel grid V can only have a maximum of one third point; and outlier points are identified and deleted from the completed third point cloud. Step 88: Based on the preset point cloud 3D modeling tool, perform 3D surface modeling according to the third point cloud and use the obtained 3D model as the 3D model of the train passing through this time. The point cloud 3D modeling tools include at least Open3D tools and PCL tools.

[0012] The second aspect of the present invention provides a system for implementing the method for train 3D modeling based on binocular disparity prediction model provided in the first aspect above. The system includes: a rule customization module, a model building module, a data acquisition module, a model training module, a binocular camera installation module, and a disparity prediction and 3D modeling module. The rule customization module is used to set the installation rules, image acquisition rules, and image stitching rules of the dual parallel line array camera to obtain the corresponding first installation rules, first acquisition rules, and first stitching rules; The model building module is used to set up a corresponding binocular disparity prediction model for the stitched images of dual parallel linear array cameras; the binocular disparity prediction model is used to perform pixel-level binocular disparity prediction based on the input images I1 and I2 and output the corresponding disparity map D. 1-2 ; The data acquisition module obtains the corresponding first dataset by training data acquisition based on the first installation rule, the first acquisition rule and the first splicing rule; The model training module trains the binocular disparity prediction model based on the first dataset; The binocular camera mounting module is used to take any train track as the ground motion trajectory of the current scanning object after the model training is completed, and take any train passing on the current train track as the current scanning object; and to install dual parallel linear array cameras on the side of the current train track based on the first mounting rule. The disparity prediction and 3D modeling module is used to acquire and stitch images based on the first acquisition rule and the first stitching rule when a train passes over two cameras on the current train track to obtain a current set of images I1 and I2; and input the current images I1 and I2 into the binocular disparity prediction model to predict the corresponding disparity map D. 1-2 Based on the image I1, the image I2, and the disparity map D... 1-2 A 3D model of the train that passed through this time was created.

[0013] This invention provides a method and system for 3D modeling of trains based on a binocular disparity prediction model. It employs dual parallel linear array cameras for scanning, utilizing the line scanning characteristics of the cameras to achieve continuous, line-by-line acquisition of data from a high-speed moving train. Furthermore, it utilizes a first acquisition rule to achieve synchronous acquisition throughout the train's passage, fundamentally solving the frame rate limitation problem of traditional solutions and ensuring stable data acquisition in high-speed scenarios. A customized binocular disparity prediction model is used for pixel-level disparity estimation. A spatiotemporal transformation network is used to learn the global correspondence of the train's surface structure, and multi-scale feature fusion is combined with a regression prediction network. It overcomes the limitations of a single viewpoint and improves the completeness of 3D reconstruction of edges and occluded areas. The binocular parallax prediction model achieves robust feature matching for low-texture areas through spatiotemporal joint coding layers and correlation coding layers. The reliability of the model output can be used as a quality assessment index. With a dedicated line laser / line LED light source illumination scheme, it can effectively suppress ambient light interference and ensure measurement accuracy under various lighting conditions. In the 3D modeling stage, the binocular parallax depth calculation method is used to reconstruct 3D coordinates and the longitudinal parallax component is used to compensate for positional offset information caused by vibration, thereby improving robustness in vibration environments. Attached Figure Description

[0014] Figure 1 This is a schematic diagram of a method for 3D modeling of trains based on a binocular parallax prediction model provided in Embodiment 1 of the present invention; Figure 2 This is a schematic diagram of the modules of the binocular disparity prediction model provided in Embodiment 1 of the present invention; Figure 3 This is a schematic diagram of the structure of a system for three-dimensional modeling of trains based on a binocular parallax prediction model, provided in Embodiment 2 of the present invention. Detailed Implementation

[0015] Embodiments of the present invention are described in detail below, examples of which are illustrated in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and intended to explain the present invention, and should not be construed as limiting the present invention.

[0016] Figure 1 This is a schematic diagram of a method for 3D modeling of trains based on a binocular parallax prediction model provided in Embodiment 1 of the present invention, as shown below. Figure 1 As shown, this method includes the following steps: Step 1: Set the installation rules, image acquisition rules, and image stitching rules for the dual parallel line array cameras to obtain the corresponding first installation rules, first acquisition rules, and first stitching rules.

[0017] Here, the first installation rule of this embodiment of the invention is as follows: two identical line scan cameras are selected to form a dual parallel line scan camera; the two cameras are installed on the same side of the ground motion trajectory of the scanned object; the horizontal distance between the center points of the two cameras and the motion trajectory of the scanned object is equal; the vertical height between the center points of the two cameras and the ground is a preset installation height H; the horizontal distance between the center points of the two cameras is fixed as a preset camera spacing L; the vertical direction of the ground motion trajectory of the scanned object is the line scan direction of the two cameras; each of the two cameras is equipped with a corresponding line laser light source or line LED light source; the image acquisition operation of each camera is synchronized with the illumination operation of its corresponding light source; the image acquisition operations of the two cameras are synchronized.

[0018] The first acquisition rule of this embodiment of the invention is as follows: when the scanned object moves toward the two cameras, the camera closer to the scanned object is recorded as the first camera and the other camera is recorded as the second camera; when the distance between the scanned object and the first camera is lower than a preset first distance threshold, the image acquisition operation of the two cameras is started simultaneously, and the illumination operation of the two corresponding light sources is started simultaneously; when the shortest distance between the scanned object and the second camera is not lower than the first distance threshold, the image acquisition operation of the two cameras is stopped simultaneously, and the illumination operation of the two corresponding light sources is stopped simultaneously; and the two image sequences acquired by the first and second cameras during the current passage time of the scanned object are recorded as the corresponding image sequences P1 and P2.

[0019] Among them, image sequences P1 and P2 each correspond to T images p 1,t p 2,t Composition, 1 ≤ index t ≤ T, where T is the total number of samples; Figure p 1,i p 2,iThe image format is the line scan format of the line scan camera. The image shape is H0×W0×D0, where H0, W0, and D0 are the corresponding line scan height, line scan width, and line scan pixel feature dimension, respectively. D0=3. The pixel features of the line scan are composed of the corresponding RGB three primary color features.

[0020] The first stitching rule of this embodiment of the invention is: according to the image stitching principle of a line scan camera, stitch the T images p of image sequence P1 along the image height direction. 1,t The corresponding image I1 is obtained by stitching the images line by line. Then, the T images p of the image sequence P2 are stitched together along the image height direction. 2,t The corresponding image I2 is obtained by stitching the images line by line.

[0021] Here, in this embodiment of the invention, the image shapes of images I1 and I2 are both H1×W1×D1, where H1, W1, and D1 are the corresponding first height, first width, and first pixel feature dimension, respectively, and H1=T×H0, W1=W0, and D1=D0; images I1 and I2 are each composed of H1×W1 corresponding pixels. , Composition; 1 ≤ horizontal coordinate x ≤ W1, 1 ≤ vertical coordinate y ≤ H1; each pixel , The pixel features are all composed of a corresponding set of three primary color features ( , , ), ( , , )composition.

[0022] Step 2: Set the corresponding binocular parallax prediction model for the stitched image of the dual parallel linear array cameras.

[0023] Here, the binocular disparity prediction model of this invention is used to perform pixel-level binocular disparity prediction based on the input images I1 and I2 and output the corresponding disparity map D. 1-2 ,like Figure 2 The diagram shows a module schematic of the binocular disparity prediction model provided in Embodiment 1 of the present invention.

[0024] Disparity map D of this invention embodiment 1-2 The image shape is H out ×W out ×D out H out W out D out These represent the corresponding disparity map height, disparity map width, and disparity map pixel feature dimension, H. out =H1, W out =W1,D out =3, from the corresponding Hout ×W out 1 pixel Composition; pixels It consists of a set of corresponding horizontal disparity Δx, vertical disparity Δy, and confidence level c; for any pixel on image I1 In other words, it corresponds to the pixels on image I2. They correspond to the same real point.

[0025] like Figure 2 As shown, the input of the binocular disparity prediction model is used to receive images I1 and I2, and the output is used to output the corresponding disparity map D. 1-2 The binocular disparity prediction model consists of a spatiotemporal transformation network and a regression prediction network connected sequentially. The spatiotemporal transformation network consists of a feature extractor, a spatiotemporal coding embedding layer, a spatiotemporal joint coding layer, and a spatiotemporal feature decoding layer connected sequentially. The regression prediction network consists of a correlation coding layer, a multi-scale feature fusion layer, a global modeling layer, and a disparity prediction layer connected sequentially.

[0026] like Figure 2 As shown, the spatiotemporal transformation network is used to extract features from images I1 and I2 to obtain corresponding feature vectors H1 and H2 respectively; then, feature fusion is performed on feature vectors H1 and H2, and spatiotemporal encoding is performed on the fused features and sequence transformation is performed on the embedded features to obtain the corresponding sequence S1; then, feature encoding is performed on sequence S1 to obtain the corresponding sequence S2; then, feature vector decomposition is performed on sequence S2 and the two decomposed vectors are upsampled to obtain the corresponding feature vectors X1 and X2, which are sent to the regression prediction network.

[0027] The functions of each submodule of the spatiotemporal transformation network are shown below.

[0028] 1) Feature Extractor: The feature extractor in this embodiment of the invention is implemented based on a CNN network or a residual network. The feature extractor is used to extract features from images I1 and I2 respectively to obtain corresponding feature vectors H1 and H2, which are then sent to the spatiotemporal coding embedding layer.

[0029] Here, in this embodiment of the invention, the vector shapes of feature vectors H1 and H2 are both H2×W2×D2, where H2, W2, and D2 are the corresponding second height, second width, and second pixel feature dimensions, respectively, and H2 < H1, W2 < W1, and D2 > D0.

[0030] 2) Spatiotemporal coding embedding layer: The spatiotemporal coding embedding layer of this invention is used to perform feature fusion on feature vectors H1 and H2 according to the feature channel concatenation method to obtain a feature vector H3 with shape H2×W2×2D2; and based on the feature data h of feature vector H3...i,j,k The height coordinates i, the second height H2, and the total number of data collected T are used to confirm the index t corresponding to the current feature data, and based on each feature data h i,j,k The corresponding camera number is set according to the correspondence between feature vector H1 or feature vector H2; and based on each feature data h i,j,k The corresponding index t, position coordinates (i,j,k), and camera number are used to initialize the corresponding time embedding code, position embedding code, and camera number embedding code. The corresponding code pe is obtained by adding the three types of embedding codes. i,j,k ; and composed of each feature data h i,j,k and its corresponding encoding pe i,j,k Generate corresponding feature data α is a preset scaling factor; and the resulting H2×W2×2D2 feature data... The corresponding embedded feature vector H4 is formed; and H2×W2 sub-vectors with a length of 2D2 are extracted from the embedded feature vector H4, and the resulting H2×W2 sub-vectors form the corresponding sequence S1, which is sent to the spatiotemporal joint coding layer.

[0031] Here, in this embodiment of the invention, the vector shapes of both the feature vector H3 and the embedded feature vector H4 are H2×W2×2D2, 1≤indexi≤H2, 1≤indexj≤W2, 1≤indexk≤2D2; the feature vector H3 consists of H2×W2×2D2 feature data h i,j,k The embedded feature vector H4 consists of two feature data points: H2×W2×2D. Composition; Sequence S1 consists of H2×W2 sub-vectors of length 2D2. , 1≤indexq≤H2×W2.

[0032] 3) Spatiotemporal joint coding layer: The spatiotemporal joint coding layer in this embodiment of the invention is implemented based on the encoder of the Transformer model. The spatiotemporal joint coding layer is used to perform multi-head self-attention coding on sequence S1 and send the resulting feature vector sequence as the corresponding sequence S2 to the spatiotemporal feature decoding layer.

[0033] Here, sequence S2 in this embodiment of the invention includes H2×W2 sub-vectors with a length of 2D2. .

[0034] 4) Spatiotemporal feature decoding layer: In this embodiment of the invention, the spatiotemporal feature decoding layer is used to convert sequence S2 into a feature vector H5 with shape H2×W2×2D2; and extract the H2×W2×D2 feature data corresponding to feature vectors H1 and H2 respectively from feature vector H5 to form two decomposition vectors H6 and H7 with shape H2×W2×D2; and upsample the decomposition vectors H6 and H7 respectively to obtain the corresponding feature vectors X1 and X2, which are then sent to the correlation coding layer.

[0035] Here, in this embodiment of the invention, the feature vectors X1 and X2 both have a shape of H3×W3×D3; H3, W3, and D3 are the corresponding third height, third width, and third pixel feature dimension, respectively, with H3=H1, W3=W1, and D3=D2; each of the feature vectors X1 and X2 has H3×W3 corresponding sub-vectors of length D3. , composition.

[0036] like Figure 2 As shown, the regression prediction network is used to identify the local relevance of feature vector X2 with feature vector X1 as the query, and obtain the corresponding relevance vector C. 1-2 And the eigenvectors X1, X2 and the correlation vector C 1-2 The corresponding feature vector X is obtained by fusion. C ; and for the eigenvector X C Multi-scale feature extraction is performed, and feature fusion is applied to obtain the corresponding feature vector H. mul ; and based on the eigenvector H mul Global feature modeling is performed to obtain the corresponding feature vector H. glb ; and based on the eigenvector H glb Disparity map prediction is performed to obtain the corresponding disparity map D. 1-2 And output it.

[0037] The functions of each sub-module of the regression prediction network are shown below.

[0038] 1) Correlation coding layer: The correlation coding layer in this embodiment of the invention is used to encode each subvector of feature vector X1. h as the current query vector que ; and the current query vector h que The corresponding coordinates x and y are denoted as x. que y que Based on two preset local incremental parameters Δx * , △y * Set the local alignment range in feature vector X2 to the corresponding [x que -△x * ,x que +△x* ]、[y que -△y * ,y que +△y * ]; and for the current query vector h que Each subvector of the local comparison range The cosine similarity is calculated, and the resulting 2(Δx) is used to calculate the similarity. * +△y * +1) similarity values ​​form the corresponding relevance vector c. x,y ; and from the obtained H3×W3 correlation vectors c x,y Form the corresponding correlation vector C 1-2 And according to the feature channel concatenation method, the feature vector X1, feature vector X2, and correlation vector C are concatenated. 1-2 Feature fusion is performed to obtain the corresponding feature vector X. C Send to the multi-scale feature fusion layer.

[0039] Here, the two local incremental parameters Δx in this embodiment of the invention * , △y * All are preset positive integers; correlation vector C 1-2 The vector shape is H4×W4×D4, where H4, W4, and D4 are the corresponding fourth height, fourth width, and fourth pixel feature dimension, respectively. H4=H3, W4=W3, and D4=2(Δx). * +△y * +1); Eigenvector X C The vector shape is H5×W5×D5, where H5, W5, and D5 are the corresponding fifth height, fifth width, and fifth pixel feature dimension, respectively. H5=H3, W5=W3, and D5=D3+D3+2(△x) * +△y * +1).

[0040] 2) Multi-scale feature fusion layer: The multi-scale feature fusion layer in this embodiment of the invention is implemented based on the lightweight U-Net model. The multi-scale feature fusion layer is used to apply the downsampling network of the U-Net model to the feature vector X. C Multi-scale feature extraction is performed, and the upsampling network of the U-Net model is used to fuse all the obtained multi-scale features to obtain the corresponding feature vector H. mul Send to the global modeling layer.

[0041] Here, the feature vector H in this embodiment of the invention mul The vector shape is H6×W6×D6, where H6, W6, and D6 are the corresponding sixth height, sixth width, and sixth pixel feature dimension, respectively. H6=H5, W6=W5, and D6≤D5.

[0042] 3) Global Modeling Layer: The global modeling layer in this embodiment of the invention is used for the feature vector H mul Global pooling yields a pooling vector H of shape D6×1. pool ; and pooling vector H pool The mapping is a weight matrix A of shape H6×W6. pool ; and use weight matrix A pool For the eigenvector H mul The corresponding feature vector H is obtained by weighting. glb Send to the disparity prediction layer.

[0043] Here, the weight matrix A of this embodiment of the invention pool The mapping method is as follows: ; Among them, W g1 W g1 B represents the first and second weight matrices corresponding to the global modeling layer. g1 B g2 The first and second bias vectors corresponding to the global modeling layer; the first weight matrix W g1 The shape is H6×D6, and the first bias vector is B. g1 The shape is H6×1, and the vector... The shape is H6×1; ReLU() is the ReLU activation function; the second weight matrix W g2 The shape is W6×1, and the second bias vector is B. g2 The shape is H6×W6; the weight matrix A pool Its shape is H6×W6.

[0044] Feature vector H in this embodiment of the invention glb The weighting method is as follows: ; Where ⊙ represents the Hadamard product, and Sigmoid() is the Sigmoid activation function. 4) Parallax prediction layer: The disparity prediction layer in this embodiment of the invention is used to employ three built-in independent convolutional layers based on the feature vector H. glb Predicting the three features of the disparity map—lateral disparity, vertical disparity, and confidence—results in three corresponding shapes, H. out ×W out The predicted vectors are ×1, and the three predicted vectors are used to form the corresponding disparity map D. 1-2 And output it.

[0045] Step 3: Based on the first installation rule, the first collection rule, and the first splicing rule, training data is collected to obtain the corresponding first dataset; The first dataset includes multiple first data records; each first data record consists of a set of images I1, I2 and their corresponding label disparity maps D. tag Composition; Label disparity map D tag The data format should be consistent with that of disparity map D1; Specifically, this includes: Step 31, selecting one or more train track segments to form a track set; taking each train track segment of the track set as the current track; installing dual parallel linear array cameras on the current track side based on the first installation rule; taking any train passing on the current track as the current scanning object; and each time a train passes the two installed cameras on the current track, performing an image acquisition and stitching based on the first acquisition rule and the first stitching rule to obtain a set of corresponding images I1 and I2 to form a corresponding first image group; and using all the obtained first image groups to form the corresponding first image set. Step 32, and take each of the first image groups in the first image set as the current image group; and use the preset 3D modeling tools and optical flow analysis tools, according to the disparity map D 1-2 The data format constructs a disparity map based on images I1 and I2 of the current image group; and performs binarization on the confidence of each pixel in the current disparity map based on a preset first confidence threshold. If the confidence is less than the first confidence threshold, it is reset to 0; if the confidence is greater than or equal to the first confidence threshold, it is reset to 1; and the current disparity map with completed confidence binarization is used as the corresponding label disparity map D. tag ; and composed of the current image group and its corresponding label disparity map D tag A corresponding first data record is formed; and all the obtained first data records form a corresponding first dataset. Here, the first confidence threshold of this embodiment of the invention is a pre-set threshold parameter; the 3D modeling tools include at least OpenCV, COLMAP, and Meshroom; the optical flow analysis tools include at least DaVis, Halcon, VisionPro, RAFT model, and FlowNet2 model.

[0046] Step 4: Train a binocular disparity prediction model based on the first dataset; Specifically, it includes: Step 41, dividing the first dataset into two sub-datasets based on a preset first segmentation ratio, denoted as the corresponding first training set and first evaluation set; Here, the first segmentation ratio in this embodiment of the invention is a preset ratio parameter, such as 8:2; both the first training set and the first evaluation set consist of multiple first data records; the ratio of the total number of records in the first training set to the total number of records in the first evaluation set satisfies the first segmentation ratio; Step 42: Take each first data record of the first training set as the corresponding current training record; and input the image I1 and image I2 of the current training record into the binocular disparity prediction model for processing to obtain the corresponding disparity map D. 1-2 ; and by the current disparity map D 1-2 The label disparity map D of the current training record tag Form a corresponding first predicted label pair; Here, the disparity map D of each first predicted label pair in this embodiment of the invention. 1-2 The lateral parallax Δx, longitudinal parallax Δy, and confidence level c are denoted as the corresponding values. , , Disparity map D (label) tag The lateral parallax Δx, longitudinal parallax Δy, and confidence level c are denoted as the corresponding values. , , ; 1 ≤ index r ≤ N tr N tr This represents the total number of records in the first dataset; Step 43: Substitute all the obtained first prediction-label pairs into the preset model loss function L. M The corresponding first loss value is obtained through calculation; Here, the model loss function L in this embodiment of the invention M From the lateral disparity loss function L △x Longitudinal disparity loss function L △y Credibility loss function L c The composition is as follows: ; ; ; ; ; ; ; , ; Where λ1, λ2, and λ3 are three preset balance coefficients; L smoothL1 () represents the smooth L1 loss function; w rβ1 and β2 are the preset first and second threshold parameters, respectively; µ △y σ △y The mean and standard deviation of longitudinal parallax; Step 44: Identify whether the first loss value meets the preset first loss value range; if it does, proceed to step 45; if not, based on the preset first model optimizer, move towards making the model loss function L... M The direction that reaches the minimum value modulates the model parameters of the binocular disparity prediction model in one round, and returns to step 42 when the modulation in this round ends; Here, the first loss value range in this embodiment of the invention is a pre-set numerical range; the first model optimizer includes the Adam optimizer and the SGD optimizer; Step 45: Take each first data record of the first evaluation set as the corresponding current evaluation record; and input the image I1 and image I2 of the current evaluation record into the binocular disparity prediction model for processing to obtain the corresponding disparity map D. 1-2 ; and by the current disparity map D 1-2 The label disparity map D of the current assessment record tag A corresponding second prediction label pair is formed; and based on all the obtained second prediction labels, the RMSE error values ​​of the horizontal and vertical disparities are calculated to obtain the corresponding first and second error values. Step 46: Identify the first and second error values; if the first error value does not meet the preset first error value range or the second error value does not meet the preset second error value range, return to step 41; if the first error value meets the first error value range and the second error value meets the second error value range, stop training and confirm that the training of the binocular disparity prediction model has ended.

[0047] Here, the first and second error value ranges in this embodiment of the invention are two preset numerical ranges.

[0048] Step 5: After the model training is completed, any train track is taken as the ground motion trajectory of the current scanning object, and any train traveling on the current train track is taken as the current scanning object; and a dual parallel linear array camera is installed on the side of the current train track based on the first installation rule.

[0049] Step 6: When a train passes two cameras on the current train track, images are acquired and stitched together based on the first acquisition rule and the first stitching rule to obtain a set of images I1 and I2; then, the current images I1 and I2 are input into the binocular disparity prediction model to predict the corresponding disparity map D. 1-2 Based on image I1, image I2, and disparity map D... 1-2 A 3D model of the train that passed through this time was created; Specifically, this includes: Step 61, when a train passes by two cameras on the current train track, images are acquired and stitched together based on the first acquisition rule and the first stitching rule to obtain a set of images I1 and I2; and the current images I1 and I2 are input into a binocular disparity prediction model to predict the corresponding disparity map D. 1-2 ; Step 62, based on image I1, image I2 and disparity map D 1-2 A 3D model of the train that passed through this time was created; Specifically, it includes: Step 621, taking the point on the current track plane where the optical center of the first camera of the two installed cameras is projected onto the current track plane as the origin, taking the forward direction of the train passing through this time as the positive X-axis, taking the upward direction perpendicular to the track plane as the positive Y-axis, and taking the straight line direction on the track plane that starts from the origin, is perpendicular to the track line on the camera side and points to the other side of the track as the positive Z-axis, thereby constructing the corresponding right-handed coordinate system; Step 622, disparity map D 1-2 Pixels with a confidence level c greater than the preset second confidence level threshold Record these as valid points; Here, the second confidence threshold in this embodiment of the invention is a pre-set numerical range; Step 623: Take image I1 as the current image; perform one round of traversal on all valid points; during this round of traversal, take the currently traversed valid point as the current point; take the pixel in the current image corresponding to the current point as the current corresponding point; take the three primary color features of the current corresponding point as the corresponding point color feature; and calculate the corresponding depth based on the lateral disparity Δx of the current point. Based on the current point's coordinates (x, y), depth Z, and preset vertical reference coordinates y... ref Camera image acquisition frequency f cam and the image coordinates (x) of the image center point cen ,y cen ), for the current point in the right-handed coordinate system, the three-dimensional coordinates (x, y, z) of the point. rh ,y rh ,z rh ) to perform calculations, , , ; and based on the current point's confidence level c, the three primary color features, and the point's three-dimensional coordinates (x, y, c),... rh ,y rh ,z rh () form a corresponding first point; and at the end of this round of traversal, all the obtained first points form a first point cloud; Step 624: Take image I2 as the current image; perform one round of traversal on all valid points; during this round of traversal, take the currently traversed valid point as the current point; take the pixel in the current image corresponding to the current point as the current corresponding point; take the three primary color features of the current corresponding point as the corresponding point color feature; and calculate the corresponding depth based on the lateral disparity Δx of the current point. Based on the current point's coordinates (x, y), depth Z, and preset vertical reference coordinates y... ref Camera image acquisition frequency f cam and the image coordinates (x) of the image center point cen ,y cen ), for the current point in the right-handed coordinate system, the three-dimensional coordinates (x, y, z) of the point. rh ,y rh ,z rh ) to perform calculations, , , ; and based on the current point's confidence level c, the three primary color features, and the point's three-dimensional coordinates (x, y, c),... rh ,y rh ,z rh A corresponding second point is formed by combining all the obtained second points; and at the end of this round of traversal, a cloud of second points is formed by combining all the obtained second points. Step 625, set the cloud coordinates z of the first and second points. rh The absolute value exceeds the preset depth coordinate threshold z hold Delete the first and second points; and obtain the six maximum and minimum boundary values ​​x of the X, Y, and Z axes from the latest cloud of the first and second points. max x min y max y min z max z min Based on the obtained six maximum and minimum boundary values ​​and the preset three-dimensional voxel mesh size, a corresponding voxel mesh network V is constructed in the right-hand coordinate system. Here, the voxel grid network V in this embodiment of the invention is composed of multiple voxel grids v; Step 626: Take each voxel mesh v containing at least one first or second point in the voxel mesh network V as the current mesh; and assign the three-dimensional coordinates (x, y, z) of all points in the current mesh to the voxel mesh network V. rh ,y rh ,z rh A new three-dimensional coordinate (x, y) of a point is obtained by averaging. rh ,y rh ,z rhThe process involves averaging the primary color features of all points in the current grid to obtain a new primary color feature. The maximum value within the confidence level *c* of all points in the current grid is then used as the new confidence level *c*. Finally, the new confidence level *c*, the primary color feature, and the three-dimensional coordinates (x, y, y) of the points are combined to form a new confidence level *c*. rh ,y rh ,z rh A corresponding third point is formed by combining all the first and second points in the current grid, and all the first and second points in the current grid are deleted; and a corresponding third point cloud is formed by combining all the resulting third points. Step 627: In the voxel grid network V, the third point cloud is completed by interpolation, and it is required that at most one third point can be set in each voxel grid V; and outlier points are identified and all outliers are deleted after completion of the third point cloud.

[0050] Step 628: Based on the preset point cloud 3D modeling tool, perform 3D surface modeling according to the third point cloud and use the obtained 3D model as the 3D model of the train passing through this time.

[0051] Here, the point cloud 3D modeling tool in this embodiment of the invention includes at least the Open3D tool and the PCL tool.

[0052] The system for implementing the method described in Embodiment 1 above has the following structure: Figure 3 The schematic diagram of a system for three-dimensional modeling of trains based on a binocular disparity prediction model provided in Embodiment 2 of the present invention is shown. The system includes: a rule customization module 201, a model construction module 202, a data acquisition module 203, a model training module 204, a binocular camera installation module 205, and a disparity prediction and three-dimensional modeling module 206.

[0053] The rule customization module 201 is used to set the installation rules, image acquisition rules, and image stitching rules of the dual parallel line array camera to obtain the corresponding first installation rule, first acquisition rule, and first stitching rule.

[0054] Model building module 202 is used to set up a corresponding binocular disparity prediction model for the stitched images of dual parallel linear array cameras; the binocular disparity prediction model is used to perform pixel-level binocular disparity prediction based on the input images I1 and I2 and output the corresponding disparity map D. 1-2 .

[0055] The data acquisition module 203 acquires training data based on the first installation rule, the first acquisition rule, and the first splicing rule to obtain the corresponding first dataset.

[0056] Model training module 204 trains a binocular disparity prediction model based on the first dataset.

[0057] The binocular camera mounting module 205 is used to take any train track as the ground motion trajectory of the current scanning object after the model training is completed, and take any train passing on the current train track as the current scanning object; and to install dual parallel linear array cameras on the side of the current train track based on the first mounting rule.

[0058] The disparity prediction and 3D modeling module 206 is used to acquire and stitch images I1 and I2 based on the first acquisition rule and the first stitching rule when a train passes over two cameras on the current train track; and inputs the current images I1 and I2 into the binocular disparity prediction model to predict the corresponding disparity map D. 1-2 Based on image I1, image I2, and disparity map D... 1-2 A 3D model of the train that passed through this time was created.

[0059] The second embodiment of the present invention provides a system for three-dimensional modeling of trains based on a binocular parallax prediction model, which can execute the method steps in the above method embodiment. Its implementation principle and technical effect are similar, and will not be repeated here.

[0060] In summary, the technical solution of the method and system for 3D modeling of trains based on a binocular parallax prediction model provided in this embodiment of the invention has at least the following technical effects or advantages: 1) It uses dual parallel linear array cameras for scanning, realizing continuous line-by-line acquisition of high-speed moving trains; 2) It realizes synchronous acquisition of the entire train passage, fundamentally solving the frame rate limitation problem of traditional solutions and ensuring stable data acquisition in high-speed scenarios; 3) It customizes a binocular parallax prediction model for pixel-level parallax estimation, and overcomes the limitations of a single viewpoint based on this model, improving the 3D reconstruction integrity of edges and occluded areas; 4) The binocular parallax prediction model achieves robust feature matching for low-texture areas through a spatiotemporal joint coding layer and a correlation coding layer. The reliability of the model output can be used as a quality assessment index. Combined with a dedicated line laser / line LED light source illumination scheme, it can effectively suppress ambient light interference and ensure measurement accuracy under various lighting conditions; 5) In the 3D modeling stage, it performs 3D coordinate reconstruction based on a binocular parallax depth calculation method and compensates for positional offset information caused by vibration through longitudinal parallax components, improving robustness in vibration environments.

[0061] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein can be implemented in hardware, a software module executed by a processor, or a combination of both. The software module can be located in random access memory (RAM), main memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium known in the art.

[0062] The specific embodiments described above further illustrate the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above description is only a specific embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A method for 3D modeling of trains based on a binocular parallax prediction model, characterized in that, The method includes: The installation rules, image acquisition rules, and image stitching rules of the dual parallel linear array cameras are set to obtain the corresponding first installation rules, first acquisition rules, and first stitching rules; A corresponding binocular disparity prediction model is set up for the stitched images of dual parallel linear array cameras; the binocular disparity prediction model is used to perform pixel-level binocular disparity prediction based on the input images I1 and I2 and output the corresponding disparity map D. 1-2 ; The first dataset is obtained by collecting training data based on the first installation rule, the first collection rule, and the first splicing rule. The binocular disparity prediction model is trained based on the first dataset; After the model training is completed, any train track is taken as the ground motion trajectory of the current scanning object, and any train passing on the current train track is taken as the current scanning object; and a dual parallel linear array camera is installed on the side of the current train track based on the first installation rule. When a train passes over two cameras on the current train track, images are acquired and stitched together based on the first acquisition rule and the first stitching rule to obtain a current set of images I1 and I2; and the current images I1 and I2 are input into the binocular disparity prediction model to predict the corresponding disparity map D. 1-2 Based on the image I1, the image I2, and the disparity map D... 1-2 A 3D model of the train that passed through this time was created.

2. The method for train 3D modeling based on a binocular parallax prediction model according to claim 1, characterized in that, The first installation rule is as follows: Select two identical line scan cameras to form a dual parallel line scan camera system; the two cameras are installed on the same side of the ground motion trajectory of the scanned object; the horizontal distance between the center points of the two cameras and the motion trajectory of the scanned object is equal; the vertical height between the center points of the two cameras and the ground is a preset installation height H; the horizontal distance between the center points of the two cameras is fixed as a preset camera spacing L; the vertical direction of the ground motion trajectory of the scanned object is the line scan direction of the two cameras; each of the two cameras is equipped with a corresponding line laser light source or line LED light source; the image acquisition operation of each camera is synchronized with the illumination operation of its corresponding light source; the image acquisition operations of the two cameras are synchronized. The first acquisition rule is as follows: when the scanned object moves toward the two cameras, the camera closer to the scanned object is recorded as the first camera and the other camera is recorded as the second camera; and when the distance between the scanned object and the first camera is lower than a preset first distance threshold, the image acquisition operation of the two cameras is started simultaneously, and the illumination operation of the two corresponding light sources is started simultaneously. When the shortest distance between the scanned object and the second camera is not less than the first distance threshold, the image acquisition operations of both cameras are simultaneously turned off, and the illumination operations of the two corresponding light sources are simultaneously turned off; the two image sequences acquired by the first and second cameras during the current passage time of the scanned object are recorded as the corresponding image sequences P1 and P2; wherein, each of the image sequences P1 and P2 is composed of T corresponding images p 1,t p 2,t Composition, 1 ≤ index t ≤ T, where T is the total number of samples collected; the graph p 1,i p 2,i The image format is the line scan format of the line scan camera. The image shape is H0×W0×D0, where H0, W0, and D0 are the corresponding line scan height, line scan width, and line scan pixel feature dimension, respectively. D0=3. The pixel features of the line scan are composed of the corresponding RGB three primary color features. The first stitching rule is: according to the image stitching principle of the line scan camera, stitch the T images p of the image sequence P1 along the image height direction. 1,t The corresponding image I1 is obtained by stitching the images line by line, and the T images p of the image sequence P2 are stitched together along the image height direction. 2,t Image I2 is obtained by stitching each row together. Both images I1 and I2 have a shape of H1×W1×D1, where H1, W1, and D1 are the first height, first width, and first pixel feature dimension, respectively. H1=T×H0, W1=W0, and D1=D0. Images I1 and I2 are each composed of H1×W1 pixels. , Composition; 1 ≤ horizontal coordinate x ≤ W1, 1 ≤ vertical coordinate y ≤ H1; each pixel , The pixel features are all composed of a corresponding set of three primary color features ( , , ), ( , , )composition; The disparity map D 1-2 The image shape is H out ×W out ×D out H out W out D out These represent the corresponding disparity map height, disparity map width, and disparity map pixel feature dimension, H. out =H1, W out =W1,D out =3, from the corresponding H out ×W out 1 pixel Composition; the pixel It consists of a set of corresponding horizontal disparity Δx, vertical disparity Δy, and confidence level c; for any pixel point on the image I1 In other words, it is related to the pixels on the image I2. Corresponding to the same real point; The first dataset includes multiple first data records; each first data record consists of a set of images I1, I2 and their corresponding label disparity maps D. tag Composition; the label disparity map D tag The data format is consistent with that of the disparity map D1.

3. The method for train 3D modeling based on a binocular parallax prediction model according to claim 1, characterized in that, The input terminal of the binocular disparity prediction model is used to receive the image I1 and the image I2, and the output terminal is used to output the corresponding disparity map D. 1-2 ; The binocular disparity prediction model is composed of a spatiotemporal transformation network and a regression prediction network connected sequentially. The spatiotemporal transformation network is composed of a feature extractor, a spatiotemporal coding embedding layer, a spatiotemporal joint coding layer, and a spatiotemporal feature decoding layer connected sequentially. The regression prediction network is composed of a correlation coding layer, a multi-scale feature fusion layer, a global modeling layer, and a disparity prediction layer connected sequentially. The spatiotemporal transformation network is used to extract features from images I1 and I2 to obtain corresponding feature vectors H1 and H2, respectively; to fuse the feature vectors H1 and H2, to embed the fused features using spatiotemporal encoding, and to perform sequence transformation on the embedded features to obtain the corresponding sequence S1; to encode the sequence S1 to obtain the corresponding sequence S2; and to decompose the sequence S2 into feature vectors and upsample the two decomposed vectors to obtain the corresponding feature vectors X1 and X2, which are then sent to the regression prediction network. The regression prediction network is used to identify the local correlation of feature vector X2 with feature vector X1 as the query, and obtain the corresponding correlation vector C. 1-2 And the feature vectors X1 and X2 and the correlation vector C are compared. 1-2 The corresponding feature vector X is obtained by fusion. C ; and for the feature vector X C Multi-scale feature extraction is performed, and feature fusion is applied to obtain the corresponding feature vector H. mul ; and based on the feature vector H mul Global feature modeling is performed to obtain the corresponding feature vector H. glb ; and based on the feature vector H glb Disparity map D is obtained by performing disparity map prediction. 1-2 And output it.

4. The method for train 3D modeling based on a binocular parallax prediction model according to claim 3, characterized in that, The feature extractor is implemented based on a CNN network or a residual network; the feature extractor is used to extract features from the image I1 and the image I2 respectively to obtain corresponding feature vectors H1 and H2, which are then sent to the spatiotemporal coding embedding layer. Wherein, the vector shapes of the feature vectors H1 and H2 are both H2×W2×D2, where H2, W2, and D2 are the corresponding second height, second width, and second pixel feature dimensions, respectively, and H2 < H1, W2 < W1, D2 > D0; The spatiotemporal coding embedding layer is used to perform feature fusion on the feature vectors H1 and H2 according to the feature channel concatenation method to obtain a feature vector H3 with shape H2×W2×2D2; and based on the feature data h of the feature vector H3... i,j,k The height coordinates i, the second height H2, and the total number of data collected T are used to confirm the index t corresponding to the current feature data, and based on each feature data h... i,j,k The corresponding camera number is set according to the correspondence between the feature vector H1 or the feature vector H2; and based on each feature data h i,j,k The corresponding index t, position coordinates (i,j,k), and camera number are used to initialize the corresponding time embedding code, position embedding code, and camera number embedding code. The corresponding code pe is obtained by adding the three types of embedding codes. i,j,k ; and composed of each of the aforementioned feature data h i,j,k and its corresponding encoding pe i,j,k Generate corresponding feature data α is a preset scaling factor; and the obtained H2×W2×2D2 feature data The corresponding embedded feature vector H4 is formed; and H2×W2 sub-vectors with a length of 2D2 are extracted from the embedded feature vector H4, and the H2×W2 sub-vectors are used to form the corresponding sequence S1 and sent to the spatiotemporal joint coding layer. Wherein, the vector shapes of both the feature vector H3 and the embedded feature vector H4 are H2×W2×2D2, 1≤indexi≤H2, 1≤indexj≤W2, 1≤indexk≤2D2; the feature vector H3 consists of H2×W2×2D2 feature data h i,j,k The embedded feature vector H4 is composed of H2×W2×2D2 feature data. Composition; the sequence S1 includes H2×W2 sub-vectors of length 2D2. , 1 ≤ index q ≤ H2 × W2; The spatiotemporal joint coding layer is implemented based on the encoder of the Transformer model; the spatiotemporal joint coding layer is used to perform multi-head self-attention coding on the sequence S1 and send the resulting feature vector sequence as the corresponding sequence S2 to the spatiotemporal feature decoding layer; The sequence S2 includes H2×W2 sub-vectors of length 2D2. ; The spatiotemporal feature decoding layer is used to convert the sequence S2 into a feature vector H5 with shape H2×W2×2D2; and extract the H2×W2×D2 feature data corresponding to the feature vectors H1 and H2 respectively from the feature vector H5 to form two decomposition vectors H6 and H7 with shape H2×W2×D2; and upsample the decomposition vectors H6 and H7 respectively to obtain the corresponding feature vectors X1 and X2, which are then sent to the correlation coding layer; The feature vectors X1 and X2 both have a shape of H3×W3×D3, where H3, W3, and D3 are the corresponding third height, third width, and third pixel feature dimension, respectively, and H3=H1, W3=W1, and D3=D2. Each feature vector X1 and X2 has three corresponding H3×W3 sub-vectors of length D3. , composition.

5. The method for train 3D modeling based on a binocular parallax prediction model according to claim 4, characterized in that, The correlation coding layer is used to encode each of the sub-vectors of feature vector X1. h as the current query vector que ; and the current query vector h que The corresponding coordinates x and y are denoted as x. que y que Based on two preset local incremental parameters Δx * , △y * Set the local alignment range in feature vector X2 to the corresponding [x que -△x * ,x que +△x * ]、[y que -△y * ,y que +△y * ]; and for the current query vector h que Each sub-vector of the local comparison range The cosine similarity is calculated, and the resulting 2(Δx) is used to calculate the similarity. * +△y * +1) similarity values ​​form the corresponding relevance vector c. x,y ; and the resulting H3×W3 correlation vectors c x,y The corresponding correlation vector C is formed. 1-2 And according to the feature channel concatenation method, the feature vector X1, the feature vector X2, and the correlation vector C are concatenated. 1-2 Feature fusion is performed to obtain the corresponding feature vector X. C Send to the multi-scale feature fusion layer; Among them, the two local incremental parameters Δx * , △y * All are preset positive integers; the correlation vector C 1-2 The vector shape is H4×W4×D4, where H4, W4, and D4 are the corresponding fourth height, fourth width, and fourth pixel feature dimension, respectively. H4=H3, W4=W3, and D4=2(Δx). * +△y * +1); the feature vector X C The vector shape is H5×W5×D5, where H5, W5, and D5 are the corresponding fifth height, fifth width, and fifth pixel feature dimension, respectively. H5=H3, W5=W3, and D5=D3+D3+2(△x) * +△y * +1); The multi-scale feature fusion layer is implemented based on the lightweight U-Net model; the multi-scale feature fusion layer is used to apply the downsampling network of the U-Net model to the feature vector X. C Multi-scale feature extraction is performed, and the upsampling network of the U-Net model is used to fuse all the obtained multi-scale features to obtain the corresponding feature vector H. mul Send to the global modeling layer; Wherein, the feature vector H mul The vector shape is H6×W6×D6, where H6, W6, and D6 are the corresponding sixth height, sixth width, and sixth pixel feature dimension, respectively. H6=H5, W6=W5, and D6≤D5. The global modeling layer is used to process the feature vector H. mul Global pooling yields a pooling vector H of shape D6×1. pool ; and the pooling vector H pool The mapping is a weight matrix A of shape H6×W6. pool ; and using the weight matrix A pool For the feature vector H mul The corresponding feature vector H is obtained by weighting. glb Send to the disparity prediction layer; Wherein, the weight matrix A pool The mapping method is as follows: ; W g1 W g1 B represents the first and second weight matrices corresponding to the global modeling layer. g1 B g2 The first and second bias vectors corresponding to the global modeling layer; the first weight matrix W g1 The shape is H6×D6, and the first bias vector is B. g1 The shape is H6×1, and the vector... The shape is H6×1; ReLU() is the ReLU activation function; the second weight matrix W g2 The shape is W6×1, and the second bias vector is B. g2 The shape is H6×W6; the weight matrix A pool Its shape is H6×W6; The feature vector H glb The weighting method is as follows: ; Sigmoid() is the Sigmoid activation function; The disparity prediction layer is used to predict the disparity using three built-in independent convolutional layers based on the feature vector H. glb Predicting the three features of the disparity map—lateral disparity, vertical disparity, and confidence—results in three corresponding shapes, H. out ×W out The three predicted vectors are multiplied by 1, and the resulting disparity map D is composed of these three predicted vectors. 1-2 And output it.

6. The method for train 3D modeling based on a binocular parallax prediction model according to claim 2, characterized in that, The first dataset obtained by training data collection based on the first installation rule, the first acquisition rule, and the first splicing rule specifically includes: Select one or more train track segments to form a track set; and use each train track segment of the track set as the current track; and install dual parallel linear array cameras on the side of the current track based on the first installation rule; and use any train passing on the current track as the current scanning object; and each time a train passes between the two installed cameras on the current track, perform an image acquisition and stitching based on the first acquisition rule and the first stitching rule to obtain a set of corresponding images I1 and I2 to form a corresponding first image group; and use all the obtained first image groups to form a corresponding first image set; And take each of the first image groups in the first image set as the current image group; and use preset 3D modeling tools and optical flow analysis tools, according to the disparity map D 1-2 The data format constructs a corresponding disparity map based on images I1 and I2 of the current image group; and performs binarization on the credibility of each pixel in the current disparity map based on a preset first credibility threshold. If the credibility is less than the first credibility threshold, it is reset to 0; if the credibility is greater than or equal to the first credibility threshold, it is reset to 1; and the current disparity map with completed credibility binarization is used as the corresponding label disparity map D. tag ; and the current image group and its corresponding label disparity map D tag The first data record is formed into a corresponding data record; and the first dataset is formed from all the first data records obtained; wherein the three-dimensional modeling tools include at least OpenCV, COLMAP, and Meshroom; and the optical flow analysis tools include at least DaVis, Halcon, VisionPro, RAFT model, and FlowNet2 model.

7. The method for train 3D modeling based on a binocular parallax prediction model according to claim 2, characterized in that, The step of training the binocular disparity prediction model based on the first dataset specifically includes: Step 71: Based on a preset first segmentation ratio, the first dataset is divided into two sub-datasets, denoted as the first training set and the first evaluation set. Wherein, both the first training set and the first evaluation set are composed of multiple first data records; the ratio of the total number of records in the first training set to the total number of records in the first evaluation set satisfies the first segmentation ratio; Step 72: Take each of the first data records in the first training set as the corresponding current training record; and input the image I1 and the image I2 of the current training record into the binocular disparity prediction model for processing to obtain the corresponding disparity map D. 1-2 ; and by the current disparity map D 1-2 and the label disparity map D of the current training record tag Form a corresponding first predicted label pair; Wherein, the disparity map D of each of the first predicted label pairs 1-2 The lateral parallax Δx, the longitudinal parallax Δy, and the confidence level c are denoted as the corresponding... , , The label disparity map D tag The lateral parallax Δx, the longitudinal parallax Δy, and the confidence level c are denoted as the corresponding... , , ; 1 ≤ index r ≤ N tr N tr The total number of records in the first dataset; Step 73: Substitute all the obtained first prediction-label pairs into the preset model loss function L. M The corresponding first loss value is obtained through calculation; Wherein, the model loss function L M From the lateral disparity loss function L △x Longitudinal disparity loss function L △y Credibility loss function L c The composition is as follows: ; ; ; ; ; ; ; , ; λ1, λ2, and λ3 are three preset balance coefficients; L smoothL1 () represents the smooth L1 loss function; w r β1 and β2 are the preset first and second threshold parameters, respectively; µ △y σ △y The mean and standard deviation of longitudinal parallax; Step 74: Identify whether the first loss value meets the preset first loss value range; if it does, proceed to step 75; if not, based on the preset first model optimizer, move towards making the model loss function L... M The direction that reaches the minimum value modulates the model parameters of the binocular disparity prediction model in one round, and returns to step 72 when the modulation ends; The first model optimizer includes the Adam optimizer and the SGD optimizer. Step 75: Take each of the first data records in the first evaluation set as the corresponding current evaluation record; and input the image I1 and the image I2 of the current evaluation record into the binocular disparity prediction model for processing to obtain the corresponding disparity map D. 1-2 ; and by the current disparity map D 1-2 and the label disparity map D of the current evaluation record tag A corresponding second prediction label pair is formed; and based on all the obtained second prediction labels, the RMSE error values ​​of the horizontal and vertical disparities are calculated to obtain the corresponding first and second error values; Step 76: Identify the first and second error values; if the first error value does not meet the preset first error value range or the second error value does not meet the preset second error value range, return to step 71; if the first error value meets the first error value range and the second error value meets the second error value range, stop training and confirm that the training of the binocular disparity prediction model has ended.

8. The method for train 3D modeling based on a binocular parallax prediction model according to claim 2, characterized in that, The image I1, the image I2, and the disparity map D are used as the basis for the following: 1-2 The 3D modeling of the train involved in this incident includes: Step 81: Take the point on the current track plane where the optical center of the first camera of the two installed cameras is projected onto the current track plane as the origin, take the forward direction of the train passing through this time as the positive X-axis, take the upward direction perpendicular to the track plane as the positive Y-axis, and take the straight line on the track plane that starts from the origin, is perpendicular to the track line on the camera side and points to the other side of the track as the positive Z-axis, thereby constructing the corresponding right-handed coordinate system. Step 82, the disparity map D 1-2 The pixels whose confidence level c is greater than the preset second confidence level threshold. Record these as valid points; Step 83: Take image I1 as the current image; perform a traversal of all valid points; during this traversal, take the currently traversed valid point as the current point; take the pixel in the current image corresponding to the current point as the current corresponding point; take the three primary color features of the current corresponding point as the corresponding point color feature; and calculate the corresponding depth based on the lateral disparity Δx of the current point. Based on the coordinates (x, y) of the current point, the depth Z, and the preset vertical reference coordinates y... ref Camera image acquisition frequency f cam and the image coordinates (x) of the image center point cen ,y cen ), for the current point in the right-handed coordinate system, the three-dimensional coordinates (x, y) of the point. rh ,y rh ,z rh ) to perform calculations, , , ; and based on the confidence level c corresponding to the current point, the three primary color features, and the three-dimensional coordinates (x, y) of the point. rh ,y rh ,z rh () form a corresponding first point; and at the end of this round of traversal, all the obtained first points form a first point cloud; Step 84: Take image I2 as the current image; perform a traversal of all valid points; during this traversal, take the currently traversed valid point as the current point; take the pixel in the current image corresponding to the current point as the current corresponding point; take the three primary color features of the current corresponding point as the corresponding point color feature; and calculate the corresponding depth based on the lateral disparity Δx of the current point. Based on the coordinates (x, y) of the current point, the depth Z, and the preset vertical reference coordinates y... ref Camera image acquisition frequency f cam and the image coordinates (x) of the image center point cen ,y cen ), for the current point in the right-handed coordinate system, the three-dimensional coordinates (x, y) of the point. rh ,y rh ,z rh ) to perform calculations, , , ; and based on the confidence level c corresponding to the current point, the three primary color features, and the three-dimensional coordinates (x, y) of the point. rh ,y rh ,z rh A corresponding second point is formed by combining all the second points obtained in this round of traversal; and at the end of this round of traversal, a cloud of second points is formed by combining all the second points obtained. Step 85, set the cloud coordinates z of the first and second points. rh The absolute value exceeds the preset depth coordinate threshold z. hold Delete the first and second points; and obtain the six maximum and minimum boundary values ​​x of the X, Y, and Z axes from the latest cloud of the first and second points. max x min y max y min z max z min Based on the obtained six maximum and minimum boundary values ​​and the preset three-dimensional voxel mesh size, a corresponding voxel mesh network V is constructed in the right-hand coordinate system. The voxel grid network V is composed of multiple voxel grids v; Step 86: Take each voxel grid v containing at least one first or second point as the current grid; and assign the three-dimensional coordinates (x, y, y) of all points in the current grid to the voxel grid network V. rh ,y rh ,z rh A new three-dimensional coordinate (x, y) of the point is obtained by averaging. rh ,y rh ,z rh The process involves averaging the primary color features of all points in the current grid to obtain a new primary color feature, and taking the maximum value within the confidence level c of all points in the current grid as a new confidence level c. This new confidence level c, the primary color feature, and the three-dimensional coordinates (x, y) of the point are then used to determine the new confidence level c. rh ,y rh ,z rh A corresponding third point is formed by combining all the first and second points in the current grid, and all the first and second points in the current grid are deleted; and a corresponding third point cloud is formed by combining all the obtained third points. Step 87: In the voxel grid network V, the third point cloud is completed by interpolation, and it is required that each voxel grid V can only have a maximum of one third point; and outlier points are identified and deleted from the completed third point cloud. Step 88: Based on the preset point cloud 3D modeling tool, perform 3D surface modeling according to the third point cloud and use the obtained 3D model as the 3D model of the train passing through this time. The point cloud 3D modeling tools include at least Open3D tools and PCL tools.

9. A system for implementing the method for three-dimensional train modeling based on a binocular parallax prediction model as described in any one of claims 1-8, characterized in that, The system includes: a rule customization module, a model building module, a data acquisition module, a model training module, a binocular camera installation module, and a disparity prediction and 3D modeling module; The rule customization module is used to set the installation rules, image acquisition rules, and image stitching rules of the dual parallel line array camera to obtain the corresponding first installation rules, first acquisition rules, and first stitching rules; The model building module is used to set up a corresponding binocular disparity prediction model for the stitched images of dual parallel linear array cameras; the binocular disparity prediction model is used to perform pixel-level binocular disparity prediction based on the input images I1 and I2 and output the corresponding disparity map D. 1-2 ; The data acquisition module obtains the corresponding first dataset by training data acquisition based on the first installation rule, the first acquisition rule and the first splicing rule; The model training module trains the binocular disparity prediction model based on the first dataset; The binocular camera mounting module is used to take any train track as the ground motion trajectory of the current scanning object after the model training is completed, and take any train passing on the current train track as the current scanning object; and to install dual parallel linear array cameras on the side of the current train track based on the first mounting rule. The disparity prediction and 3D modeling module is used to acquire and stitch images based on the first acquisition rule and the first stitching rule when a train passes over two cameras on the current train track to obtain a current set of images I1 and I2; and input the current images I1 and I2 into the binocular disparity prediction model to predict the corresponding disparity map D. 1-2 Based on the image I1, the image I2, and the disparity map D... 1-2 A 3D model of the train that passed through this time was created.