Train speed measurement method and system based on track-side double parallel arrangement linear array camera
Patent Information
- Application Number
- CN202511924087.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-18
- Publication Date
- 2026-09-22
- Estimated Expiration
- 2045-12-18
AI Technical Summary
然而,传统的轨旁测速技术(磁钢测速、雷达测速)在通行车速偏低时出错率会增大,基于视觉目标检测模型的智能测速方案在光照条件不佳、通行车速过快等情况下出错率也会增大
[0014]本发明实施例提供的一种基于轨道侧双平行排列线阵相机的列车测速方法和系统,通过双平行排列线阵相机的同步图像采集机制来避免发生信号丢失、通过语义引导来抑制低纹理区域的噪声干扰、通过匀速运动模型和RANSAC算法来消除低速抖动影响,从而提高了低速状态下的预测准确度;通过线激光光源补充车身表面纹理来解决运动模糊问题,通过RAFT模型的4D相关性金字塔模块来捕捉高速运动下的微小位移变化,从而提高了高速状态下的预测准确度;通过视觉语义分割模型提高了对多类光照环境的适应性,通过线激光光源提升了复杂光照条件下的测速鲁棒性,从而提高了对各类复杂光照环境的适应能力。
Smart Images

Figure CN121883836B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of rail transit technology, and in particular to a train speed measurement method and system based on a dual parallel linear array camera on the track side. Background Technology
[0002] In intelligent operation and maintenance systems for rail transit, train speed is a core parameter supporting multiple trackside visual inspection systems. However, traditional trackside speed measurement technologies (magnetic speed measurement, radar speed measurement) have a higher error rate when the train speed is low, and intelligent speed measurement schemes based on visual target detection models also have a higher error rate under poor lighting conditions and excessively high train speeds. Currently, there is an urgent need for a highly reliable speed measurement method that can adapt to both high and low train speeds and various complex lighting environments. Summary of the Invention
[0003] The purpose of this invention is to address the shortcomings of existing technologies by providing a train speed measurement method and system based on a dual parallel linear array camera on the trackside. Under low-speed train conditions, this invention can capture continuous displacement information and avoid signal loss through the synchronous image acquisition mechanism of the dual parallel linear array camera; it can suppress noise interference in low-texture areas through semantic guidance to ensure the stability of optical flow matching at low speeds; and it can eliminate the effects of low-speed jitter through a uniform motion model and the RANSAC algorithm. Under high-speed train conditions, it can supplement the surface texture of the train body with a line laser light source to solve the motion blur problem; it can capture minute displacement changes under high-speed motion through the 4D correlation pyramid module of the RAFT model to avoid matching loss at a single scale; and under complex lighting environments, it can improve adaptability to various lighting conditions through a visual semantic segmentation model, and provide active illumination through a line laser light source to improve the robustness of speed measurement.
[0004] In view of this, a first aspect of the present invention provides a train speed measurement method based on a track-side dual parallel linear array camera, the method comprising: A semantic guidance model is constructed based on a visual semantic segmentation model; a pre-trained RAFT model is used as a displacement field prediction model; a vehicle speed prediction model is constructed based on the semantic guidance model and the displacement field prediction model; the semantic guidance model is trained based on a preset first dataset, and the displacement field prediction model is fine-tuned based on the first dataset; the semantic guidance model is used to perform semantic segmentation of key parts of the vehicle body on the input images I1 and I2, respectively, and to perform semantic weight fusion on the two segmentation images, and output corresponding semantic guidance maps M1 and M2 based on the fusion result; the displacement field prediction model is used to predict the pixel displacement field between the input semantic enhancement maps X1 and X2 and output the corresponding displacement field feature map F0; the vehicle speed prediction model is used to perform end-to-end speed prediction based on the input images I1 and I2 and output the corresponding predicted vehicle speed V; Based on the installation rules of dual parallel line array cameras, the two line array cameras are installed on the same side of the first train track; Each time a train passes the two line array cameras along the first train track, corresponding image sequences P1 and P2 are obtained based on the image acquisition rules of the dual parallel line array cameras. The images are then stitched together according to the image stitching principle of the line array cameras, using the image sequences P1 and P2 as the corresponding images I1 and I2. Noise reduction and distortion correction are then performed on images I1 and I2 respectively. The processed images I1 and I2 are then input into the vehicle speed prediction model to predict the corresponding predicted vehicle speed V. The predicted vehicle speed V, along with the corresponding images I1 and I2 and the image sequences P1 and P2, constitute the train speed measurement record for this train and are saved.
[0005] Preferably, the installation rules for the dual parallel line array cameras are as follows: the line scanning direction of the two line array cameras is perpendicular to the track line of the first train track; the vertical height of the two camera center points of the two line array cameras from the ground is equal, both being a preset installation height H; the horizontal distance between the two camera center points is a preset camera spacing L; the vertical distance between the two camera center points and the track line of the first train track is equal; each of the two line array cameras is equipped with a corresponding line laser light source, and the angle between the camera axis of each line array camera and its corresponding light source line is the same, both being a preset angle θ; the image acquisition operation of each line array camera is synchronized with the illumination operation of its corresponding line laser light source; the image acquisition operations of the two line array cameras are synchronized, and the image time alignment error of the two line array cameras is lower than a preset time error threshold; The image acquisition rules for the dual parallel linear array cameras are as follows: when a train approaches the two linear array cameras on the first train track, the one closer to the front of the train is designated as the first camera, and the one farther from the front is designated as the second camera; when the distance between the train and the first camera is lower than a preset first distance threshold, the two linear array cameras are simultaneously activated for image acquisition, and the two corresponding line laser light sources are simultaneously activated for illumination; when the shortest distance between the train and the second camera is not lower than the first distance threshold, the image acquisition operations of the two linear array cameras are simultaneously deactivated, and the illumination operations of the two line laser light sources are simultaneously deactivated; and the two image sequences acquired by the first and second cameras during the train's passage are designated as the corresponding image sequences P1 and P2; wherein each of the image sequences P1 and P2 is associated with N linear array images p. 1,i p 2,i Composition, 1 ≤ index i ≤ N, where N is the total number of images; the linear array p 1,i p 2,i The image formats are all line scan format of line scan cameras; the height, width and pixel resolution of the line scan images of the first and second cameras are consistent; The image stitching principle of the line scan camera is as follows: all the line scan images in the image sequence obtained according to the image acquisition rules of the dual parallel line scan camera are horizontally stitched together according to the sorting index of the current sequence to obtain the corresponding stitched image. The first dataset includes multiple first data records; each first data record includes a first training image, a second training image, a first semantic label map, a second semantic label map, and a first displacement field label map; the first and second training images are two images I1 and I2 used for training, which have undergone denoising and distortion correction processing. The two images I1 and I2 are currently obtained by stitching together two corresponding image sequences P1 and P2 according to the image stitching principle of the line scan camera. The two image sequences P1 and P2 are currently acquired by two line scan cameras used for data acquisition according to the image acquisition rules of the dual parallel line scan camera. The two line scan cameras used for data acquisition are installed according to the dual parallel line scan camera installation rules. Installation; the first and second semantic label images are semantic segmentation images of key parts of the vehicle body corresponding to the first and second training images, respectively; the image height, width, and pixel feature dimensions of the first and second training images are consistent; the image height and width of the first and second semantic label images are consistent with the corresponding first and second training images; the pixel feature dimension of the first and second semantic label images is 4, corresponding to four semantic types; the four semantic types include bogie, window, door, and others; the image height and width of the first displacement field label image are consistent with the first and second training images; the feature dimension of the first displacement field label image is 2, composed of lateral displacement Δx and longitudinal displacement Δy.
[0006] Preferably, the input terminal of the semantic guidance model is used to receive the images I1 and I2, and the output terminal is used to output the semantic guidance maps M1 and M2. The image shapes of images I1 and I2 are both H1×W1×D1, where H1 and W1 are the corresponding first image height and first image width, respectively, and D1 is a preset first pixel feature dimension, D1=3; images I1 or I2 are composed of H1×W1 first pixel feature vectors with a length of D1, and each first pixel feature vector is composed of RGB three primary color feature values. The semantic guidance maps M1 and M2 have image shapes of H1×W1×D3, where D3 is a preset third pixel feature dimension and D3=1; The semantic guidance model is composed of the visual semantic segmentation model and the weight fusion layer connected sequentially. The visual semantic segmentation model is implemented based on a CNN network model or a U-Net model; the visual semantic segmentation model is used to perform semantic segmentation of key parts of the vehicle body on images I1 and I2 to obtain corresponding semantic segmentation maps O1 and O2; The semantic segmentation maps O1 and O2 have an image shape of H1×W1×D2, where D2 is a preset second pixel feature dimension, and D2=4. Both semantic segmentation maps O1 and O2 are composed of H1×W1 second pixel feature vectors with a length of D2. Each second pixel feature vector is composed of four types of semantic probabilities. The sum of the four types of semantic probabilities is 1 and corresponds one-to-one with the four types of semantic types. The weighted fusion layer is used to take the semantic segmentation maps O1 and O2 as the current segmentation maps; take each of the second pixel feature vectors of the current segmentation maps as the current vectors; take the semantic type corresponding to the maximum probability of the current vectors as the current type; set the corresponding third pixel features based on the current type; and form the corresponding semantic guidance map M1 or M2 by the H1×W1 third pixel features corresponding to the current segmentation map; and output the obtained semantic guidance maps M1 and M2. Among them, the preset semantic guidance parameters corresponding to the bogie, windows, doors and others are w. 转向架 w 车窗 w 车门 w 其他 , 0≤w 其他 <w 车门 <w 车窗 <w 转向架 ≤1; If the current type is bogie, window, door or other, then the corresponding third pixel feature is set to the preset semantic guidance parameter corresponding to the current type.
[0007] Preferably, the first and second model input terminals of the displacement field prediction model are used to receive the corresponding semantic enhancement maps X1 and X2, respectively, and the model output terminal is used to output the displacement field feature map F0. The semantic enhancement images X1 and X2 have an image shape of H1×W1×D4, where D4 is a preset fourth pixel feature dimension, and D4=D1+D3=4. Both semantic enhancement images X1 and X2 are composed of H1×W1 fourth pixel feature vectors with a length of D4. Each fourth pixel feature vector is composed of the corresponding RGB three primary color feature values and semantic guidance parameters. The displacement field feature map F0 has an image shape of H2×W2×D5, where H2 and W2 are the corresponding second image height and second image width, respectively, H2 The displacement field prediction model has the same model structure as the RAFT model, consisting of a context encoder, a feature encoder, a 4D correlation pyramid module, and an iterative update module. The input of the context encoder is connected to the first model input of the displacement field prediction model, and the output is connected to the first input of the iterative update module; the input of the feature encoder is connected to the first and second model inputs of the displacement field prediction model, and the output is connected to the input of the 4D correlation pyramid module; the output of the 4D correlation pyramid module is connected to the second input of the iterative update module; the output of the iterative update module is connected to the model output of the displacement field prediction model. The context encoder is used to downsample the semantic enhancement map X1 to obtain the corresponding first downsampled map, and to perform feature encoding on the first downsampled map to obtain the corresponding first feature map, which is then sent to the iterative update module. The first downsampled image has an image shape of H2×W2×D4; the first feature image has an image shape of H3×W3×D4. con H3 and W3 are the height and width of the corresponding third image, respectively, H3 < H2, W3 < W2, D con D represents the pixel feature dimension output by the context encoder. con >D4; The feature encoder is used to downsample the semantic enhancement maps X1 and X2 respectively to obtain the corresponding second and third downsampled maps, and to perform feature encoding on the second and third downsampled maps respectively to obtain the corresponding second and third feature maps, which are then sent to the 4D correlation pyramid module. The image shapes of the second and third downsampled images are H2×W2×D4; the image shapes of the second and third feature images are H3×W3×D4. fea D fea D represents the pixel feature dimension output by the feature encoder. fea =D con ; The 4D correlation pyramid module is used to take the pixel feature vectors of the second feature map as the current vector; calculate the vector similarity between the current vector and the pixel feature vectors of the third feature map to obtain the corresponding first similarity; and form a pyramid of shape H3×W3×D using the H3×W3 first similarities corresponding to the current vector. cor The correlation feature map is obtained; and a corresponding 4D correlation tensor, denoted as feature tensor C0, is constructed from the H3×W3 correlation feature maps corresponding to the second feature map; and the feature map C0 is subjected to three-level pooling according to the pyramid pooling method to obtain the corresponding feature tensors C1, C2, and C3; and the obtained four-scale feature tensors C0, C1, C2, and C3 are sent to the iterative update module. The image shapes of the feature tensors C0, C1, C2, and C3 are as follows: , , , ; D cor D represents the pixel feature dimension output by the 4D correlation pyramid module. cor =1; The iterative update module is used to predict the lateral and longitudinal displacements corresponding to each pixel in the first feature map based on the first feature map and the feature tensors C0, C1, C2, and C3, to obtain a low-resolution displacement field feature map F with shape H3×W3×D5. low ; and the displacement field feature map F low Upsampling yields a dense displacement field feature map F with shape H2×W2×D5. high ; and the dense displacement field feature map F high The corresponding displacement field feature map F0 is output.
[0008] Preferably, the model input terminal of the vehicle speed prediction model is used to receive the corresponding images I1 and I2, and the model output terminal is used to output the predicted vehicle speed V; The vehicle speed prediction model includes the semantic guidance model, the semantic enhancement layer, the displacement field prediction model, and the post-processing layer; The input of the semantic guidance model is connected to the model input of the vehicle speed prediction model, and the first and second outputs are connected to the first input of the post-processing layer and the first input of the semantic enhancement layer, respectively. The second input of the semantic enhancement layer is connected to the model input of the vehicle speed prediction model, and its output is connected to the input of the displacement field prediction model. The output of the displacement field prediction model is connected to the second input of the post-processing layer. The output of the post-processing layer is connected to the model output of the vehicle speed prediction model. The semantic guidance model is used to perform semantic segmentation of key parts of the vehicle body on images I1 and I2 respectively, and to perform semantic weight fusion on the two segmentation maps respectively, and output the corresponding semantic guidance maps M1 and M2 based on the fusion result; and to send the semantic guidance map M1 to the post-processing layer; and to send the semantic guidance maps M1 and M2 to the semantic enhancement layer. The semantic enhancement layer is used to perform feature fusion on the images I1 and I2 and the corresponding semantic guidance maps M1 and M2 according to the feature channel splicing method to obtain the corresponding semantic enhancement maps X1 and X2, which are then sent to the displacement field prediction model. The displacement field prediction model is used to predict the pixel displacement field between the semantic enhancement maps X1 and X2 and output the corresponding displacement field feature map F0; and send the displacement field feature map F0 to the post-processing layer; The post-processing layer is used to downsample the semantic guidance map M1 to obtain a downsampled guidance map M with shape H2×W2×D3. down ; and based on the downsampling guidance map M down The confidence level is calculated to obtain a confidence feature map M with shape H2×W2×D3. conf ; and based on the credibility feature map M conf The preset credibility threshold c hold The corresponding reliable point set G is confirmed based on the displacement field feature map F0; and the interior point set G in the reliable point set G that conforms to the uniform translation model is determined based on the RANSAC algorithm. iner Make an estimate; and based on the interior point set G iner Estimate the median of the corresponding lateral displacement Δx * ; and based on the median of the lateral displacement Δx * The predicted vehicle speed V is calculated based on the camera spacing L and the image acquisition frequency f1 of the first camera, and then output as V = L × f1 / Δx.* ; Among them, the downsampling guidance map M down Each of the aforementioned third pixel features is denoted as feature w. u,v , 1 ≤ index u ≤ H2, 1 ≤ index v ≤ W2; the feature vectors of each fifth pixel of the displacement field feature map F0 are denoted as feature f. u,v (△x u,v ,△y u,v ); The credibility feature map M conf Considered as having H2×W2 confidence levels c u,v composition; The credibility c u,v The calculation method is as follows: , , , , ; µ x µ y Let σ be the average value of the lateral displacement Δx and the longitudinal displacement Δy, respectively. x σ y These are the standard deviations of the lateral displacement Δx and the longitudinal displacement Δy, respectively. The credibility feature map M conf Each c in u,v ≥c hold For each trusted point (u,v), the total number of trusted points is denoted as the total number N. G A set of lateral displacements Δx corresponding to each reliable point. u,v Longitudinal displacement Δy u,v Features w u,v Let be the corresponding △x k , △y k w k Form a corresponding credible point g k (△x k ,△y k ,w k ), 1 ≤ index k ≤ total number N G N G The aforementioned credible point g k (△x k ,△y k ,w k The set of trusted points G is composed of these points. The interior point set G iner Consisting of multiple interior points g q (△x q ,△y q,w q ), 1 ≤ index q ≤ total number of interior points N iner .
[0009] Furthermore, the RANSAC algorithm is used to process the interior point set G in the trusted point set G that conforms to the uniform translation model. iner The estimation includes: Step 61: Initialize the iterative calculator A to 0; and set the corresponding lateral error threshold ε. hold Longitudinal error threshold δ hold ; Where, ε hold >δ hold >0; Default setting ε hold =1.5, δ hold =0.3; Step 62, set the corresponding temporary interior point set G A Empty; Step 63: Randomly select three trusted points g from the trusted point set G. k (△x k ,△y k ,w k ) as the corresponding trust point g a (△x a ,△y a ,w a ), g b (△x b ,△y b ,w b ), g c (△x c ,△y c ,w c ); and according to the uniform translation model, the three lateral displacements Δx a , △x b , △x c The median is used as the corresponding current horizontal median △x mid The current vertical median △y mid Set to 0; Step 64, based on the current horizontal median Δx mid The current longitudinal median Δy mid The lateral error threshold ε hold and the longitudinal error threshold δ hold For each of the trusted points g in the trusted point set G k (△x k ,△y k ,w k The corresponding set of transverse errors ε k and longitudinal error δ kCalculations are performed based on the lateral error ε described for each group. k and the longitudinal error δ k Set the corresponding judgment flag r k ; Wherein, the lateral error ε k and the longitudinal error δ k The calculation method is as follows: , ; The judgment flag r k The setup method is as follows: ; Step 65, set each of the aforementioned judgment flags r k The credibility point g is 1 k (△x k ,△y k ,w k ) as a corresponding interior point g l (△x l ,△y l ,w l Add to the currently described temporary interior set G A In the middle; and for the current temporary interior point set G. A The total number N is obtained by counting the total number of interior points. A ; and based on the current temporary interior point set G A Calculate the corresponding point set integral S A ; Where 1 ≤ index l ≤ N A ; The point set integral S A The calculation method is as follows: ; Step 66: Increment the iterative calculator A by 1; and check whether the iterative calculator A after incrementing by 1 is less than the preset maximum number of iterations N. max Perform identification; if yes, return to step 62; if no, proceed to step 67. Step 67, integrate all the obtained point sets S A Let A be the iteratively calculated value corresponding to the maximum integral in the equation. * And record the corresponding temporary interior point set and the total number of interior points as . , And calculate the corresponding interior point ratio. ; and for the aforementioned interior point rate Is it less than the preset interior rate threshold? If yes, then reset the iterative calculator A to 0 and return to step 62; if no, then reset the current temporary interior point set. As the corresponding interior point set G iner Output; Wherein, the interior point rate The calculation method is as follows: .
[0010] Furthermore, the statement based on the interior point set G iner Estimate the median of the corresponding lateral displacement Δx * Specifically, it includes: According to feature w q In ascending order, the interior point set G... iner N iner The aforementioned interior point g q (△x q ,△y q ,w q The index q of the internal point set G is reordered; and the reordered internal point set G is used as the basis for the reordering. iner Calculate the corresponding N iner Each cumulative weight C q ; and based on N iner The cumulative weight C q Confirm the corresponding median weight index q * ; and will be compared with the median weight index q * Corresponding lateral displacement As the median value of the lateral displacement Δx * And output; Among them, each of the cumulative weights C q The calculation method is as follows: For each of the aforementioned indices q, the range of the corresponding index o is 1 ≤ o ≤ q; Nth iner Cumulative weights Let be the corresponding total weight. ; The median weight index q * The confirmation method is as follows: .
[0011] Preferably, training the semantic guidance model based on a preset first dataset specifically includes: Step 81: Based on a preset first segmentation ratio, the first dataset is divided into two sub-datasets, denoted as the first training set and the first evaluation set. Wherein, both the first training set and the first evaluation set are composed of multiple first data records; the ratio of the total number of records in the first training set and the first evaluation set satisfies the first segmentation ratio; Step 82: Take each of the first data records in the first training set as the current training record; and input the first and second training images of the current training record as the corresponding images I1 and I2 into the semantic guidance model for processing, and take the semantic segmentation maps O1 and O2 obtained in this processing as the corresponding first and second prediction maps; and form two corresponding first prediction-label pairs with the current first and second prediction maps and the corresponding first and second semantic label maps in the current training record. Step 83: Substitute all the obtained first prediction-label pairs into the preset first model loss function to calculate the corresponding first loss value; The first model loss function is implemented based on the cross-entropy loss function and the DICE loss function; Step 84: Identify whether the first loss value meets the preset first loss value range; if yes, proceed to step 85; if no, modulate the model parameters of the visual semantic segmentation model of the semantic guidance model in one round based on the preset first model optimizer in the direction of minimizing the first model loss function, and return to step 82 when the modulation ends. The first model optimizer includes the Adam optimizer and the SGD optimizer. Step 85: Take each of the first data records in the first evaluation set as the current evaluation record; and input the first and second training images of the current evaluation record as the corresponding images I1 and I2 into the semantic guidance model for processing, and take the semantic segmentation maps O1 and O2 obtained in this processing as the corresponding third and fourth prediction maps; and form two corresponding second prediction-label pairs with the current third and fourth prediction maps and the corresponding first and second semantic label maps in the current evaluation record. Step 86: Based on all the obtained second prediction-label pairs, perform statistical calculations on the four F1 scores corresponding to the four semantic types to obtain the corresponding bogie classification F1 score, window classification F1 score, door classification F1 score, and other classification F1 score; Step 87: Identify the four F1 scores; if the F1 score of the bogie category does not meet the preset first F1 score range, or the F1 score of the window category does not meet the preset second F1 score range, or the F1 score of the door category does not meet the preset third F1 score range, or the F1 score of the other categories does not meet the preset fourth F1 score range, then return to step 81; if the F1 score of the bogie category meets the first F1 score range, and the F1 score of the window category meets the second F1 score range, and the F1 score of the door category meets the third F1 score range, and the F1 score of the other categories meets the fourth F1 score range, then stop training and confirm that the training of the semantic guidance model has ended.
[0012] Preferably, the step of fine-tuning the displacement field prediction model based on the first dataset specifically includes: Step 91: Based on the preset second segmentation ratio, the first dataset is divided into two sub-datasets, which are denoted as the corresponding second training set and second evaluation set. The second training set and the second evaluation set are both composed of multiple first data records; the ratio of the total number of records in the second training set and the second evaluation set satisfies the second segmentation ratio. Step 92: Take each of the first data records in the second training set as the current training record; and input the first and second training images of the current training record as the corresponding images I1 and I2 into the vehicle speed prediction model for processing, and take the displacement field feature map F0 obtained in this processing as the corresponding fifth prediction map; and form the corresponding third prediction-label pair with the current fifth prediction map and the first displacement field label map of the current training record. Step 93: Substitute all the obtained third prediction-label pairs into the preset second model loss function to calculate the corresponding second loss value; The second model loss function is implemented based on either the L1 loss function or the L2 loss function. Step 94: Identify whether the second loss value meets the preset second loss value range; if yes, proceed to step 95; if no, perform a round of fine-tuning of the model parameters of the displacement field prediction model of the vehicle speed prediction model based on the preset second model optimizer in the direction of minimizing the second model loss function, and return to step 92 when the fine-tuning ends. The second model optimizer includes the Adam optimizer and the SGD optimizer; Step 95: Take each of the first data records in the second evaluation set as the current evaluation record; and input the first and second training images of the current evaluation record as the corresponding images I1 and I2 into the vehicle speed prediction model for processing, and take the displacement field feature map F0 obtained in this processing as the corresponding sixth prediction map; and form the corresponding fourth prediction-label pair with the current sixth prediction map and the first displacement field label map of the current evaluation record. Step 96: Substitute all the obtained fourth prediction-labels into the preset first model evaluation function to calculate the corresponding first evaluation value; The first model evaluation function is implemented based on the RMSE function; Step 97: Identify whether the first evaluation value meets the preset first evaluation value range; if not, return to step 91; if it does, stop training and confirm that the training of the displacement field prediction model has ended.
[0013] A second aspect of the present invention provides a system for implementing the train speed measurement method based on a track-side dual parallel linear array camera provided in the first aspect above. The system includes: a model building and training module, a linear array camera installation module, and a model speed measurement module. The model building and training module constructs a semantic guidance model based on a visual semantic segmentation model; uses a pre-trained RAFT model as a displacement field prediction model; constructs a vehicle speed prediction model based on the semantic guidance model and the displacement field prediction model; trains the semantic guidance model based on a preset first dataset, and fine-tunes the displacement field prediction model based on the first dataset; the semantic guidance model is used to perform semantic segmentation of key parts of the vehicle body on the input images I1 and I2, respectively, and performs semantic weight fusion on the two segmentation images, and outputs corresponding semantic guidance maps M1 and M2 based on the fusion result; the displacement field prediction model is used to predict the pixel displacement field between the input semantic enhancement maps X1 and X2 and output the corresponding displacement field feature map F0; the vehicle speed prediction model is used to perform end-to-end speed prediction based on the input images I1 and I2 and output the corresponding predicted vehicle speed V; The line array camera mounting module is based on the installation rule of dual parallel line array cameras, and the two line array cameras are installed on the same side of the first train track. The model speed measurement module is used to obtain corresponding image sequences P1 and P2 based on the image acquisition rules of the dual parallel line array cameras each time a train passes the two line array cameras via the first train track; and to perform image stitching according to the image stitching principle of the line array cameras based on the image sequences P1 and P2, and use the stitching result as the corresponding images I1 and I2; and to perform denoising and distortion correction processing on the images I1 and I2 respectively; and to input the processed images I1 and I2 into the vehicle speed prediction model to predict the corresponding predicted vehicle speed V; and to form and save the train speed measurement record corresponding to this train by the predicted vehicle speed V, the corresponding images I1 and I2, and the image sequences P1 and P2.
[0014] This invention provides a train speed measurement method and system based on a trackside dual parallel linear array camera. It avoids signal loss through a synchronous image acquisition mechanism using the dual parallel linear array camera, suppresses noise interference in low-texture areas through semantic guidance, and eliminates the impact of low-speed jitter through a uniform motion model and the RANSAC algorithm, thereby improving prediction accuracy at low speeds. It addresses motion blur by supplementing the vehicle surface texture with a line laser light source, and captures minute displacement changes at high speeds using the 4D correlation pyramid module of the RAFT model, thus improving prediction accuracy at high speeds. Furthermore, it enhances adaptability to various lighting environments through a visual semantic segmentation model and improves the robustness of speed measurement under complex lighting conditions through the line laser light source, thereby improving adaptability to various complex lighting environments. Attached Figure Description
[0015] Figure 1 This is a schematic diagram of a train speed measurement method based on a track-side dual parallel linear array camera provided in Embodiment 1 of the present invention; Figure 2 This is a schematic diagram of the semantic guidance model provided in Embodiment 1 of the present invention; Figure 3 This is a schematic diagram of the displacement field prediction model provided in Embodiment 1 of the present invention. Figure 4 This is a schematic diagram of the vehicle speed prediction model provided in Embodiment 1 of the present invention. Figure 5 This is a top and side view schematic diagram of the installation position of the line scan camera provided in Embodiment 1 of the present invention; Figure 6 This is a schematic diagram of the horizontal stitching of an image sequence from a line scan camera provided in Embodiment 1 of the present invention; Figure 7 This is a schematic diagram of a train speed measurement system based on a track-side dual parallel linear array camera, provided in Embodiment 2 of the present invention. Detailed Implementation
[0016] Embodiments of the present invention are described in detail below, examples of which are illustrated in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and intended to explain the present invention, and should not be construed as limiting the present invention.
[0017] Figure 1 This is a schematic diagram of a train speed measurement method based on a track-side dual parallel linear array camera provided in Embodiment 1 of the present invention, as shown below. Figure 1 As shown, this method includes the following steps: Step 1: Construct a semantic guidance model based on the visual semantic segmentation model; use the pre-trained RAFT model as the displacement field prediction model; construct a vehicle speed prediction model based on the semantic guidance model and the displacement field prediction model; train the semantic guidance model based on the preset first dataset, and fine-tune the displacement field prediction model based on the first dataset.
[0018] Specifically, it includes: Step 11: Construct a semantic guidance model based on the visual semantic segmentation model.
[0019] Here, the semantic guidance model of this embodiment of the invention is used to perform semantic segmentation of key parts of the vehicle body on the input images I1 and I2, respectively, and to perform semantic weight fusion on the two segmentation images, and output the corresponding semantic guidance images M1 and M2 based on the fusion result.
[0020] like Figure 2 As shown in the schematic diagram of the semantic guidance model provided in Embodiment 1 of the present invention, the model input end of the semantic guidance model is used to receive images I1 and I2, and the model output end is used to output semantic guidance maps M1 and M2.
[0021] In this context, images I1 and I2 both have a shape of H1×W1×D1, where H1 and W1 are the corresponding first image height and width, respectively, and D1 is a preset first pixel feature dimension, with D1=3. Each image I1 or I2 is composed of H1×W1 vectors of length D1, each first pixel feature vector consisting of RGB three primary color feature values. Semantic guidance maps M1 and M2 have a shape of H1×W1×D3, where D3 is a preset third pixel feature dimension, with D3=1.
[0022] like Figure 2 As shown, the semantic guidance model is composed of a visual semantic segmentation model and a weight fusion layer connected sequentially; the functional descriptions of each model component are as follows.
[0023] 1) The visual semantic segmentation model is implemented based on the CNN network model or the U-Net model; the visual semantic segmentation model is used to perform semantic segmentation of key parts of the vehicle body on images I1 and I2 to obtain the corresponding semantic segmentation maps O1 and O2.
[0024] The semantic segmentation maps O1 and O2 have an image shape of H1×W1×D2, where D2 is the preset second pixel feature dimension, and D2=4. Both semantic segmentation maps O1 and O2 are composed of H1×W1 second pixel feature vectors with a length of D2. Each second pixel feature vector is composed of four types of semantic probabilities, and the sum of the four types of semantic probabilities is 1, which corresponds one-to-one with the four types of semantic types. The four types of semantic types include bogie, window, door, and others.
[0025] 2) The weight fusion layer is used to take the semantic segmentation maps O1 and O2 as the current segmentation maps respectively; take each second pixel feature vector of the current segmentation map as the current vector; take the semantic type corresponding to the highest probability of the current vector as the current type; set the corresponding third pixel feature based on the current type; and form the corresponding semantic guidance map M1 or M2 by the H1×W1 third pixel features corresponding to the current segmentation map; and output the obtained semantic guidance maps M1 and M2.
[0026] Here, in this embodiment of the invention, four preset semantic guidance parameters are pre-set, corresponding to the bogie, windows, doors, and others respectively: w 转向架 w 车窗 w 车门 w 其他 , 0≤w 其他 <w 车门 <w 车窗 <w 转向架 ≤1; When the current type is bogie, window, door or other, the corresponding third pixel feature will be set to the preset semantic guidance parameter corresponding to the current type.
[0027] Step 12, and use the pre-trained RAFT model as the displacement field prediction model.
[0028] Here, the Recurrent All-Pairs Field Transforms for Optical Flow (RAFT) model is an intelligent model for predicting motion displacement fields using the optical flow method. This model comes from the publicly available technical document A, "RAFT: Recurrent All-Pairs Field Transforms for Optical Flow". The RAFT model can be pre-trained using the dataset provided in technical document A.
[0029] The displacement field prediction model in this embodiment of the invention is based on the RAFT model, and its model structure is consistent with the standard model structure in technical document A. This displacement field prediction model is used to predict the pixel displacement field between the semantically enhanced maps X1 and X2 as input to the model and output the corresponding displacement field feature map F0, such as... Figure 3 The diagram shows a module schematic of the displacement field prediction model provided in Embodiment 1 of the present invention.
[0030] like Figure 3 As shown, the first and second model inputs of the displacement field prediction model are used to receive the corresponding semantic enhancement maps X1 and X2, respectively, and the model output is used to output the displacement field feature map F0.
[0031] The semantically enhanced images X1 and X2 have an image shape of H1×W1×D4, where D4 is the preset fourth pixel feature dimension, and D4=D1+D3=4. Both semantically enhanced images X1 and X2 are composed of H1×W1 fourth pixel feature vectors with a length of D4. Each fourth pixel feature vector is composed of the corresponding RGB three primary color feature values and semantic guidance parameters.
[0032] The shape of the displacement field feature map F0 is H2×W2×D5, where H2 and W2 are the corresponding second image height and second image width, respectively, H2
[0033] From technical document A and Figure 3 As can be seen, the model structure of the displacement field prediction model is consistent with that of the RAFT model, consisting of a context encoder, a feature encoder, a 4D correlation pyramid module, and an iterative update module. The connection relationships between the modules are as follows: the input of the context encoder is connected to the first input of the displacement field prediction model, and its output is connected to the first input of the iterative update module; the input of the feature encoder is connected to the first and second inputs of the displacement field prediction model, and its output is connected to the input of the 4D correlation pyramid module; the output of the 4D correlation pyramid module is connected to the second input of the iterative update module; and the output of the iterative update module is connected to the output of the displacement field prediction model.
[0034] The functions of each model component are briefly described below. For more detailed information, please refer to technical document A.
[0035] 1) The context encoder is used to downsample the semantic enhancement map X1 to obtain the corresponding first downsampled map, and to encode the first downsampled map to obtain the corresponding first feature map, which is then sent to the iterative update module.
[0036] The first downsampled image has a shape of H2×W2×D4; the first feature image has a shape of H3×W3×D4. con H3 and W3 are the height and width of the corresponding third image, respectively, H3 < H2, W3 < W2, D con D represents the pixel feature dimension output by the context encoder. con >D4.
[0037] 2) The feature encoder is used to downsample the semantic enhancement maps X1 and X2 respectively to obtain the corresponding second and third downsampled maps, and to perform feature encoding on the second and third downsampled maps respectively to obtain the corresponding second and third feature maps, which are then sent to the 4D correlation pyramid module.
[0038] The second and third downsampled images have an image shape of H2×W2×D4; the second and third feature images have an image shape of H3×W3×D4. fea D fea D represents the pixel feature dimension output by the feature encoder. fea =D con .
[0039] 3) The 4D correlation pyramid module uses the pixel feature vectors of the second feature map as the current vector; calculates the vector similarity between the current vector and the pixel feature vectors of the third feature map to obtain the corresponding first similarity; and forms a pyramid of shape H3×W3×D using the H3×W3 first similarity values corresponding to the current vector. cor The correlation feature map is obtained; and a corresponding 4D correlation tensor, denoted as feature tensor C0, is constructed from the H3×W3 correlation feature maps corresponding to the second feature map; and the feature map C0 is subjected to three-level pooling according to the pyramid pooling method to obtain the corresponding feature tensors C1, C2, and C3; and the obtained four-scale feature tensors C0, C1, C2, and C3 are sent to the iterative update module.
[0040] Here, the image shapes of the feature tensors C0, C1, C2, and C3 in this embodiment of the invention are as follows: , , , ; Among them, D cor D represents the pixel feature dimension output by the 4D correlation pyramid module. cor =1.
[0041] 4) The iterative update module is used to predict the lateral and longitudinal displacements corresponding to each pixel in the first feature map based on the first feature map and the feature tensors C0, C1, C2, and C3, to obtain a low-resolution displacement field feature map F with shape H3×W3×D5. low ; and the displacement field characteristic map F low Upsampling yields a dense displacement field feature map F with shape H2×W2×D5. high ; and the dense displacement field characteristic map F high The corresponding displacement field feature map F0 is output.
[0042] Step 13, and construct a vehicle speed prediction model based on the semantic guidance model and the displacement field prediction model.
[0043] Here, the vehicle speed prediction model in this embodiment of the invention is used to perform end-to-end speed prediction based on the input images I1 and I2 and output the corresponding predicted vehicle speed V, such as... Figure 4 The diagram shows a module schematic of the vehicle speed prediction model provided in Embodiment 1 of the present invention.
[0044] like Figure 4 As shown, the input end of the vehicle speed prediction model is used to receive the corresponding images I1 and I2, and the output end is used to output the predicted vehicle speed V.
[0045] like Figure 4 As shown, the vehicle speed prediction model consists of the following components: a semantic guidance model, a semantic enhancement layer, a displacement field prediction model, and a post-processing layer.
[0046] like Figure 4 As shown, the connection relationships of the model components of the vehicle speed prediction model are as follows: the input end of the semantic guidance model is connected to the model input end of the vehicle speed prediction model; the first and second output ends are connected to the first input end of the post-processing layer and the first input end of the semantic enhancement layer, respectively; the second input end of the semantic enhancement layer is connected to the model input end of the vehicle speed prediction model, and its output end is connected to the input end of the displacement field prediction model; the output end of the displacement field prediction model is connected to the second input end of the post-processing layer; and the output end of the post-processing layer is connected to the model output end of the vehicle speed prediction model.
[0047] The model components of the vehicle speed prediction model are shown below.
[0048] 1) The semantic guidance model is used to perform semantic segmentation of key parts of the vehicle body on images I1 and I2 respectively, and to perform semantic weight fusion on the two segmentation maps respectively, and output corresponding semantic guidance maps M1 and M2 based on the fusion result; and send semantic guidance map M1 to the post-processing layer; and send semantic guidance maps M1 and M2 to the semantic enhancement layer.
[0049] 2) The semantic enhancement layer is used to perform feature fusion on images I1 and I2 and the corresponding semantic guidance maps M1 and M2 according to the feature channel splicing method to obtain the corresponding semantic enhancement maps X1 and X2, which are then sent to the displacement field prediction model.
[0050] 3) The displacement field prediction model is used to predict the pixel displacement field between semantically enhanced images X1 and X2 and output the corresponding displacement field feature map F0; and send the displacement field feature map F0 to the post-processing layer.
[0051] 4) The post-processing layer is used to downsample the semantic guidance map M1 to obtain a downsampled guidance map M with shape H2×W2×D3. down ; and based on the downsampling guidance map M down The confidence level is calculated to obtain a confidence feature map M with shape H2×W2×D3. conf ; and based on the credibility feature map M conf The preset credibility threshold c hold The corresponding reliable point set G is confirmed by the displacement field feature map F0; and the interior point set G in the reliable point set G that conforms to the uniform translation model is then analyzed based on the RANSAC algorithm. iner Estimate; and based on the interior set G iner Estimate the median of the corresponding lateral displacement Δx * Based on the median lateral displacement Δx * The predicted vehicle speed V is calculated and output based on the camera spacing L and the image acquisition frequency f1 of the first camera, where V = L × f1 / △x. * .
[0052] Here, the downsampling guidance diagram M of this embodiment of the invention... down The features of each third pixel are denoted as feature w. u,v , 1 ≤ index u ≤ H2, 1 ≤ index v ≤ W2; the feature vectors of each fifth pixel of the displacement field feature map F0 are denoted as feature f. u,v (△x u,v ,△y u,v ); Credibility threshold c hold This is a pre-set threshold parameter.
[0053] Credibility Feature Map M of this Invention Embodiment conf Considered as having H2×W2 confidence levels c u,v Composition, credibility c u,v The calculation method is as follows: , , , , ; Where, µx µ y Let σ be the average value of the lateral displacement Δx and the longitudinal displacement Δy, respectively. x σ y These are the standard deviations of the lateral displacement Δx and the longitudinal displacement Δy, respectively.
[0054] Credibility Feature Map M of this Invention Embodiment conf Each c in u,v ≥c hold For each trusted point (u,v), the total number of trusted points is denoted as the total number N. G A set of lateral displacements Δx corresponding to each reliable point. u,v Longitudinal displacement Δy u,v Features w u,v Let be the corresponding △x k , △y k w k Form a corresponding credible point g k (△x k ,△y k ,w k ), 1 ≤ index k ≤ total number N G The obtained N G A credible point g k (△x k ,△y k ,w k These form the corresponding set of trust points G.
[0055] Interior point set G in this embodiment of the invention iner Consisting of multiple interior points g q (△x q ,△y q ,w q ), 1 ≤ index q ≤ total number of interior points N iner .
[0056] The post-processing layer uses the RANSAC algorithm to process the interior point set G in the set of trustworthy points G that conforms to the uniform translation model. iner The specific steps for making an estimate include: Step A1: Initialize the iterative calculator A to 0; and set the corresponding lateral error threshold ε. hold Longitudinal error threshold δ hold ; Where, ε hold >δ hold >0; Default setting ε hold =1.5, δ hold =0.3; Step A2, set the corresponding temporary interior point set G A Empty; Step A3: Randomly select three trustworthy points g from the trustworthy point set G. k (△x k ,△y k ,w k ) as the corresponding trust point g a (△x a ,△y a ,w a ), g b (△x b ,△y b ,w b ), g c (△x c ,△y c ,w c ); and according to the uniform translation model, the three lateral displacements Δx a , △x b , △x c The median is used as the corresponding current horizontal median △x mid The current vertical median △y mid Set to 0; Step A4, based on the current horizontal median Δx mid Current vertical median Δy mid Lateral error threshold ε hold and longitudinal error threshold δ hold For each trust point g in the trust point set G k (△x k ,△y k ,w k The corresponding set of transverse errors ε k and longitudinal error δ k Calculations were performed, and the results were based on the transverse error ε of each group. k and longitudinal error δ k Set the corresponding judgment flag r k ; Here, the lateral error ε in the embodiment of the present invention k and longitudinal error δ k The calculation method is as follows: , ; The judgment flag r in the embodiment of the present invention k The setup method is as follows: ; Step A5, set each judgment flag r k The confidence point g is 1 k (△x k ,△y k ,w k ) as a corresponding interior point g l(△x l ,△y l ,w l Add to the current temporary interior set G A In the middle; and for the current temporary interior set G A The total number N is obtained by counting the total number of interior points. A ; and based on the current temporary interior set G A Calculate the corresponding point set integral S A ; Where 1 ≤ index l ≤ N A ; Point set integral S A The calculation method is as follows: ; Step A6: Increment the iterative calculator A by 1; and check whether the incremented iterative calculator A is less than the preset maximum number of iterations N. max Perform identification; if yes, return to step A2; if no, proceed to step A7. Here, the maximum number of iterations N in the embodiments of the present invention max It is a pre-set positive integer; Step A7, integrate the obtained point sets S A Let A be the iteratively calculated value corresponding to the maximum integral in the equation. * And record the corresponding temporary interior point set and the total number of interior points as . , And calculate the corresponding interior point ratio. ; and the internal point rate Is it less than the preset interior rate threshold? If yes, reset the iterative calculator A to 0 and return to step A2; otherwise, set the current temporary interior point set to 0. As the corresponding interior set G iner Output; Here, the interior point rate of the present invention embodiment The calculation method is as follows: ; Interior point rate threshold This is a pre-set threshold parameter, such as 92%.
[0057] The post-processing layer is based on the interior point set G iner Estimate the median of the corresponding lateral displacement Δx * The specific processing steps include: According to feature w q In ascending order, for the interior point set G iner N iner Internal point g q (△x q ,△y q ,wq The index q of the internal point set G is reordered; and the reordered internal point set G is used as the basis for the reordering. iner Calculate the corresponding N iner Each cumulative weight C q ; and based on N iner Each cumulative weight C q Confirm the corresponding median weight index q * ; and will be compared with the median weight index q * Corresponding lateral displacement As the median of the lateral displacement Δx * And output; Here, the cumulative weights C in each embodiment of the present invention q The calculation method is as follows: ; For each index q, the range of the corresponding index o is 1≤o≤q; The Nth embodiment of the present invention iner Cumulative weights Let be the corresponding total weight. The median weight index q in this embodiment of the invention * The confirmation method is as follows: .
[0058] Step 14, and train the semantic guidance model based on the preset first dataset, and fine-tune the displacement field prediction model based on the first dataset.
[0059] Here, the first dataset in this embodiment of the invention includes multiple first data records. Each first data record includes a first training image, a second training image, a first semantic label map, a second semantic label map, and a first displacement field label map.
[0060] The first and second training images are two images I1 and I2 used for training, which have undergone denoising and distortion correction. These two images I1 and I2 are obtained by stitching together two corresponding image sequences P1 and P2 according to the image stitching principle of a line scan camera. The two image sequences P1 and P2 were acquired by two line scan cameras used for data acquisition according to the image acquisition rules of a dual parallel line scan camera setup. The two line scan cameras used for data acquisition are installed according to the dual parallel line scan camera installation rules. The height, width, and pixel feature dimensions of the first and second training images are consistent. It should be noted that the data records in the first dataset cover various weather types in different seasons and multiple time periods throughout the day, meaning that the data records in the first dataset can cover various complex lighting environments.
[0061] The first and second semantic label images are semantic segmentation maps of key parts of the vehicle body corresponding to the first and second training images, respectively. The height and width of the first and second semantic label images are consistent with the corresponding first and second training images. The pixel feature dimension of the first and second semantic label images is 4, corresponding to the four semantic types.
[0062] The height and width of the first displacement field label image are consistent with those of the first and second training images; the feature dimension of the first displacement field label image is 2, consisting of the horizontal displacement Δx and the vertical displacement Δy.
[0063] It should be noted that the installation rule for the dual parallel line scan cameras in this embodiment of the invention is as follows: the line scanning direction of the two line scan cameras is perpendicular to the track line of the first train track, such as... Figure 5 The diagram shows a top view and a side view of the installation position of the line scan cameras provided in Embodiment 1 of the present invention; the vertical height of the two camera center points of the two line scan cameras from the ground is equal to a preset installation height H, the horizontal distance between the two camera center points is a preset camera spacing L, and the vertical distance between the two camera center points and the track line of the first train track is equal. Figure 5 As shown, each of the two line scan cameras is equipped with a corresponding line laser light source. The angle between the camera axis of each line scan camera and its corresponding light source line is the same, which is a preset angle θ. The image acquisition operation of each line scan camera is synchronized with the illumination operation of its corresponding line laser light source. The image acquisition operations of the two line scan cameras are synchronized, and the image time alignment error between the two line scan cameras is lower than a preset time error threshold. It should be noted that the installation height H, camera spacing L, angle θ, and time error threshold can all be set according to the actual application, for example, H=1.5 meters, L=10 meters, θ=30°, and time error threshold=10μs.
[0064] It should be noted that the image acquisition rules of the dual parallel line array cameras in this embodiment of the invention are as follows: when a train approaches the two line array cameras on the first train track, the one closer to the front of the train is designated as the first camera, and the one farther from the front of the train is designated as the second camera; when the distance between the train and the first camera is lower than a preset first distance threshold, the two line array cameras are simultaneously activated to perform image acquisition operations, and the two corresponding line laser light sources are simultaneously activated to perform illumination operations; when the shortest distance between the train and the second camera is not lower than the first distance threshold, the image acquisition operations of the two line array cameras are simultaneously deactivated, and the illumination operations of the two line laser light sources are simultaneously deactivated; and the two image sequences acquired by the first and second cameras during the period when the train passes are designated as the corresponding image sequences P1 and P2.
[0065] Here, in this embodiment of the invention, the image sequences P1 and P2 each correspond to N linear array images p. 1,i p 2,iComposition, 1 ≤ index i ≤ N, where N is the total number of images, such as Figure 6 The diagram shows a horizontal stitching of an image sequence from a line scan camera provided in Embodiment 1 of the present invention; the line scan image p of the embodiment of the present invention 1,i p 2,i The image format is the same as that of a line scan camera; the height, width, and pixel resolution of the line scan images from the first and second cameras are consistent. The first distance threshold in this embodiment is a pre-set threshold parameter.
[0066] It should be noted that the image stitching principle of the line scan camera in this embodiment of the invention is as follows: all line scan images in the image sequence obtained according to the image acquisition rules of the dual parallel line scan camera are horizontally stitched together according to the sorting index of the current sequence to obtain the corresponding stitched image, such as... Figure 6 As shown.
[0067] The current step 14 specifically includes: Step 141: Train a semantic guidance model based on a pre-set first dataset; Specifically, it includes: Step 1411, dividing the first dataset into two sub-datasets based on a preset first segmentation ratio, denoted as the corresponding first training set and first evaluation set; Here, the first segmentation ratio in this embodiment of the invention is a preset ratio parameter, such as 8:2; both the first training set and the first evaluation set consist of multiple first data records; the ratio of the total number of records in the first training set and the first evaluation set satisfies the first segmentation ratio; Step 1412: Take each first data record of the first training set as the current training record; and take the first and second training images of the current training record as the corresponding images I1 and I2 as inputs into the semantic guidance model for processing, and take the semantic segmentation maps O1 and O2 obtained in this processing as the corresponding first and second prediction maps; and take the current first and second prediction maps and the corresponding first and second semantic label maps in the current training record as two corresponding first prediction-label pairs. Step 1413: Substitute all the obtained first prediction-label pairs into the preset first model loss function to calculate the corresponding first loss value; Here, the first model loss function in this embodiment of the invention is implemented based on the cross-entropy loss function and the DICE loss function; Step 1414: Identify whether the first loss value meets the preset first loss value range; if yes, proceed to step 1415; if no, based on the preset first model optimizer, perform a round of modulation on the model parameters of the visual semantic segmentation model of the semantic guidance model in the direction of minimizing the first model loss function, and return to step 1412 when the modulation ends. Here, the first loss value range in this embodiment of the invention is a pre-set numerical range; the first model optimizer includes the Adam optimizer and the SGD optimizer; Step 1415: Take each first data record of the first evaluation set as the current evaluation record; and take the first and second training images of the current evaluation record as the corresponding images I1 and I2 as inputs into the semantic guidance model for processing, and take the semantic segmentation maps O1 and O2 obtained in this processing as the corresponding third and fourth prediction maps; and take the current third and fourth prediction maps and the corresponding first and second semantic label maps in the current evaluation record to form two corresponding second prediction-label pairs. Step 1416: Based on all the obtained second prediction-label pairs, perform statistical calculations on the four F1 scores corresponding to the four semantic types to obtain the corresponding F1 scores for bogie classification, window classification, door classification, and other classifications; Step 1417: Identify the F1 scores for the four categories; if the F1 score for the bogie category does not meet the preset first F1 score range, or the F1 score for the window category does not meet the preset second F1 score range, or the F1 score for the door category does not meet the preset third F1 score range, or the F1 score for other categories does not meet the preset fourth F1 score range, then return to step 1411; if the F1 score for the bogie category meets the first F1 score range, and the F1 score for the window category meets the second F1 score range, and the F1 score for the door category meets the third F1 score range, and the F1 score for other categories meets the fourth F1 score range, then stop training and confirm that the training of the semantic guidance model has ended; Here, the first, second, third, and fourth F1 score ranges in this embodiment of the invention are four pre-set numerical ranges; Step 142: Fine-tune the displacement field prediction model based on the first dataset; Specifically, it includes: Step 1421, dividing the first dataset into two sub-datasets based on a preset second segmentation ratio, denoted as the corresponding second training set and second evaluation set; Here, the second segmentation ratio in this embodiment of the invention is a preset ratio parameter, such as 8:2; both the second training set and the second evaluation set are composed of multiple first data records; the ratio of the total number of records in the second training set and the second evaluation set satisfies the second segmentation ratio; Step 1422: Take each of the first data records in the second training set as the current training record; and take the first and second training images of the current training record as the corresponding images I1 and I2 and input them into the vehicle speed prediction model for processing, and take the displacement field feature map F0 obtained in this processing as the corresponding fifth prediction map; and take the current fifth prediction map and the first displacement field label map of the current training record to form the corresponding third prediction-label pair. Step 1423: Substitute all the obtained third prediction-label pairs into the preset second model loss function to calculate the corresponding second loss value; Here, the second model loss function in this embodiment of the invention is implemented based on the L1 loss function or the L2 loss function; Step 1424: Identify whether the second loss value meets the preset range of the second loss value; if yes, proceed to step 1425; if no, perform a round of fine-tuning of the model parameters of the displacement field prediction model of the vehicle speed prediction model based on the preset second model optimizer in the direction of minimizing the second model loss function, and return to step 1422 when the fine-tuning is completed. Here, the second loss value range in this embodiment of the invention is a pre-set numerical range; the second model optimizer includes the Adam optimizer and the SGD optimizer; Step 1425: Take each of the first data records in the second evaluation set as the current evaluation record; and take the first and second training images of the current evaluation record as the corresponding images I1 and I2 and input them into the vehicle speed prediction model for processing, and take the displacement field feature map F0 obtained in this processing as the corresponding sixth prediction map; and take the current sixth prediction map and the first displacement field label map of the current evaluation record to form the corresponding fourth prediction-label pair. Step 1426: Substitute all the obtained fourth prediction-label pairs into the preset first model evaluation function to calculate the corresponding first evaluation value; Here, the first model evaluation function in this embodiment of the invention is implemented based on the RMSE function; Step 1427: Identify whether the first evaluation value meets the preset first evaluation value range; if not, return to step 1421; if it does, stop training and confirm that the training of the displacement field prediction model has ended.
[0068] Here, the first evaluation value range of this embodiment of the invention is a pre-set numerical range.
[0069] Step 2: Based on the installation rules for dual parallel line array cameras, install the two line array cameras on the same side of the first train track.
[0070] Step 3: Each time a train passes the two line array cameras on the first train track, the corresponding image sequences P1 and P2 are obtained based on the image acquisition rules of the dual parallel line array cameras; the images are then stitched together according to the image stitching principle of the line array cameras based on the image sequences P1 and P2, and the stitched results are used as the corresponding images I1 and I2; the images I1 and I2 are then subjected to denoising and distortion correction processing respectively; the processed images I1 and I2 are then input into the train speed prediction model to predict the corresponding predicted train speed V; and the predicted train speed V, the corresponding images I1 and I2, and the image sequences P1 and P2 constitute the train speed measurement record for this train and are saved.
[0071] The system for implementing the method described in Embodiment 1 above has the following structure: Figure 7 The schematic diagram of a train speed measurement system based on a dual parallel linear array camera on the track side provided in Embodiment 2 of the present invention is shown. The system includes: a model building and training module 201, a linear array camera installation module 202, and a model speed measurement module 203.
[0072] The model building and training module 201 constructs a semantic guidance model based on the visual semantic segmentation model; and uses the pre-trained RAFT model as the displacement field prediction model; and constructs a vehicle speed prediction model based on the semantic guidance model and the displacement field prediction model; and trains the semantic guidance model based on the preset first dataset, and fine-tunes the displacement field prediction model based on the first dataset; the semantic guidance model is used to perform semantic segmentation of key parts of the vehicle body on the input images I1 and I2 respectively, and performs semantic weight fusion on the two segmentation maps respectively, and outputs the corresponding semantic guidance maps M1 and M2 based on the fusion result; the displacement field prediction model is used to predict the pixel displacement field between the semantic enhancement maps X1 and X2 input to the model and outputs the corresponding displacement field feature map F0; the vehicle speed prediction model is used to perform end-to-end speed prediction based on the input images I1 and I2 and outputs the corresponding predicted vehicle speed V.
[0073] The line array camera mounting module 202 is based on the installation rule of dual parallel line array cameras, and installs two line array cameras on the same side of the first train track.
[0074] The model speed measurement module 203 is used to obtain the corresponding image sequences P1 and P2 based on the image acquisition rules of the dual parallel line array cameras each time a train passes through the two line array cameras on the first train track; and to perform image stitching according to the image stitching principle of the line array cameras based on the image sequences P1 and P2, and use the stitching result as the corresponding images I1 and I2; and to perform denoising and distortion correction processing on images I1 and I2 respectively; and to input the processed images I1 and I2 into the vehicle speed prediction model to predict the corresponding predicted vehicle speed V; and to form and save the train speed measurement record corresponding to this train by the predicted vehicle speed V, the corresponding images I1 and I2, and the image sequences P1 and P2.
[0075] The train speed measurement system based on a dual parallel linear array camera on the track side provided in Embodiment 2 of the present invention can execute the method steps in the above method embodiments. Its implementation principle and technical effect are similar, and will not be repeated here.
[0076] In summary, the technical solution of the train speed measurement method and system based on a dual parallel linear array camera on the track side provided in this embodiment of the invention has at least the following technical effects or advantages: 1) By using the synchronous image acquisition mechanism of the dual parallel linear array camera, signal loss is avoided; by using semantic guidance, noise interference in low-texture areas is suppressed; and by using a uniform motion model and the RANSAC algorithm, the influence of low-speed jitter is eliminated, thereby improving the prediction accuracy at low speeds; 2) By using a line laser light source to supplement the surface texture of the vehicle body to solve the motion blur problem, and by using the 4D correlation pyramid module of the RAFT model to capture small displacement changes under high-speed motion, the prediction accuracy at high speeds is improved; 3) By using a visual semantic segmentation model, the adaptability to various lighting environments is improved; and by using a line laser light source, the robustness of speed measurement under complex lighting conditions is enhanced, thereby further improving the adaptability to various complex lighting environments.
[0077] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein can be implemented in hardware, a software module executed by a processor, or a combination of both. The software module can be located in random access memory (RAM), main memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium known in the art.
[0078] The specific embodiments described above further illustrate the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above description is only a specific embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A train speed measurement method based on a track-side dual parallel linear array camera, characterized in that, The method includes: A semantic guidance model is constructed based on a visual semantic segmentation model; a pre-trained RAFT model is used as a displacement field prediction model; a vehicle speed prediction model is constructed based on the semantic guidance model and the displacement field prediction model; the semantic guidance model is trained based on a preset first dataset, and the displacement field prediction model is fine-tuned based on the first dataset; the semantic guidance model is used to perform semantic segmentation of key parts of the vehicle body on the input images I1 and I2, respectively, and to perform semantic weight fusion on the two segmentation images, and output corresponding semantic guidance maps M1 and M2 based on the fusion result; the displacement field prediction model is used to predict the pixel displacement field between the input semantic enhancement maps X1 and X2 and output the corresponding displacement field feature map F0; the vehicle speed prediction model is used to perform end-to-end speed prediction based on the input images I1 and I2 and output the corresponding predicted vehicle speed V; Based on the installation rules of dual parallel line array cameras, the two line array cameras are installed on the same side of the first train track; Each time a train passes the two line array cameras along the first train track, corresponding image sequences P1 and P2 are obtained based on the image acquisition rules of the dual parallel line array cameras. The images are then stitched together according to the image stitching principle of the line array cameras, using the image sequences P1 and P2 as the corresponding images I1 and I2. Denoising and distortion correction are then performed on images I1 and I2 respectively. The processed images I1 and I2 are then input into the vehicle speed prediction model to predict the corresponding predicted vehicle speed V. The predicted vehicle speed V, along with the corresponding images I1 and I2 and the image sequences P1 and P2, constitute the train speed measurement record for this train and are saved. The semantic guidance model has an input terminal for receiving images I1 and I2, and an output terminal for outputting semantic guidance maps M1 and M2. The image shapes of images I1 and I2 are both H1×W1×D1, where H1 and W1 are the corresponding first image height and first image width, respectively, and D1 is a preset first pixel feature dimension, D1=3; images I1 or I2 are composed of H1×W1 first pixel feature vectors with a length of D1, and each first pixel feature vector is composed of RGB three primary color feature values. The semantic guidance maps M1 and M2 have image shapes of H1×W1×D3, where D3 is a preset third pixel feature dimension and D3=1; The semantic guidance model is composed of the visual semantic segmentation model and the weight fusion layer connected sequentially. The visual semantic segmentation model is implemented based on a CNN network model or a U-Net model; the visual semantic segmentation model is used to perform semantic segmentation of key parts of the vehicle body on images I1 and I2 to obtain corresponding semantic segmentation maps O1 and O2; The semantic segmentation maps O1 and O2 have an image shape of H1×W1×D2, where D2 is a preset second pixel feature dimension, and D2=4. Both semantic segmentation maps O1 and O2 are composed of H1×W1 second pixel feature vectors with a length of D2. Each second pixel feature vector is composed of four types of semantic probabilities. The sum of the four types of semantic probabilities is 1 and corresponds one-to-one with the four types of semantic types. The weighted fusion layer is used to take the semantic segmentation maps O1 and O2 as the current segmentation maps; take each of the second pixel feature vectors of the current segmentation maps as the current vectors; take the semantic type corresponding to the maximum probability of the current vectors as the current type; set the corresponding third pixel features based on the current type; and form the corresponding semantic guidance map M1 or M2 by the H1×W1 third pixel features corresponding to the current segmentation map; and output the obtained semantic guidance maps M1 and M2. Among them, the preset semantic guidance parameters corresponding to the bogie, windows, doors and others are w. 转向架 w 车窗 w 车门 w 其他 , 0≤w 其他 <w 车门 <w 车窗 <w 转向架 ≤1; If the current type is bogie, window, door or other, then the corresponding third pixel feature is set to the preset semantic guidance parameter corresponding to the current type.
2. The train speed measurement method based on track-side dual parallel linear array cameras according to claim 1, characterized in that, The installation rules for the dual parallel line array cameras are as follows: the line scanning direction of the two line array cameras is perpendicular to the track line of the first train track; the vertical height of the two camera center points of the two line array cameras from the ground is equal, both being a preset installation height H; the horizontal distance between the two camera center points is a preset camera spacing L; the vertical distance between the two camera center points and the track line of the first train track is equal; each of the two line array cameras is equipped with a corresponding line laser light source, and the angle between the camera axis of each line array camera and its corresponding light source line is the same, both being a preset angle θ; the image acquisition operation of each line array camera is synchronized with the illumination operation of its corresponding line laser light source; the image acquisition operations of the two line array cameras are synchronized, and the image time alignment error of the two line array cameras is lower than a preset time error threshold. The image acquisition rules for the dual parallel linear array cameras are as follows: when a train approaches the two linear array cameras on the first train track, the one closer to the front of the train is designated as the first camera, and the one farther from the front is designated as the second camera; when the distance between the train and the first camera is lower than a preset first distance threshold, the two linear array cameras are simultaneously activated for image acquisition, and the two corresponding line laser light sources are simultaneously activated for illumination; when the shortest distance between the train and the second camera is not lower than the first distance threshold, the image acquisition operations of the two linear array cameras are simultaneously deactivated, and the illumination operations of the two line laser light sources are simultaneously deactivated; and the two image sequences acquired by the first and second cameras during the train's passage are designated as the corresponding image sequences P1 and P2; wherein each of the image sequences P1 and P2 is associated with N linear array images p. 1,i p 2,i Composition, 1 ≤ index i ≤ N, where N is the total number of images; the linear array p 1,i p 2,i The image formats are all line scan format of line scan cameras; the height, width and pixel resolution of the line scan images of the first and second cameras are consistent; The image stitching principle of the line scan camera is as follows: all the line scan images in the image sequence obtained according to the image acquisition rules of the dual parallel line scan camera are horizontally stitched together according to the sorting index of the current sequence to obtain the corresponding stitched image. The first dataset includes multiple first data records; each first data record includes a first training image, a second training image, a first semantic label map, a second semantic label map, and a first displacement field label map; the first and second training images are two images I1 and I2 used for training, which have undergone denoising and distortion correction processing. The two images I1 and I2 are currently obtained by stitching together two corresponding image sequences P1 and P2 according to the image stitching principle of the line scan camera. The two image sequences P1 and P2 are currently acquired by two line scan cameras used for data acquisition according to the image acquisition rules of the dual parallel line scan camera. The two line scan cameras used for data acquisition are installed according to the dual parallel line scan camera installation rules. Installation; the first and second semantic label images are semantic segmentation images of key parts of the vehicle body corresponding to the first and second training images, respectively; the image height, width, and pixel feature dimensions of the first and second training images are consistent; the image height and width of the first and second semantic label images are consistent with the corresponding first and second training images; the pixel feature dimension of the first and second semantic label images is 4, corresponding to four semantic types; the four semantic types include bogie, window, door, and others; the image height and width of the first displacement field label image are consistent with the first and second training images; the feature dimension of the first displacement field label image is 2, composed of lateral displacement Δx and longitudinal displacement Δy.
3. The train speed measurement method based on a track-side dual parallel linear array camera according to claim 1, characterized in that, The first and second model input terminals of the displacement field prediction model are used to receive the corresponding semantic enhancement maps X1 and X2, respectively, and the model output terminal is used to output the displacement field feature map F0. The semantic enhancement images X1 and X2 have an image shape of H1×W1×D4, where D4 is a preset fourth pixel feature dimension, and D4=D1+D3=4. Both semantic enhancement images X1 and X2 are composed of H1×W1 fourth pixel feature vectors with a length of D4. Each fourth pixel feature vector is composed of the corresponding RGB three primary color feature values and semantic guidance parameters. The displacement field feature map F0 has an image shape of H2×W2×D5, where H2 and W2 are the corresponding second image height and second image width, respectively, H2<H1, W2<W1, and D5 is a preset fifth pixel feature dimension, D5=2; the displacement field feature map F0 is composed of H2×W2 fifth pixel feature vectors with a length of D5, and each fifth pixel feature vector is composed of the corresponding horizontal displacement Δx and vertical displacement Δy; The displacement field prediction model has the same model structure as the RAFT model, consisting of a context encoder, a feature encoder, a 4D correlation pyramid module, and an iterative update module. The input of the context encoder is connected to the first model input of the displacement field prediction model, and the output is connected to the first input of the iterative update module; the input of the feature encoder is connected to the first and second model inputs of the displacement field prediction model, and the output is connected to the input of the 4D correlation pyramid module; the output of the 4D correlation pyramid module is connected to the second input of the iterative update module; the output of the iterative update module is connected to the model output of the displacement field prediction model. The context encoder is used to downsample the semantic enhancement map X1 to obtain the corresponding first downsampled map, and to perform feature encoding on the first downsampled map to obtain the corresponding first feature map, which is then sent to the iterative update module. The first downsampled image has an image shape of H2×W2×D4; the first feature image has an image shape of H3×W3×D4. con H3 and W3 are the height and width of the corresponding third image, respectively, H3 < H2, W3 < W2, D con D represents the pixel feature dimension output by the context encoder. con >D4; The feature encoder is used to downsample the semantic enhancement maps X1 and X2 respectively to obtain the corresponding second and third downsampled maps, and to perform feature encoding on the second and third downsampled maps respectively to obtain the corresponding second and third feature maps, which are then sent to the 4D correlation pyramid module. The image shapes of the second and third downsampled images are H2×W2×D4; the image shapes of the second and third feature images are H3×W3×D4. fea D fea D represents the pixel feature dimension output by the feature encoder. fea =D con ; The 4D correlation pyramid module is used to take the pixel feature vectors of the second feature map as the current vector; calculate the vector similarity between the current vector and the pixel feature vectors of the third feature map to obtain the corresponding first similarity; and form a pyramid of shape H3×W3×D using the H3×W3 first similarities corresponding to the current vector. cor The correlation feature map is obtained; and a corresponding 4D correlation tensor, denoted as feature tensor C0, is constructed from the H3×W3 correlation feature maps corresponding to the second feature map; and the feature map C0 is subjected to three-level pooling according to the pyramid pooling method to obtain the corresponding feature tensors C1, C2, and C3; and the obtained four-scale feature tensors C0, C1, C2, and C3 are sent to the iterative update module. The image shapes of the feature tensors C0, C1, C2, and C3 are as follows: 、 、 、 ; D cor D represents the pixel feature dimension output by the 4D correlation pyramid module. cor =1; The iterative update module is used to predict the lateral and longitudinal displacements corresponding to each pixel in the first feature map based on the first feature map and the feature tensors C0, C1, C2, and C3, to obtain a low-resolution displacement field feature map F with shape H3×W3×D5. low ; and the displacement field feature map F low Upsampling yields a dense displacement field feature map F with shape H2×W2×D5. high ; and the dense displacement field feature map F high The corresponding displacement field feature map F0 is output.
4. The train speed measurement method based on a track-side dual parallel linear array camera according to claim 3, characterized in that, The model input terminal of the vehicle speed prediction model is used to receive the corresponding images I1 and I2, and the model output terminal is used to output the predicted vehicle speed V. The vehicle speed prediction model includes the semantic guidance model, the semantic enhancement layer, the displacement field prediction model, and the post-processing layer; The input of the semantic guidance model is connected to the model input of the vehicle speed prediction model, and the first and second outputs are connected to the first input of the post-processing layer and the first input of the semantic enhancement layer, respectively. The second input of the semantic enhancement layer is connected to the model input of the vehicle speed prediction model, and its output is connected to the input of the displacement field prediction model. The output of the displacement field prediction model is connected to the second input of the post-processing layer. The output of the post-processing layer is connected to the model output of the vehicle speed prediction model. The semantic guidance model is used to perform semantic segmentation of key parts of the vehicle body on images I1 and I2 respectively, and to perform semantic weight fusion on the two segmentation maps respectively, and output the corresponding semantic guidance maps M1 and M2 based on the fusion result; and to send the semantic guidance map M1 to the post-processing layer; and to send the semantic guidance maps M1 and M2 to the semantic enhancement layer. The semantic enhancement layer is used to perform feature fusion on the images I1 and I2 and the corresponding semantic guidance maps M1 and M2 according to the feature channel splicing method to obtain the corresponding semantic enhancement maps X1 and X2, which are then sent to the displacement field prediction model. The displacement field prediction model is used to predict the pixel displacement field between the semantic enhancement maps X1 and X2 and output the corresponding displacement field feature map F0; and send the displacement field feature map F0 to the post-processing layer; The post-processing layer is used to downsample the semantic guidance map M1 to obtain a downsampled guidance map M with shape H2×W2×D3. down ; and based on the downsampling guidance map M down The confidence level is calculated to obtain a confidence feature map M with shape H2×W2×D3. conf ; and based on the credibility feature map M conf The preset credibility threshold c hold The corresponding reliable point set G is confirmed based on the displacement field feature map F0; and the interior point set G in the reliable point set G that conforms to the uniform translation model is determined based on the RANSAC algorithm. iner Make an estimate; and based on the interior point set G iner Estimate the median of the corresponding lateral displacement Δx * ; and based on the median of the lateral displacement Δx * The predicted vehicle speed V is calculated based on the camera spacing L and the image acquisition frequency f1 of the first camera, and then output as V = L × f1 / Δx. * ; Among them, the downsampling guidance map M down Each of the aforementioned third pixel features is denoted as feature w. u,v , 1 ≤ index u ≤ H2, 1 ≤ index v ≤ W2; the feature vectors of each fifth pixel of the displacement field feature map F0 are denoted as feature f. u,v (△x u,v ,△y u,v ); The credibility feature map M conf Considered as having H2×W2 confidence levels c u,v composition; The credibility c u,v The calculation method is as follows: , , , , ; µ x µ y Let σ be the average value of the lateral displacement Δx and the longitudinal displacement Δy, respectively. x σ y These are the standard deviations of the lateral displacement Δx and the longitudinal displacement Δy, respectively. The credibility feature map M conf Each c in u,v ≥c hold For each trusted point (u,v), the total number of trusted points is denoted as the total number N. G A set of lateral displacements Δx corresponding to each reliable point. u,v Longitudinal displacement Δy u,v Features w u,v Let be the corresponding △x k , △y k w k Form a corresponding credible point g k (△x k ,△y k ,w k ), 1 ≤ index k ≤ total number N G N G The aforementioned credible point g k (△x k ,△y k ,w k The set of trusted points G is composed of these points. The interior point set G iner Consisting of multiple interior points g q (△x q ,△y q ,w q ), 1 ≤ index q ≤ total number of interior points N iner .
5. The train speed measurement method based on a track-side dual parallel linear array camera according to claim 4, characterized in that, The RANSAC algorithm is used to analyze the interior point set G in the trusted point set G that conforms to the uniform translation model. iner The estimation includes: Step 61: Initialize the iterative calculator A to 0; and set the corresponding lateral error threshold ε. hold Longitudinal error threshold δ hold ; Where, ε hold >δ hold >0; Default setting ε hold =1.5, δ hold =0.3; Step 62, set the corresponding temporary interior point set G A Empty; Step 63: Randomly select three trusted points g from the trusted point set G. k (△x k ,△y k ,w k ) as the corresponding trust point g a (△x a ,△y a ,w a ), g b (△x b ,△y b ,w b ), g c (△x c ,△y c ,w c ); and according to the uniform translation model, the three lateral displacements Δx a , △x b , △x c The median is used as the corresponding current horizontal median △x mid The current vertical median △y mid Set to 0; Step 64, based on the current horizontal median Δx mid The current longitudinal median Δy mid The lateral error threshold ε hold and the longitudinal error threshold δ hold For each of the trusted points g in the trusted point set G k (△x k ,△y k ,w k The corresponding set of transverse errors ε k and longitudinal error δ k Calculations are performed based on the lateral error ε described for each group. k and the longitudinal error δ k Set the corresponding judgment flag r k ; Wherein, the lateral error ε k and the longitudinal error δ k The calculation method is as follows: , ; The judgment flag r k The setup method is as follows: ; Step 65, set each of the aforementioned judgment flags r k The credibility point g is 1 k (△x k ,△y k ,w k ) as a corresponding interior point g l (△x l ,△y l ,w l Add to the currently described temporary interior set G A In; and for the currently described temporary interior point set G A The total number N is obtained by counting the total number of interior points. A ; and based on the current temporary interior point set G A Calculate the corresponding point set integral S A ; Where 1 ≤ index l ≤ N A ; The point set integral S A The calculation method is as follows: ; Step 66: Increment the iterative calculator A by 1; and check whether the iterative calculator A after incrementing by 1 is less than the preset maximum number of iterations N. max Perform identification; if yes, return to step 62; if no, proceed to step 67. Step 67, integrate S of all the obtained point sets. A Let A be the iteratively calculated value corresponding to the maximum integral in the equation. * And record the corresponding temporary interior point set and the total number of interior points as . , And calculate the corresponding interior point ratio. ; and for the aforementioned interior point rate Is it less than the preset interior rate threshold? If yes, then reset the iterative calculator A to 0 and return to step 62; if no, then reset the current temporary interior point set. As the corresponding interior point set G iner Output; Wherein, the interior point rate The calculation method is as follows: .
6. The train speed measurement method based on a track-side dual parallel linear array camera according to claim 4, characterized in that, The basis of the interior point set G iner Estimate the median of the corresponding lateral displacement Δx * Specifically, it includes: According to feature w q In ascending order, the interior point set G... iner N iner The aforementioned interior point g q (△x q ,△y q ,w q The index q of the internal point set G is reordered; and the reordered internal point set G is used as the basis for the reordering. iner Calculate the corresponding N iner Each cumulative weight C q ; and based on N iner The cumulative weight C q Confirm the corresponding median weight index q * ; and will be compared with the median weight index q * Corresponding lateral displacement As the median value of the lateral displacement Δx * And output; Among them, each of the cumulative weights C q The calculation method is as follows: For each of the aforementioned indices q, the range of the corresponding index o is 1 ≤ o ≤ q; Nth iner Cumulative weights Let be the corresponding total weight. ; The median weight index q * The confirmation method is as follows: .
7. The train speed measurement method based on a track-side dual parallel linear array camera according to claim 2, characterized in that, The process of training the semantic guidance model based on a preset first dataset specifically includes: Step 81: Based on a preset first segmentation ratio, the first dataset is divided into two sub-datasets, denoted as the first training set and the first evaluation set. Wherein, both the first training set and the first evaluation set are composed of multiple first data records; the ratio of the total number of records in the first training set and the first evaluation set satisfies the first segmentation ratio; Step 82: Take each of the first data records in the first training set as the current training record; and input the first and second training images of the current training record as the corresponding images I1 and I2 into the semantic guidance model for processing, and take the semantic segmentation maps O1 and O2 obtained in this processing as the corresponding first and second prediction maps; and form two corresponding first prediction-label pairs with the current first and second prediction maps and the corresponding first and second semantic label maps in the current training record. Step 83: Substitute all the obtained first prediction-label pairs into the preset first model loss function to calculate the corresponding first loss value; The first model loss function is implemented based on the cross-entropy loss function and the DICE loss function; Step 84: Identify whether the first loss value meets the preset first loss value range; if yes, proceed to step 85; if no, modulate the model parameters of the visual semantic segmentation model of the semantic guidance model in one round based on the preset first model optimizer in the direction of minimizing the first model loss function, and return to step 82 when the modulation ends. The first model optimizer includes the Adam optimizer and the SGD optimizer. Step 85: Take each of the first data records in the first evaluation set as the current evaluation record; and input the first and second training images of the current evaluation record as the corresponding images I1 and I2 into the semantic guidance model for processing, and take the semantic segmentation maps O1 and O2 obtained in this processing as the corresponding third and fourth prediction maps; and form two corresponding second prediction-label pairs with the current third and fourth prediction maps and the corresponding first and second semantic label maps in the current evaluation record. Step 86: Based on all the obtained second prediction-label pairs, perform statistical calculations on the four F1 scores corresponding to the four semantic types to obtain the corresponding bogie classification F1 score, window classification F1 score, door classification F1 score, and other classification F1 score; Step 87: Identify the four F1 scores; if the F1 score of the bogie category does not meet the preset first F1 score range, or the F1 score of the window category does not meet the preset second F1 score range, or the F1 score of the door category does not meet the preset third F1 score range, or the F1 score of the other categories does not meet the preset fourth F1 score range, then return to step 81; if the F1 score of the bogie category meets the first F1 score range, and the F1 score of the window category meets the second F1 score range, and the F1 score of the door category meets the third F1 score range, and the F1 score of the other categories meets the fourth F1 score range, then stop training and confirm that the training of the semantic guidance model has ended.
8. The train speed measurement method based on a track-side dual parallel linear array camera according to claim 1, characterized in that, The fine-tuning of the displacement field prediction model based on the first dataset specifically includes: Step 91: Based on the preset second segmentation ratio, the first dataset is divided into two sub-datasets, which are denoted as the corresponding second training set and second evaluation set; The second training set and the second evaluation set are both composed of multiple first data records; the ratio of the total number of records in the second training set and the second evaluation set satisfies the second segmentation ratio. Step 92: Take each of the first data records in the second training set as the current training record; and input the first and second training images of the current training record as the corresponding images I1 and I2 into the vehicle speed prediction model for processing, and take the displacement field feature map F0 obtained in this processing as the corresponding fifth prediction map; and form the corresponding third prediction-label pair with the current fifth prediction map and the first displacement field label map of the current training record. Step 93: Substitute all the obtained third prediction-label pairs into the preset second model loss function to calculate the corresponding second loss value; The second model loss function is implemented based on either the L1 loss function or the L2 loss function. Step 94: Identify whether the second loss value meets the preset second loss value range; if yes, proceed to step 95; if no, perform a round of fine-tuning of the model parameters of the displacement field prediction model of the vehicle speed prediction model based on the preset second model optimizer in the direction of minimizing the second model loss function, and return to step 92 when the fine-tuning ends. The second model optimizer includes the Adam optimizer and the SGD optimizer; Step 95: Take each of the first data records in the second evaluation set as the current evaluation record; and input the first and second training images of the current evaluation record as the corresponding images I1 and I2 into the vehicle speed prediction model for processing, and take the displacement field feature map F0 obtained in this processing as the corresponding sixth prediction map; and form the corresponding fourth prediction-label pair with the current sixth prediction map and the first displacement field label map of the current evaluation record. Step 96: Substitute all the obtained fourth prediction-labels into the preset first model evaluation function to calculate the corresponding first evaluation value; The first model evaluation function is implemented based on the RMSE function; Step 97: Identify whether the first evaluation value meets the preset first evaluation value range; if not, return to step 91; if it does, stop training and confirm that the training of the displacement field prediction model has ended.
9. A system for implementing the train speed measurement method based on track-side dual parallel linear array cameras as described in any one of claims 1-8, characterized in that, The system includes: a model building and training module, a line scan camera mounting module, and a model speed measurement module; The model building and training module constructs a semantic guidance model based on a visual semantic segmentation model; uses a pre-trained RAFT model as a displacement field prediction model; constructs a vehicle speed prediction model based on the semantic guidance model and the displacement field prediction model; trains the semantic guidance model based on a preset first dataset, and fine-tunes the displacement field prediction model based on the first dataset; the semantic guidance model is used to perform semantic segmentation of key parts of the vehicle body on the input images I1 and I2, respectively, and performs semantic weight fusion on the two segmentation images, and outputs corresponding semantic guidance maps M1 and M2 based on the fusion result; the displacement field prediction model is used to predict the pixel displacement field between the input semantic enhancement maps X1 and X2 and output the corresponding displacement field feature map F0; the vehicle speed prediction model is used to perform end-to-end speed prediction based on the input images I1 and I2 and output the corresponding predicted vehicle speed V; The line array camera mounting module is based on the installation rule of dual parallel line array cameras, and the two line array cameras are installed on the same side of the first train track. The model speed measurement module is used to obtain corresponding image sequences P1 and P2 based on the image acquisition rules of the dual parallel line array cameras each time a train passes the two line array cameras via the first train track; and to perform image stitching according to the image stitching principle of the line array cameras based on the image sequences P1 and P2, and use the stitching result as the corresponding images I1 and I2; and to perform denoising and distortion correction processing on the images I1 and I2 respectively; and to input the processed images I1 and I2 into the vehicle speed prediction model to predict the corresponding predicted vehicle speed V; and to form and save the train speed measurement record corresponding to this train by the predicted vehicle speed V, the corresponding images I1 and I2, and the image sequences P1 and P2.
Citation Information
Patent Citations
Vehicle bottom transparent registration method and system, medium and program product
CN119169064A
Scene space three-dimensional model dynamic modeling method based on multi-modal data
CN119339008A