Train chassis registration method and system based on deep learning and DFT technology

By combining deep learning and DFT technology, the acceleration feature vector and image stitching are used to generate low-dimensional spatial deformation domains, which solves the problem of high-resolution train chassis image registration speed and accuracy, and achieves efficient image registration.

CN120388055APending Publication Date: 2025-07-29LI CHUANG ZHI HENG ELECTRONICS TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510475158.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-16
Publication Date
2025-07-29

AI Technical Summary

Technical Problem

The existing deep learning-based image registration method has high computational complexity when processing high-resolution train chassis images, resulting in low registration speed and accuracy, making it difficult to meet the real-time requirements of online train detection.

Method used

Using a method based on deep learning and DFT technology, the channel stitching is performed by obtaining acceleration feature vectors and images, and a low-dimensional spatial deformation domain is generated by an encoder, and the spatial deformation field is reconstructed through frequency domain conversion and decoder to achieve deformation correction of the image.

Benefits of technology

It improves the speed and accuracy of image registration, meets the real-time requirements of online train detection, and reduces computing complexity and memory usage.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120388055A_ABST
    Figure CN120388055A_ABST
Patent Text Reader

Abstract

The invention provides a train chassis registration method and system based on deep learning and a DFT technology, and the method comprises the steps: obtaining an acceleration feature vector which is generated through a train acceleration signal, obtaining a to-be-registered image and a template image, carrying out the channel splicing of the to-be-registered image and the template image, and obtaining a to-be-registered image; and inputting the acceleration feature vector and the spliced image into an encoder, performing prediction by the encoder to generate a low-dimensional space deformation domain, performing frequency domain conversion on the low-dimensional space deformation domain to obtain a low-frequency space deformation field, and reconstructing the low-frequency space deformation field through a decoder to generate a space deformation field. And performing deformation correction on the to-be-registered image based on the spatial deformation field, and outputting a target registered image. The DFT technology is introduced into image registration, a low-frequency space deformation field is obtained through a small-scale neural network, then IDFT is carried out to obtain a space deformation domain, the image registration speed can be increased, and the image registration quality can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of deep learning technology, and in particular, to a train chassis registration method and system based on deep learning and DFT technology. Background Art

[0002] In many fields, it is necessary to align images taken at different times, different angles, or by different devices. Image registration technology finds a spatial transformation to match and align two or more images of the same scene obtained at different times, different perspectives, or by different sensors, so that the corresponding feature points in the images can be accurately coincident.

[0003] Image registration methods are based on feature extraction and matching. For example, algorithms such as SIFT and SURF are used to identify key points in images, and then the transformation matrix between the images is calculated through these key points to achieve registration. Deep learning technology has also been applied to image registration, and the transformation matrix between images is automatically learned by means of a convolutional neural network, thereby achieving more accurate registration.

[0004] However, registration methods based on deep learning adopt network architectures such as GAN or ResNet. When processing high-resolution images collected by linear array cameras, they will face the problem of increased computational complexity, resulting in low registration speed and accuracy. Summary of the Invention

[0005] This application provides a train chassis registration method and system based on deep learning and DFT technology to solve the problem of low registration speed and accuracy.

[0006] In a first aspect, this application provides a train chassis registration method based on deep learning and DFT technology, including:

[0007] Obtaining an acceleration feature vector, where the acceleration feature vector is generated from a train acceleration signal;

[0008] Obtaining a to-be-registered image and a template image, and performing channel splicing on the to-be-registered image and the template image to generate a spliced image;

[0009] Inputting the acceleration feature vector and the spliced image into an encoder to predict and generate a low-dimensional space deformation domain through the encoder;

[0010] Performing frequency domain conversion on the low-dimensional space deformation domain to obtain a low-frequency space deformation field;

[0011] Reconstructing the low-frequency space deformation field through a decoder to generate a space deformation field;

[0012] Performing deformation correction on the to-be-registered image based on the space deformation field, and outputting a target registered image.

[0013] In some feasible embodiments, the obtaining of the acceleration feature vector includes:

[0014] Obtain an acceleration signal;

[0015] Convert the acceleration signal into an embedding vector through an embedding layer;

[0016] Input the embedding vector into a multi-layer perceptron layer to generate an embedding vector. The number of multi-layer perceptron layers is a preset number, and the lengths of multiple embedding vectors are different;

[0017] Perform a vector replication operation on the embedding vector to expand the dimension of the embedding vector to a preset dimension to obtain an acceleration feature vector.

[0018] In some feasible embodiments, the encoder includes a neural network layer and a frequency domain feature extraction layer. The neural network layer includes a convolutional layer and an attention module;

[0019] Inputting the acceleration feature vector and the spliced image into the encoder to predict and generate a low-dimensional space deformation domain through the encoder includes:

[0020] Input the spliced image into the convolutional layer to perform downsampling through the convolutional layer and generate a multi-scale feature map;

[0021] Add the acceleration feature vector and the multi-scale feature map vectorially to output a channel-adjusted feature map;

[0022] Extract local features of the channel-adjusted feature map through the attention module to generate a low-dimensional space deformation domain.

[0023] In some feasible embodiments, the extracting of local features of the channel-adjusted feature map to generate a low-dimensional space deformation domain includes:

[0024] Divide the channel-adjusted feature map into local regions, and the sizes of multiple local regions are the same;

[0025] Expand the local regions into feature units, and the feature units are in one-dimensional form;

[0026] Perform self-attention calculation on the feature units to output a feature representation. The self-attention calculation is to pass the feature units through a preset number of linear layers;

[0027] Based on the feature representation, perform cross-attention operation on the feature units to output a fusion result. The cross-attention operation is to perform information fusion of different scales on the feature units;

[0028] Based on the fusion result, output a deformation domain in the low-dimensional space.

[0029] In some feasible embodiments, performing frequency domain conversion on the deformation domain in the low-dimensional space to obtain a deformation field in the low-frequency space includes:

[0030] Performing frequency domain conversion on the deformation domain in the low-dimensional space through a discrete Fourier transform layer to obtain a frequency domain deformation domain, where the frequency domain deformation domain contains low-frequency information;

[0031] Performing frequency domain conversion on the frequency domain deformation domain through a discrete Fourier transform layer to obtain a deformation field in the low-frequency space.

[0032] In some feasible embodiments, the decoder includes a zero-padding layer and an inverse discrete Fourier transform layer;

[0033] Reconstructing the deformation field in the low-frequency space through the decoder to generate a deformation field in the space includes:

[0034] Performing zero-padding on the deformation field in the low-frequency space through the zero-padding layer and generating a restored image, where the size of the restored image is the same as that of the image to be registered;

[0035] Inputting the restored image into the inverse discrete Fourier transform layer to perform inverse discrete Fourier transform through the inverse discrete Fourier transform layer, and restoring the deformation field in the low-frequency space to a deformation field in the space, where the deformation field in the space is a deformation field with full resolution.

[0036] In some feasible embodiments, the deformation field in the space includes a first coordinate axis displacement amount and a second coordinate axis displacement amount;

[0037] Performing deformation correction on the image to be registered based on the deformation field in the space and outputting a target registered image includes:

[0038] Based on the first coordinate axis displacement amount and the second coordinate axis displacement amount, calculate the original image coordinates of the pixel point, where the pixel point is a pixel point in the target registered image;

[0039] If the original image coordinates are non-integer coordinates, obtain adjacent coordinates, where the adjacent coordinates are the pixel values of the four adjacent integer coordinate points around the non-integer coordinates;

[0040] Based on the adjacent coordinates, calculate an interpolation result through bilinear interpolation, where the interpolation result is the pixel value of the non-integer coordinates;

[0041] Use the interpolation result as the pixel value of the corresponding pixel point in the target registered image to generate the target registered image.

[0042] In some feasible embodiments, calculating the interpolation result by the bilinear interpolation method includes:

[0043] Calculating a horizontal direction weight value according to the horizontal direction distance ratio between the non-integer coordinate and the adjacent coordinates; and calculating a vertical direction weight value according to the vertical direction distance ratio between the non-integer coordinate and the adjacent coordinates;

[0044] Performing two linear interpolations on the adjacent coordinates based on the horizontal direction weight value and the vertical direction weight value to obtain an interpolation result.

[0045] In some feasible embodiments, obtaining the image to be registered includes:

[0046] Recording the acquisition time of the acceleration feature vector;

[0047] Generating a preset time period based on the acquisition time;

[0048] Obtaining the image to be registered within the preset time period so that the acquisition time of the image to be registered is synchronized with the acquisition time of the acceleration feature vector.

[0049] In a second aspect, the present application provides a train chassis registration system based on deep learning and DFT technology, including:

[0050] An acquisition unit for acquiring an acceleration feature vector, where the acceleration feature vector is generated from a train acceleration signal; and acquiring an image to be registered and a template image, and performing channel splicing on the image to be registered and the template image to generate a spliced image;

[0051] A processing unit for inputting the acceleration feature vector and the spliced image into an encoder;

[0052] The encoder is used to predict and generate a low-dimensional space deformation domain; and perform a frequency domain conversion on the low-dimensional space deformation domain to obtain a low-frequency space deformation field;

[0053] A decoder for reconstructing the low-frequency space deformation field to generate a space deformation field;

[0054] A reconstruction unit for performing deformation correction on the image to be registered based on the space deformation field and outputting a target registered image.

[0055] As can be seen from the above technical solutions, the present application provides a train chassis registration method and system based on deep learning and DFT technology, including: obtaining an acceleration feature vector, which is generated from the train acceleration signal, then obtaining the image to be registered and the template image, and performing channel splicing on the image to be registered and the template image to generate a spliced image, inputting the acceleration feature vector and the spliced image into an encoder, the encoder predicts and generates a low-dimensional space deformation domain, then performs a frequency domain conversion on the low-dimensional space deformation domain to obtain a low-frequency space deformation field, reconstructs the low-frequency space deformation field through a decoder to generate a space deformation field, and performs deformation correction on the image to be registered based on the space deformation field to output the target registered image. By introducing the DFT technology into image registration, obtaining its low-frequency space deformation field with a small-scale neural network, and then performing an IDFT transformation to obtain the space deformation field, the speed and quality of image registration can be accelerated. BRIEF DESCRIPTION OF THE DRAWINGS

[0056] In order to more clearly illustrate the technical solutions of the present application, the drawings required for use in the embodiments will be briefly introduced below. Obviously, for those of ordinary skill in the art, other drawings can also be obtained based on these drawings without creative efforts.

[0057] Figure 1 It is a schematic flowchart of the train chassis registration method based on deep learning and DFT technology provided by the embodiment of the present application;

[0058] Figure 2 It is a schematic diagram of the train chassis registration method based on deep learning and DFT technology provided by the embodiment of the present application;

[0059] Figure 3 It is a schematic structural diagram of the attention module provided by the embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0060] The embodiments will be described in detail below, and the examples are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following examples do not represent all embodiments consistent with the present application. They are only examples of systems and methods consistent with some aspects of the present application detailed in the claims.

[0061] In the railway transportation industry, as the load-bearing system of a train, the train chassis needs to be regularly evaluated for its safety and reliability. In actual operation, an online real-time detection method can be adopted, and a line array camera is used to collect images of the chassis during the train's operation. Due to the high and unstable speed of the train and the imaging characteristics of the line array camera itself, the speed fluctuations during the high-speed operation of the train and the imaging characteristics of the line array camera will cause low-frequency deformation (such as stretching and distortion) in the collected images, and these images need to be registered to correct the distortion in order to accurately evaluate the state of the chassis.

[0062] To solve the image registration problem, in some embodiments, it can be carried out around feature extraction and matching. For example, the SIFT (Scale-Invariant Feature Transform) algorithm detects key points in the image and calculates the scale-invariant feature descriptions of these key points to identify the same key points in different images. Furthermore, a transformation matrix between the images is calculated based on these matched key points to achieve image registration.

[0063] The SURF (Speeded-Up Robust Features) algorithm is an improvement based on the SIFT algorithm. It uses integral images to accelerate the detection and description of key points, improving the speed of feature extraction, and also completes image registration based on key point matching.

[0064] In some embodiments, it is also possible to be based on deep learning technology. Through the feature learning ability of a convolutional neural network, the transformation matrix between images is automatically learned. Compared with the above methods, more accurate registration can be achieved.

[0065] For example, a Siamese network contains two or more sub-network branches with shared weights. In image registration, the two images to be registered are respectively input into the sub-networks. The network learns to extract the feature representations of the images, and determines the transformation relationship between the images by comparing the similarity between the feature vectors output by the two sub-networks.

[0066] The Siamese network infers the transformation relationship by comparing the feature similarities of two images. However, its sub-networks with shared weights are usually based on convolutional operations and are good at capturing local features, such as edges and textures. But for large-scale low-frequency deformations, such as the overall stretching or compression of the train chassis, its global modeling ability is insufficient, resulting in a decrease in registration accuracy. High-resolution images, such as 4096×4096 images collected by a line array camera, need to be input into two sub-networks for parallel processing, resulting in a doubling of memory occupancy and computational complexity, making it difficult to meet the real-time requirements of train online detection.

[0067] Applying a generative adversarial network (GAN) to image registration, it consists of a generator and a discriminator. The generator is responsible for generating the transformed image to make the generated image similar to the target image, and the discriminator judges whether the generated image is real (i.e., whether it matches the target image). During the training process, the generator is continuously optimized to generate a transformed image that better conforms to the characteristics of the target image, thereby learning the registration transformation between images.

[0068] The generator and discriminator of GAN need to be balanced and optimized in the confrontation. During the training process, mode collapse or oscillation is likely to occur, resulting in unstable prediction results of the deformation field and affecting the registration reliability. GAN tends to generate images with rich details but containing high-frequency noise, while the deformation of the train chassis is mainly composed of low-frequency components. The high-frequency noise will interfere with the registration result, and additional post-processing (such as filtering) is required to increase the complexity.

[0069] Based on the residual network (ResNet), the problem of gradient disappearance and gradient explosion during the training of deep neural networks is solved by introducing residual blocks, enabling the network to learn image features at a deeper level. In image registration, the powerful feature extraction ability of ResNet is utilized to extract multi-scale and multi-level features from the image. These features can better describe the structure and details of the image, and then more accurately calculate the transformation matrix between images.

[0070] ResNet constructs a deep network through residual blocks to extract multi-scale features. However, when processing high-resolution images, the number of parameters and computational complexity of the deep network increase sharply, resulting in excessive memory occupancy and significant inference latency, making it difficult to be deployed on embedded devices. The feature extraction of ResNet focuses on progressive learning from local to global, but lacks targeted modeling of the global law of low-frequency deformation, and additional modules need to be combined to improve the accuracy.

[0071] In the train chassis detection scenario, the images collected by the line array camera have a high resolution. For example, 4096×4096, and the speed fluctuation of the train results in the deformation being mainly composed of low-frequency components. For example, overall stretching or compression. Due to the limited receptive field of some networks, it is difficult to accurately predict large-scale low-frequency deformation; at the same time, the high-resolution input leads to a sharp increase in memory occupancy, unable to meet the real-time requirements. Although the global attention of some networks can model low-frequency deformation, the computational complexity is quadratic with the image resolution, resulting in the inference speed being unable to match the real-time requirements of on-train detection, leading to low registration speed and accuracy.

[0072] To solve the problem of low registration speed and accuracy, some embodiments of the present application provide a train chassis registration method based on deep learning and DFT technology. Since the driving direction of the train is almost constant, the deformation of the captured images is along the same direction. Therefore, the deformation of the captured images is low-frequency, and it is necessary to predict the low-frequency deformation. The method inputs the captured image, its corresponding template, and the acceleration collected by the sensor, predicts the low-frequency deformation field in the frequency domain through a neural network, then zero-pads the frequency domain deformation field to make the size of the deformation field matrix equal to that of the original image, and performs IDFT transformation to obtain the deformation field in the spatial domain. The deformation field is input into the deformation layer to perform deformation registration on the captured image. By introducing the DFT technology into image registration, a small-scale neural network is used to obtain the low-frequency deformation field in the frequency domain, and then IDFT transformation is performed to obtain the deformation field, which can accelerate the speed of image registration.

[0073] As Figure 1 , Figure 2 shown, the method includes:

[0074] S100: Obtain an acceleration feature vector.

[0075] The acceleration feature vector is generated from the train acceleration signal. The acceleration signal is generated during the operation of the train, which characterizes the motion state of the train. The acceleration feature vector is a certain piece of information extracted from the acceleration signal, and it is also a set of feature information extracted from the acceleration signal, which is used in the image registration process. The acceleration feature vector can include information such as the speed change of the train, the magnitude and direction of the acceleration.

[0076] To obtain the acceleration feature vector, in some embodiments, it includes: obtaining the acceleration signal, then converting the acceleration signal into an embedding vector through an embedding layer, inputting the embedding vector into a multi-layer perceptron layer to generate an embedding vector, and the lengths of multiple said embedding vectors are different; performing a vector replication operation on the embedding vector to expand the dimension of the embedding vector to a preset dimension to obtain the acceleration feature vector.

[0077] Among them, the acceleration signal can be obtained through an acceleration sensor set on the train chassis or the car body, which characterizes the motion state of the train at different time periods, represents the instantaneous value of the train speed change, is used to quantify the fluctuation degree of the train driving speed, and indirectly characterizes the deformation amplitude when the line array camera captures images.

[0078] The Embedding Layer is a neural network layer used to map discrete or continuous signals into low-dimensional dense vectors (embedding vectors). The Embedding Layer includes an input layer, a fully connected layer, and an activation function. The input layer is used to receive the original acceleration signal, i.e., one-dimensional time-series data. The fully connected layer maps the input to the target dimension through a linear transformation. The activation function, such as ReLU, increases the non-linear expression ability.

[0079] The acceleration signal is converted into a representation in a high-dimensional space, i.e., an embedding vector, through the Embedding Layer. Each embedding vector contains feature information of different aspects of the train motion state, so as to be able to accurately match the train chassis images at different times or from different perspectives.

[0080] Then, the embedding vectors are input into a Multi-Layer Perceptron (MLP) layer to further extract and fuse the feature information. The number of layers in the Multi-Layer Perceptron layer is preset. The preset number can be 3 layers, and the lengths of multiple embedding vectors may be different. A vector replication operation is performed on the embedding vectors to expand their dimension to a preset dimension, where the preset dimension is two-dimensional, thereby obtaining the final acceleration feature vector.

[0081] During the train chassis image registration process, by setting specific conditions, such as the number of Embedding Layers, the configuration of the Multi-Layer Perceptron layer, and vector replication, to generate the acceleration feature vector, it can ensure that the feature vector can accurately represent the train motion state and improve the accuracy and efficiency of image registration.

[0082] S200: Obtain the image to be registered and the template image, and perform channel splicing on the image to be registered and the template image to generate a spliced image.

[0083] The image to be registered is a deformed image that needs to be adjusted to align with the template image. The template image is a reference image used to guide the adjustment of the image to be registered. The image to be registered and the template image may have differences due to factors such as shooting time, angle, and lighting. Therefore, the differences need to be eliminated through the registration process.

[0084] In some embodiments, to obtain the registered image, record the acquisition time of the acceleration feature vector; generate a preset time period based on the acquisition time; and acquire the image to be registered within the preset time period so that the acquisition time of the image to be registered is synchronized with the acquisition time of the acceleration feature vector.

[0085] To align the time of image acquisition with the acceleration signal, a preset time period can be set. The preset time period is a continuous time window defined based on the time obtained from the acceleration feature vector and is used to synchronously acquire the images to be registered. Exemplarily, assuming the acceleration signal acquisition time is t0, the preset time period is [t0 - Δt, t0 + Δt], where Δt is the allowable time deviation, and the value of Δt is determined based on the train running speed and the camera exposure time. For example, when the train speed is 120 km / h, the displacement within 50 ms is about 1.67 meters, and it is necessary to ensure that the image acquisition covers this displacement range.

[0086] By limiting the time range of image acquisition, the acquisition time of the images to be registered can be made to be at the same physical moment or within a very short time window as the acquisition time of the acceleration feature vector, eliminating the correlation error between acceleration and deformation field caused by time asynchronization, improving the registration accuracy, and Δt can also be adaptively adjusted according to the real-time train speed, increasing the time window in high-speed scenarios to cover a larger displacement range.

[0087] After obtaining the registered images aligned with the acceleration signal time, image stitching is then performed. Channel stitching is an image processing technique that combines different channels of multiple images into one image. The stitched image contains all the information of the images to be registered and the template image. In some embodiments, first determine the number of channels of the images to be registered and the template image. If both are single-channel images, the channel data of the image to be registered and the channel data of the template image are arranged in sequence in the channel dimension.

[0088] If the images to be registered and the template image are multi-channel images, for example, three-channel color images, the red, green, and blue channel data of the image to be registered and the corresponding red, green, and blue channel data of the template image can be arranged in sequence in the channel dimension according to the same channel order. They can be sorted in the order of red channel data, green channel data, and blue channel data to generate the stitched image.

[0089] S300: Input the acceleration feature vector and the stitched image into the encoder to predict and generate a low-dimensional space deformation domain through the encoder.

[0090] Input the acceleration feature vector and the stitched image into the encoder. The encoder is a neural network model (Transformer layer) that can map the input data into a low-dimensional space to extract key information and remove redundancy. In some embodiments, the encoder includes a neural network layer and a frequency domain feature extraction layer, and the neural network layer includes a convolutional layer and an attention module.

[0091] The neural network layer consists of a convolutional layer and an attention module. The convolutional layer is used for spatial downsampling and local feature extraction. The local details of the image are captured through the convolutional layer, and the long-range dependencies are modeled through the attention module to improve the accuracy of deformation prediction.

[0092] The convolutional layer is a neural network layer that extracts image features by sliding a convolutional kernel for calculation. The input is the spliced image, that is, the channel splicing result of the image to be registered and the template image. Multiple convolutional kernels, such as 3×3 and 5×5, are used for convolution operations, and downsampling is achieved in cooperation with the stride to generate multi-scale feature maps. For example, the resolution is halved layer by layer. The multi-scale feature maps are output, such as four scales, 1 / 4, 1 / 8, 1 / 16, and 1 / 32 of the original image size. The convolutional layer compresses the image resolution and extracts semantic information at different scales.

[0093] In some embodiments, the spliced image is input into the convolutional layer to perform downsampling through the convolutional layer and generate multi-scale feature maps; the acceleration feature vector and the multi-scale feature maps are added vectorially at corresponding positions to output a channel-adjusted feature map; through the attention module, local features of the channel-adjusted feature map are extracted to generate a low-dimensional spatial deformation domain.

[0094] The convolutional kernel in the convolutional layer slides on the spliced image for convolution operations. Downsampling can be achieved through convolutional kernels of different sizes and strides, that is, reducing the size of the image. At the same time, each convolutional kernel extracts different features of the image to generate multiple feature maps, and the feature maps contain image feature information at different scales. Downsampling can reduce the dimension of the data and the computational complexity, and the multi-scale feature maps can capture the features of the image at different scales.

[0095] The previously obtained acceleration feature vector and the multi-scale feature maps output by the convolutional layer are subjected to a vector addition operation at corresponding positions. Before the addition, the dimension of the acceleration feature vector needs to be adjusted to match the dimension of the multi-scale feature maps. Incorporating the acceleration information of the train into the image features enables the feature maps to contain not only the visual information of the image but also the motion information of the train, which helps to improve the accuracy of subsequent deformation prediction.

[0096] Exemplarily, the image to be registered and the template image are spliced in the channel dimension, and four-scale feature maps are generated by gradually downsampling through four convolutional blocks. Each convolutional block includes a convolutional layer, batch normalization, and ReLU activation, and can extract multi-scale features, retaining local details and global structures. The acceleration feature vector is extended to a size matching the current-scale feature map through a replication operation, and the extended acceleration feature and the feature map are added element-wise to generate a channel-adjusted feature map, embedding the train speed change information into the image features and enhancing the physical relevance of deformation prediction.

[0097] To extract local features of the channel-adjusted feature map to generate a low-dimensional spatial deformation domain, in some embodiments, it includes:

[0098] Divide the channel-adjusted feature map into local regions, and the sizes of multiple said local regions are the same;

[0099] Expand the local region into a feature unit, and the feature unit is in one-dimensional form;

[0100] Perform self-attention calculation on the feature unit to output a feature representation, and the self-attention calculation is to pass the feature unit through a preset number of linear layers;

[0101] Based on the feature representation, perform cross-attention operation on the feature unit to output a fusion result, and the cross-attention operation is to perform information fusion of different scales on the feature unit;

[0102] Based on the fusion result, output a low-dimensional spatial deformation domain.

[0103] The local region refers to several small blocks obtained by dividing the channel-adjusted feature map. Each small block is called a local region. These regions have the same size and are used to capture the local features of the image. A fixed-size window can be slid on the channel-adjusted feature map to divide the entire feature map into multiple non-overlapping local regions, and the size of each region is the same. The feature vector obtained after expanding the local region into one-dimensional form is the feature unit, and the feature unit contains all the information of the local region.

[0104] As Figure 3 shown, the attention module includes a self-attention unit (Self-Attention) and a cross-attention unit (Cross-Attention), which are used for global feature correlation and multi-scale information fusion. The self-attention unit calculates the correlation between positions within the same feature map. The cross-attention unit transmits information between feature maps of different scales to achieve cross-scale feature fusion.

[0105] Input the channel-adjusted feature map into the attention module. The attention module processes the channel-adjusted feature map, calculates the correlation between the features at each position and the features at other positions, and obtains a self-attention representation. The cross-attention layer fuses the self-attention representations of different scales and integrates information of different scales. According to the fusion result, a low-dimensional spatial deformation domain is output, and this deformation domain represents the possible deformation situation of the image to be registered relative to the template image.

[0106] The attention module can automatically focus on important local features in the channel-adjusted feature map and enhance the expression ability of features. By integrating information of different scales, the generated low-dimensional spatial deformation domain can more accurately characterize the deformation situation of the image.

[0107] Self-attention calculation is used to capture the dependencies between feature units. Through a preset number of linear layers, the feature units are mapped to different spaces, and the attention weights in different spaces are calculated, thereby realizing the information interaction between feature units.

[0108] Cross-attention operation is a method for fusing information of different scales of feature units. Through cross-attention operation, feature units of different scales can be fused to obtain a more abundant feature representation.

[0109] Exemplarily, for a channel-adjusted feature map (such as size H / 4×W / 4×256), the feature map is divided into 8×8 patches, a total of (H / 4) / 8×(W / 4) / 8 patches are generated, and each patch is unfolded into a one-dimensional token with a length of 8×8×256 = 16,384, converting the high-dimensional spatial features into a serialized representation for easy processing by the attention mechanism. For each feature map (token) sequence, Q1, K1, and V1 matrices are generated through three linear layers. Calculate the attention weights and weighted fusion value vectors, and output the enhanced feature sequence to enhance the consistency of local deformation patterns (such as the displacement correlation in the wheel area), as shown in the following formula:

[0110]

[0111] where Q1, K1, and V1 are obtained after the same feature map passes through three different linear layers, and d k is the embedding dimension of the feature map, and Softmax is the normalization operation.

[0112] Take the token sequence of the current scale as Q2, and the token sequences of the previous scale as K2 and V2. Calculate the cross-scale attention weights to fuse the coarse-grained and fine-grained features. Integrate multi-scale information (such as the correlation between the overall deformation of the frame and the local displacement of the bolt), as shown in the following formula:

[0113]

[0114] where Q2 is obtained after the current layer feature map passes through a linear layer, and K2 and V2 are obtained after the previous layer feature map passes through two linear layers.

[0115] Input the fused feature sequence into the MLP layer to compress it into a low-dimensional space (such as 64 channels). Output a low-dimensional deformation field with a size of H / 32×W / 32×64. Reduce the computational complexity of frequency domain conversion while retaining the core deformation information. Through this process, the local features of the channel-adjusted feature map can be extracted, and a low-dimensional space deformation domain can be generated. The low-dimensional space deformation domain can accurately represent the deformation relationship between images, which helps to improve the accuracy and efficiency of image registration.

[0116] S400: Perform frequency-domain conversion on the low-dimensional space deformation domain to obtain a low-frequency space deformation field.

[0117] Frequency-domain conversion is a mathematical transformation that can convert a signal in the time domain or spatial domain into a signal in the frequency domain. In this embodiment, frequency-domain conversion is used to extract low-frequency information in the deformation domain. The low-frequency information characterizes the deformation trend of the image, while the high-frequency information may contain noise and details. The low-frequency space deformation field is the result of frequency-domain conversion and contains the main deformation information in the deformation domain.

[0118] In some embodiments, for the low-dimensional space deformation domain, frequency-domain conversion is performed through a discrete Fourier transform layer to obtain a frequency-domain deformation domain, and the frequency-domain deformation domain contains low-frequency information; for the frequency-domain deformation domain, frequency-domain conversion is performed through a discrete Fourier transform layer to obtain a low-frequency space deformation field.

[0119] The low-dimensional space deformation domain is a deformation feature map obtained after feature extraction and compression, which characterizes the displacement vectors of each region of the image and has a lower dimension than the original image. Frequency-domain conversion (FFT layer) converts the spatial-domain signal into a frequency-domain signal through discrete Fourier transform (DFT), maps the deformation field from the spatial domain to the frequency domain, and separates the low-frequency global deformation and high-frequency noise.

[0120] The frequency-domain deformation domain is a complex matrix representing the deformation field in the frequency domain, which contains amplitude and phase information. The data is compressed through frequency-domain sparsity to retain the main deformation mode, that is, the low-frequency component.

[0121] The low-frequency space deformation field intercepts the low-frequency component from the frequency-domain deformation domain and inverse-transforms it back to the spatial-domain deformation field, only retaining the central region of the frequency domain, filtering out high-frequency noise, eliminating high-frequency interference, and retaining the main low-frequency displacements related to the deformation of the train chassis.

[0122] Exemplarily, perform a discrete Fourier transform on the low-dimensional space deformation domain to obtain a frequency-domain deformation domain, convert the spatial-domain signal into a frequency-domain signal, then screen the frequency-domain deformation domain to extract low-frequency information. Noise and detail interference can be removed, the main deformation information can be retained, and finally the low-frequency information is combined into a low-frequency space deformation field.

[0123] Furthermore, the signal can be converted from the spatial domain to the frequency domain through Fourier transform (DFT). The frequency-domain representation can more efficiently capture the main features of the signal, especially the low-frequency components, which contain most of the information of the signal. By retaining the low-frequency components, information compression can be achieved, reducing the computational amount and storage requirements. In the field of computer vision, the edges and textures of an image usually appear as high-frequency components in the frequency domain, while the background, large-scale structures, and semantic information appear as low-frequency components. By processing these features in the frequency domain, feature extraction and analysis can be performed more effectively. Since the driving speed and direction of the train change very little, most of the deformations of the chassis pictures collected are low-frequency.

[0124] By performing frequency domain conversion through a Fourier transform layer and extracting low-frequency information, a low-frequency spatial deformation field representing the main deformation relationship between images can be obtained, which helps to improve the accuracy and efficiency of image registration. At the same time, noise and detail interference can be removed, and the main deformation information can be retained, thereby improving the accuracy of deformation analysis.

[0125] S500: Reconstruct the low-frequency spatial deformation field through a decoder to generate a spatial deformation field.

[0126] The decoder is the inverse process of the encoder and can restore the information in the low-dimensional space to the high-dimensional space. In this embodiment, the decoder takes the low-frequency spatial deformation field as the input and reconstructs a complete spatial deformation field, and the spatial deformation field is a full-resolution spatial deformation field.

[0127] In some embodiments, the decoder gradually restores the image resolution through upsampling, for example, transposed convolution or interpolation, to improve the local accuracy of the deformation field. However, the upsampling of the decoder may introduce artifacts, and at the same time, the computational amount and memory occupation increase sharply.

[0128] In this embodiment, the decoder includes a zero-padding layer (pad layer) and an inverse discrete Fourier transform layer (IDFT layer). The zero-padding layer is an operation layer that expands the matrix size by filling zeros and is used to restore the spatial resolution of the frequency domain deformation field. The inverse discrete Fourier transform layer is a mathematical operation module that inverse-transforms the frequency domain signal back to the spatial domain and is implemented based on the inverse discrete Fourier transform.

[0129] The input of the zero-padding layer is the low-frequency spatial deformation field. Centering on the low-frequency deformation field, zeros are filled in the periphery to expand its size to be the same as that of the original image, such as H×W×64. The filling rule is that if the size of the original image is H×W, then (H - H / 32) / 2 and (W - W / 32) / 2 zeros are filled on the top, bottom, left, and right respectively. By zero-padding, the low-frequency deformation information is retained, high-frequency noise interference is avoided, and at the same time, the spatial resolution of the deformation field is restored.

[0130] The input of the inverse discrete Fourier transform layer is the frequency domain deformation field after zero-padding, with a size of H×W×64. Perform two-dimensional IDFT on the frequency domain signal of each channel to convert the low-frequency deformation information from the frequency domain back to the spatial domain, generating a full-resolution displacement field aligned with the original image, that is, the spatial deformation field. Both the zero-padding layer and the IDFT layer are differentiable. Therefore, the Fourier network can be optimized through standard backpropagation, converting the displacement field into a spatial deformation field. The whole process is based on mathematical rules and does not require training parameters. Therefore, it is fast and stable and can be jointly optimized with the neural network.

[0131] The spatial deformation field is the displacement field matrix output by the IDFT layer, which characterizes the deformation displacement vector of each pixel in the image. It includes a real tensor containing horizontal and vertical displacement components and is used to perform deformation correction on the input image to eliminate geometric distortion during the acquisition process.

[0132] In some embodiments, zero-padding is performed on the low-frequency spatial deformation field through the zero-padding layer to generate a restored image with the same size as the image to be registered. The restored image is input into the inverse discrete Fourier transform layer to perform the inverse discrete Fourier transform through the inverse discrete Fourier transform layer, and the low-frequency spatial deformation field is restored to the spatial deformation field.

[0133] Exemplarily, if the size of the low-frequency spatial deformation field is H / 32×W / 32×64, zero values are symmetrically filled around the frequency-domain deformation field to expand its size to H×W×64. If the original image is 1024×1024 and the low-frequency deformation field is 32×32, then (1024 - 32) / 2 = 496 zero values are filled on each side. This can keep the low-frequency deformation information complete, and the high-frequency region is filled with zero values to suppress noise interference. The size of the deformation field is consistent with the original image, providing an input basis for subsequent IDFT.

[0134] The inverse discrete Fourier transform (IDFT) performs two-dimensional IFFT on each channel after zero-padding to generate a spatial-domain displacement field. Only the real part is retained (the imaginary part is close to zero due to symmetry), and the output size is H×W×64. The 64 channels are compressed into 2 channels (Δx, Δy) through a linear projection layer to obtain the final spatial deformation field. It can restore the full-resolution displacement field, characterize the deformation displacement of each pixel, and is computationally efficient. The complexity of IFFT is logarithmically related to the image size.

[0135] Normalize the displacement field to limit the displacement range (e.g., Δx, Δy ∈ [-10, 10] pixels). Optionally, Gaussian smoothing filtering can be used to eliminate the small oscillations that may be introduced by the IDFT. Ensure that the displacement field is physically reasonable and avoid image tearing caused by extreme deformations. Improve the smoothness and stability of the registration result.

[0136] Through the processing of the zero-padding layer and the inverse discrete Fourier transform layer, the decoder restores the low-frequency spatial deformation field to the spatial deformation field, which can improve the accuracy and efficiency of image registration.

[0137] S600: Perform deformation correction on the image to be registered based on the spatial deformation field and output the target registered image.

[0138] The spatial deformation field is the full-resolution deformation field or differential deformation of the image to be registered. Deformation correction adjusts the image according to the information in the deformation field to eliminate the deformation. Deformation correction adjusts each pixel point in the image to be registered according to the information in the spatial deformation field to align it with the template image. The finally output target registered image is the result of the registration process, which is consistent with the template image in content but has the deformation eliminated after adjustment.

[0139] In this embodiment, through a reconstruction unit, namely the Warp layer (spatial distortion layer), deformation correction of the image to be registered is performed based on the spatial deformation field, and the target registered image is output. The Warp layer converts the geometric transformation of the deformation field into pixel-level image reconstruction through coordinate mapping and linear interpolation processes.

[0140] The spatial deformation field includes a first coordinate axis displacement amount and a second coordinate axis displacement amount. In some embodiments,

[0141] Based on the first coordinate axis displacement amount and the second coordinate axis displacement amount, the original image coordinates of the pixel point, where the pixel point is a pixel point in the target registered image, are calculated;

[0142] If the original image coordinates are non-integer coordinates, adjacent coordinates are obtained, where the adjacent coordinates are the pixel values of the four adjacent integer coordinate points around the non-integer coordinates;

[0143] Based on the adjacent coordinates, the interpolation result is calculated by the bilinear interpolation method, where the interpolation result is the pixel value of the non-integer coordinates;

[0144] The interpolation result is used as the pixel value of the corresponding pixel point in the target registered image to generate the target registered image.

[0145] In this embodiment, the spatial deformation field is defined as the deformation amount of each point in the target registered image relative to the original image, and is composed of matrices of two channels, with each channel representing the displacement amounts (dx and dy) on the x-axis and y-axis respectively. For example, for the target registered image, where each point is (i′, j′), the corresponding coordinates in the image to be registered are (x, y), which can be calculated by the following formula:

[0146] x = i′ + dx[i′, j′];

[0147] y = j′ + dy[i′, j′];

[0148] Therefore, as long as (i′, j′) is obtained, based on the deformation field, (x, y) can be calculated.

[0149] Since the train chassis is a continuously deformable feature, directly taking the nearest integer point, that is, nearest neighbor interpolation, will cause image jaggedness. In order to make the bilinear interpolation transition smooth, in some embodiments, the horizontal weight value is calculated based on the horizontal distance ratio between the non-integer coordinate and the adjacent coordinate; and the vertical weight value is calculated based on the vertical distance ratio between the non-integer coordinate and the adjacent coordinate; then, based on the horizontal weight value and the vertical weight value, two linear interpolations are performed on the adjacent coordinates to obtain the interpolation result.

[0150] Bilinear interpolation solves the problem of not being able to directly obtain pixel values from non-integer coordinates. For example, after deformation correction, the original image coordinates corresponding to a pixel point might be (10.5, 9.6). In this case, the pixel value at that location is calculated through a weighted calculation based on the pixel values of the four surrounding integer coordinates (such as 10, 9; 10, 10; 11, 9; 11, 10).

[0151] For example, in the horizontal direction, for a non-integer coordinate (10.5, 9.6), the ratio of the distance to the left integer point (10, 9.6) on the horizontal x-axis is 0.5 (10.5-10=0.5), and the ratio to the right integer point (11, 9.6) is also 0.5. The weight of the left point is 1-0.5=0.5, and the weight of the right point is 0.5.

[0152] For example, in the vertical direction, i.e., on the y-axis, the distance ratio of the coordinate to the upper integer point (10.5, 9) is 0.6 (9.6-9=0.6), and the distance ratio to the lower integer point (10.5, 10) is 0.4 (10-9.6=0.4). The weight of the upper point is 1-0.6=0.4, and the weight of the lower point is 0.6.

[0153] After obtaining the weight values in the horizontal and vertical directions, bilinear interpolation is performed. The first interpolation is in the horizontal direction. The horizontal interpolation results are calculated for the upper and lower rows respectively. For the upper row, y=9, the pixel value of the left point × 0.5 + the pixel value of the right point × 0.5; for the lower row, y=10, the pixel value of the left point × 0.5 + the pixel value of the right point × 0.5.

[0154] The second interpolation is in the vertical direction. The interpolation results of the upper and lower rows are mixed according to the vertical weight. The final pixel value = upper interpolation result × 0.4 + lower interpolation result × 0.6.

[0155] By performing interpolation horizontally and then vertically, a two-dimensional interpolation result is synthesized, thereby smoothly restoring the pixel values of non-integer coordinates. This can not only eliminate image deformation but also avoid the risk of false detection due to interpolation errors.

[0156] Based on the above train chassis registration method based on deep learning and DFT technology, some embodiments of the present application also provide a train chassis registration system based on deep learning and DFT technology, including:

[0157] An acquisition unit, configured to acquire an acceleration feature vector, where the acceleration feature vector is generated from a train acceleration signal; and, acquire a to-be-registered image and a template image, and perform channel splicing on the to-be-registered image and the template image to generate a spliced image;

[0158] A processing unit, configured to input the acceleration feature vector and the spliced image into an encoder;

[0159] An encoder, configured to predict and generate a low-dimensional space deformation domain; and perform frequency domain conversion on the low-dimensional space deformation domain to obtain a low-frequency space deformation field;

[0160] A decoder, configured to reconstruct the low-frequency space deformation field to generate a space deformation field;

[0161] A reconstruction unit, configured to perform deformation correction on the to-be-registered image based on the space deformation field and output a target registered image.

[0162] As can be seen from the above technical solutions, the present application provides a train chassis registration method and system based on deep learning and DFT technology, including: acquiring an acceleration feature vector, where the acceleration feature vector is generated from a train acceleration signal, then acquiring a to-be-registered image and a template image, performing channel splicing on the to-be-registered image and the template image to generate a spliced image, inputting the acceleration feature vector and the spliced image into an encoder, the encoder predicting and generating a low-dimensional space deformation domain, then performing frequency domain conversion on the low-dimensional space deformation domain to obtain a low-frequency space deformation field, reconstructing the low-frequency space deformation field by a decoder to generate a space deformation field, performing deformation correction on the to-be-registered image based on the space deformation field, and outputting a target registered image. By introducing DFT technology into image registration, obtaining its low-frequency space deformation field with a small-scale neural network, and then performing IDFT transformation to obtain the space deformation field, the speed and quality of image registration can be accelerated.

[0163] For the similar parts between the embodiments provided in the present application, reference can be made to each other. The specific embodiments provided above are only several examples under the general concept of the present application and do not constitute a limitation on the protection scope of the present application. For those skilled in the art, any other embodiments extended based on the solution of the present application without creative efforts belong to the protection scope of the present application.

Claims

1. A train chassis registration method based on deep learning and DFT technology, characterized in that Including: Obtaining an acceleration feature vector, where the acceleration feature vector is generated from a train acceleration signal; Obtaining an image to be registered and a template image, and performing channel splicing on the image to be registered and the template image to generate a spliced image; Inputting the acceleration feature vector and the spliced image into an encoder to predict and generate a low-dimensional space deformation domain through the encoder; Performing a frequency domain conversion on the low-dimensional space deformation domain to obtain a low-frequency space deformation field; Reconstructing the low-frequency space deformation field through a decoder to generate a space deformation field; Performing deformation correction on the image to be registered based on the space deformation field and outputting a target registered image.

2. The train chassis registration method based on deep learning and DFT technology according to claim 1, wherein, The obtaining of the acceleration feature vector includes: Obtaining an acceleration signal; Converting the acceleration signal into an embedding vector through an embedding layer; Inputting the embedding vector into a multi-layer perceptron layer to generate an embedding vector, where the number of multi-layer perceptron layers is a preset number, and the lengths of multiple embedding vectors are different; Performing a vector replication operation on the embedding vector to expand the dimension of the embedding vector to a preset dimension to obtain an acceleration feature vector.

3. The train chassis registration method based on deep learning and DFT technology according to claim 2, characterized in that, The encoder includes a neural network layer and a frequency domain feature extraction layer, and the neural network layer includes a convolutional layer and an attention module; The inputting of the acceleration feature vector and the spliced image into the encoder to predict and generate a low-dimensional space deformation domain through the encoder includes: Inputting the spliced image into the convolutional layer to perform downsampling through the convolutional layer and generate a multi-scale feature map; Adding the acceleration feature vector and the multi-scale feature map by vector to output a channel-adjusted feature map; Extracting local features of the channel-adjusted feature map through the attention module to generate a low-dimensional space deformation domain.

4. The method for registering a train chassis based on deep learning and DFT technology according to claim 3, wherein The extracting of local features of the channel-adjusted feature map to generate a low-dimensional space deformation domain includes: Dividing the channel-adjusted feature map into local regions, where the sizes of multiple local regions are the same; Unfolding the local regions into feature units, where the feature units are in one-dimensional form; Performing self-attention calculation on the feature units to output a feature representation, where the self-attention calculation is to pass the feature units through a preset number of linear layers; Performing cross-attention operation on the feature units based on the feature representation to output a fusion result, where the cross-attention operation is to perform information fusion of different scales on the feature units; Outputting a low-dimensional space deformation domain based on the fusion result.

5. The train chassis registration method based on deep learning and DFT technology according to claim 1, characterized in that The performing of frequency domain conversion on the low-dimensional space deformation domain to obtain a low-frequency space deformation field includes: Performing frequency domain conversion on the low-dimensional space deformation domain through a discrete Fourier transform layer to obtain a frequency domain deformation domain, where the frequency domain deformation domain contains low-frequency information; Performing frequency domain conversion on the frequency domain deformation domain through a discrete Fourier transform layer to obtain a low-frequency space deformation field.

6. The train chassis registration method based on deep learning and DFT technology according to claim 1, characterized in that The decoder includes a zero-padding layer and an inverse discrete Fourier transform layer; The reconstructing of the low-frequency space deformation field through the decoder to generate a space deformation field includes: Zero-padding is performed on the low-frequency spatial deformation field through the zero-padding layer to generate a restored image, and the restored image has the same size as the image to be registered; The restored image is input into the inverse discrete Fourier transform layer to perform an inverse discrete Fourier transform through the inverse discrete Fourier transform layer, and the low-frequency spatial deformation field is restored to a spatial deformation field, and the spatial deformation field is a full-resolution spatial deformation field.

7. The method for registering a train chassis based on deep learning and DFT technology according to claim 1, wherein The spatial deformation field includes a first coordinate axis displacement amount and a second coordinate axis displacement amount; Performing deformation correction on the image to be registered based on the spatial deformation field and outputting a target registered image includes: Calculating the original image coordinates of a pixel point based on the first coordinate axis displacement amount and the second coordinate axis displacement amount, where the pixel point is a pixel point in the target registered image; If the original image coordinates are non-integer coordinates, obtaining adjacent coordinates, where the adjacent coordinates are the pixel values of the four adjacent integer coordinate points around the non-integer coordinates; Calculating an interpolation result based on the adjacent coordinates through bilinear interpolation, where the interpolation result is the pixel value of the non-integer coordinates; Using the interpolation result as the pixel value of the corresponding pixel point in the target registered image to generate the target registered image.

8. The method for registering a train chassis based on deep learning and DFT technology according to claim 7, characterized in that Calculating the interpolation result through bilinear interpolation includes: Calculating a horizontal direction weight value according to the horizontal direction distance ratio between the non-integer coordinates and the adjacent coordinates; and calculating a vertical direction weight value according to the vertical direction distance ratio between the non-integer coordinates and the adjacent coordinates; Performing two linear interpolations on the adjacent coordinates based on the horizontal direction weight value and the vertical direction weight value to obtain an interpolation result.

9. The method for registering a train chassis based on deep learning and DFT technology according to claim 1, wherein Obtaining the image to be registered includes: Recording the acquisition time of the acceleration feature vector; Generating a preset time period based on the acquisition time; Obtaining the image to be registered within the preset time period so that the acquisition time of the image to be registered is synchronized with the acquisition time of the acceleration feature vector.

10. A train chassis registration system based on deep learning and DFT technology, characterized in that, Including: An acquisition unit for acquiring an acceleration feature vector, where the acceleration feature vector is generated from a train acceleration signal; And acquiring the image to be registered and a template image, and performing channel splicing on the image to be registered and the template image to generate a spliced image; A processing unit for inputting the acceleration feature vector and the spliced image into an encoder; The encoder is used for predicting and generating a low-dimensional spatial deformation domain; And performing a frequency domain conversion on the low-dimensional spatial deformation domain to obtain a low-frequency spatial deformation field; A decoder for reconstructing the low-frequency spatial deformation field to generate a spatial deformation field; A reconstruction unit for performing deformation correction on the image to be registered based on the spatial deformation field and outputting a target registered image.