Track anomaly detection method and device
By monitoring ambient light intensity in real time and using visible and infrared light image acquisition modules combined with super-resolution reconstruction technology, the problem of insufficient light in track anomaly detection systems under complex environments has been solved, enabling high-quality image acquisition and accurate detection in any lighting environment.
Patent Information
- Application Number
- CN202511202433.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-26
- Publication Date
- 2026-01-23
AI Technical Summary
Existing track anomaly detection systems struggle to acquire high-quality track images in complex environments (such as nighttime, rainy days, tunnels, foggy days, or dust obstruction) due to insufficient ambient light, resulting in a significant drop in detection accuracy.
By monitoring ambient light intensity in real time, a visible light image acquisition module is used to acquire visible light trajectory images when the light is sufficient, and an infrared image acquisition module is used to acquire infrared trajectory images when the light is insufficient. The images are then super-resolution reconstructed to ensure high-quality image data is obtained in any lighting environment.
It can acquire high-quality image data under any external lighting conditions, improving the accuracy and reliability of track anomaly detection.
Smart Images

Figure CN121391702A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of track monitoring, and in particular to a track anomaly detection method and device. BACKGROUND
[0002] With the rapid development of urban rail transit systems, real-time monitoring of track structure safety and operating conditions has become a key link in ensuring train operation safety.
[0003] Currently, automatic track anomaly detection is realized by an intelligent identification system based on track image recognition. The intelligent identification system acquires a photographed track image and identifies track anomalies through image recognition technology. High-quality track images are a key factor in accurately detecting track anomalies. However, in complex environments (such as at night, in the rain, in tunnels, in fog, or with dust obstruction), it is difficult to obtain high-quality track images due to insufficient external environmental light, resulting in a significant decrease in the accuracy of track anomaly detection. SUMMARY
[0004] The present application provides a track anomaly detection method and device to solve the problem of insufficient external environmental light, which makes it difficult to obtain high-quality track images, resulting in a significant decrease in the accuracy of track anomaly detection in the prior art.
[0005] The present application provides a track anomaly detection method, comprising the following steps: acquiring real-time light intensity of external environmental light; in the case that the real-time light intensity is greater than a light intensity threshold, acquiring a visible light track image based on visible light collection, determining that the visible light track image is a target track image, and in the case that the real-time light intensity is less than or equal to the light intensity threshold, acquiring an infrared track image based on infrared light collection, determining that the infrared track image is a target track image; performing super-resolution reconstruction on the target track image to obtain a to-be-detected track image with a higher resolution than the target track image; performing anomaly detection on the track based on the to-be-detected track image.
[0006] According to the track anomaly detection method provided by the present application, the target track image is the visible light track image; performing super-resolution reconstruction on the target track image to obtain a to-be-detected track image with a higher resolution than the target track image, comprising: inputting the visible light track image into a visible light image super-resolution reconstruction model to obtain a to-be-detected track image with a higher resolution than the visible light track image output by the visible light image super-resolution reconstruction model; The visible light image super-resolution reconstruction model is trained based on a sample visible light track image and a label visible light track image corresponding to the sample visible light track image, and the resolution of the label visible light track image is higher than that of the sample visible light track image.
[0007] According to the track anomaly detection method provided in the application, the target track image is the infrared track image. The target track image is subjected to super-resolution reconstruction to obtain a to-be-detected track image with a higher resolution than the target track image, including: The infrared track image is input into an infrared image super-resolution reconstruction model to obtain a to-be-detected track image with a higher resolution than the infrared track image output by the infrared image super-resolution reconstruction model; The infrared image super-resolution reconstruction model is trained based on a sample infrared track image and a label infrared track image corresponding to the sample infrared track image, and the resolution of the label infrared track image is higher than that of the sample infrared track image.
[0008] According to the track anomaly detection method provided in the application, the infrared image super-resolution reconstruction model includes: A shallow feature extraction module is configured to extract features of the infrared track image to obtain shallow features. A deep feature extraction module is configured to extract features based on the shallow features to obtain deep features, and combine the shallow features and the deep features to obtain combined features. An image reconstruction module is configured to up-sample the combined features to obtain the to-be-detected track image after super-resolution reconstruction.
[0009] According to the track anomaly detection method provided in the application, the shallow feature extraction module includes a first convolution module, and the deep feature extraction module includes a plurality of cascaded second convolution modules, a large kernel attention module and a third convolution module. The first convolution module is configured to input the extracted shallow features into the plurality of cascaded second convolution modules to sequentially extract features to obtain first intermediate features. The plurality of cascaded second convolution modules input the first intermediate features into the large kernel attention module to obtain second intermediate features output by the large kernel attention module. The third convolution module extracts features of the second intermediate features to obtain the deep features.
[0010] According to the track anomaly detection method provided in the application, the first convolution module, the second convolution module and the third convolution module are all edge-oriented convolution blocks.
[0011] According to the track anomaly detection method provided by the application, the large core attention module comprises a deep convolution set layer, a deep hole convolution layer, a target convolution layer and a multiplier, the deep convolution set layer is connected with the deep hole convolution layer, the deep hole convolution layer is connected with the target convolution layer, and the target convolution layer is connected with the multiplier; The first intermediate feature sequentially passes through the deep convolution set layer, the deep hole convolution layer and the target convolution layer for feature extraction, and forms a third intermediate feature; The multiplier is used for multiplying the first intermediate feature and the third intermediate feature to obtain the second intermediate feature; In the model training stage, the deep convolution set layer comprises an adder and a plurality of first deep convolution layers arranged in parallel, the plurality of first deep convolution layers are connected with the adder, and the adder is connected with the deep hole convolution layer; each first deep convolution layer is used for feature extraction on the first intermediate feature, and the adder inputs the features extracted by the first deep convolution layers into the deep hole convolution layer after adding the features; In the model inference stage, the deep convolution set layer comprises a second deep convolution layer, and the second deep convolution layer is connected with the deep hole convolution layer; The weight of the second deep convolution layer is the sum of the weights of the first deep convolution layers, and the bias of the second deep convolution layer is the sum of the biases of the first deep convolution layers.
[0012] The application further provides a track anomaly detection device, comprising the following units: A real-time light intensity acquisition unit is configured to acquire real-time light intensity of ambient environment light. A target track image determination module is configured to, when the real-time light intensity is greater than a light intensity threshold, acquire a visible light track image based on visible light collection and determine the visible light track image as a target track image, and when the real-time light intensity is less than or equal to the light intensity threshold, acquire an infrared track image based on infrared light collection and determine the infrared track image as a target track image. A super-resolution reconstruction unit is configured to perform super-resolution reconstruction on the target track image to obtain a to-be-detected track image with a higher resolution than the target track image. An anomaly detection unit is configured to perform anomaly detection on a track based on the to-be-detected track image.
[0013] The application further provides an electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, wherein the processor implements the track anomaly detection method according to any one of the above when executing the program.
[0014] The application further provides a non-transitory computer-readable storage medium, which stores a computer program, and the computer program is executed by a processor to implement the track anomaly detection method according to any one of the above.
[0015] The track anomaly detection method and device provided by the application can ensure that high-quality image data can be obtained under any external light environment, so that an accurate anomaly detection result can be obtained, and whether the visible light track image or the infrared track image is obtained, the image data obtained through the super-resolution reconstruction has higher resolution and richer image details, so that the track anomaly result detected based on the image data with higher resolution and richer details is more accurate. BRIEF DESCRIPTION OF DRAWINGS
[0016] In order to more clearly illustrate the technical solutions of the application or the prior art, the following will briefly introduce the drawings needed to be used in the embodiments or the prior art description. Obviously, the drawings in the following description are some embodiments of the application, and those skilled in the art can obtain other drawings according to these drawings without creative labor.
[0017] Figure 1 is a flowchart of the track anomaly detection method provided by the application.
[0018] Figure 2 is a visible light image super-resolution reconstruction model structure diagram in the track anomaly detection method provided by the application.
[0019] Figure 3 is an infrared image super-resolution reconstruction model structure diagram in the track anomaly detection method provided by the application.
[0020] Figure 4 is an edge-oriented convolution block structure diagram of the infrared image super-resolution reconstruction model in the track anomaly detection method provided by the application.
[0021] Figure 5 is a large kernel attention module structure diagram of the infrared image super-resolution reconstruction model in the track anomaly detection method provided by the application.
[0022] Figure 6 is a structure diagram of the track anomaly detection device provided by the application.
[0023] Figure 7 is a structure diagram of the electronic device provided by the application. DETAILED DESCRIPTION
[0024] In order to make the objects, technical solutions and advantages of the present application clearer, the technical solutions in the present application will be described clearly and completely below with reference to the drawings in the present application. Obviously, the described embodiments are only some, but not all of the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative work fall within the protection scope of the present application.
[0025] The track anomaly detection method of the embodiment of the present application, as shown in Figure 1 includes the following steps S110 to S140.
[0026] Step S110: acquiring real-time light intensity of external environment light. Specifically, the real-time light intensity of the external environment can be detected by an ambient light perception module, such as a brightness sensor, installed on a track detection vehicle or a patrol platform. In this step, the real-time light intensity of the external environment light can be acquired from the ambient light perception module.
[0027] Step S120: acquiring a visible light track image based on visible light collection in the case that the real-time light intensity is greater than a light intensity threshold, and determining the visible light track image as a target track image; acquiring an infrared track image based on infrared light collection in the case that the real-time light intensity is less than or equal to the light intensity threshold, and determining the infrared track image as the target track image.
[0028] Specifically, a visible light image collection module (such as a visible light camera) and an infrared image collection module (such as an infrared camera) are installed on the track detection vehicle or the patrol platform. In the case that the real-time light intensity is greater than a light intensity threshold (for example, 1000 lux), it indicates that the light is sufficient, and the visible light image collection module is triggered to collect a visible light track image, which is determined as the target track image. In the case that the real-time light intensity is less than or equal to the light intensity threshold (for example, 1000 lux), it indicates that the light is insufficient, and the image collected by the visible light may not be clear. Therefore, the infrared image collection module is triggered to collect an infrared track image, which is determined as the target track image.
[0029] The ambient light perception module is used to monitor the light intensity of the current environment in real time. When the light intensity is greater than a set light intensity threshold (such as a sunny day), the visible light image collection module is triggered to collect a visible light track image. When the light intensity is less than or equal to the set light intensity threshold (such as an overcast day, a rainy day, night or a tunnel scene), the infrared image collection module is triggered to collect an infrared track image, so that better quality image data can be acquired in any external light environment, and thus an accurate anomaly detection result can be obtained.
[0030] Step S130: super-resolution reconstruction is performed on the target track image to obtain a to-be-detected track image with a higher resolution than the target track image.
[0031] Most current track anomaly detection systems still use original resolution images, which cannot capture small abnormalities such as track cracks, foreign matter intrusion, and missing bolts, affecting detection accuracy and diagnosis depth. In order to more accurately detect track anomalies, in this step, the target track image is super-resolution reconstructed, that is, whether it is a visible light track image or an infrared track image, a higher resolution and more detailed image data is obtained through super-resolution reconstruction, so that in subsequent track anomaly detection, the track anomaly result detected based on the higher resolution and more detailed image data is more accurate.
[0032] Step S140: performing anomaly detection on the track based on the to-be-detected track image. Specifically, the detectable track anomaly types include but are not limited to track cracks (vertical cracks, horizontal cracks, micro cracks), missing spikes, loose bolts, tie breakage, foreign matter intrusion (stones, tools, debris), and tie deviation.
[0033] For example, the to-be-detected track image can be input into a track anomaly detection model to obtain a detection result output by the track anomaly detection model, and the track anomaly detection model is trained based on sample to-be-detected track images and result labels (including anomaly labels and no anomaly labels) corresponding to the sample to-be-detected track images.
[0034] For example, the track anomaly detection model is a target detection network YOLOv5 based on deep learning, and the YOLOv5 network structure includes three parts: a backbone network, a neck structure, and a head structure. Among them, the backbone network is responsible for feature extraction; the neck structure is responsible for feature fusion; the head structure includes three detection heads, which are respectively responsible for outputting detection information of three size levels of abnormal target sizes from large to small, and each detection head can detect multiple abnormalities of its corresponding size level.
[0035] In the track anomaly detection method of this embodiment, the light intensity of the current environment is monitored in real time, when the light intensity is greater than the set light intensity threshold, the visible light track image is collected by triggering the visible light image collection module; when the light intensity is less than or equal to the set light intensity threshold, the infrared track image is collected by triggering the infrared image collection module, so that better quality image data can be obtained in any external light environment, thereby obtaining accurate anomaly detection results, and whether it is a visible light track image or an infrared track image, a higher resolution and more detailed image data is obtained through super-resolution reconstruction, so that in subsequent track anomaly detection, the track anomaly result detected based on the higher resolution and more detailed image data is more accurate.
[0036] In some embodiments, the target track image is the visible light track image, based on which, in step S130, the target track image is super-resolution reconstructed to obtain a to-be-detected track image with a higher resolution than the target track image, specifically including: inputting the visible light track image into a visible light image super-resolution reconstruction model to obtain a to-be-detected track image with a higher resolution than the visible light track image output by the visible light image super-resolution reconstruction model.
[0037] wherein the visible light image super-resolution reconstruction model is trained based on a sample visible light track image and a label visible light track image corresponding to the sample visible light track image, the resolution of the label visible light track image being higher than that of the sample visible light track image.
[0038] Exemplarily, the visible light image super-resolution reconstruction model can adopt a convolutional neural network model structure of Figure 2 , including a plurality of 3x3 convolutional layers, an adder, and an image reconstruction module (PixelShuffle), the plurality of 3x3 convolutional layers being connected with the adder in cascade, and the adder being connected with the image reconstruction module. The low-resolution visible light track image LR is subjected to feature extraction through the plurality of 3x3 convolutional layers (conv-3x3) to obtain track image features, the track image features are combined with the low-resolution visible light track image LR to obtain combined track image features, and the image reconstruction module reconstructs the combined track image features to obtain a high-resolution track image HR.
[0039] During training of the model, first, a public image super-resolution dataset DIV2K is used to pre-train the model to obtain a basic visual resolution enhancement capability, and then, track image data is used to perform fine-tuning on the model to adapt to the structural features and texture distribution of track images, so as to further improve the actual performance of the visible light image super-resolution reconstruction model in track anomaly detection. During the fine-tuning, high-resolution track image labels are down-sampled to generate low-resolution track image samples to construct sample label pairs for training, an L1 loss function is used as an objective function to optimize the reconstruction capability of the model, and finally, 2, 4, or 8 times or even higher super-resolution images are generated from a single frame of track image, and the specific multiple depends on the resolution multiple of the sample label pair.
[0040] In this embodiment, the super-resolution reconstruction of the visible light track image is realized through the trained visible light image super-resolution reconstruction model.
[0041] In some embodiments, the target track image is the infrared track image, based on which, in step S130, the target track image is super-resolution reconstructed to obtain a to-be-detected track image with a higher resolution than the target track image, including: inputting the infrared track image into an infrared image super-resolution reconstruction model to obtain a to-be-detected track image with a higher resolution than the infrared track image output by the infrared image super-resolution reconstruction model.
[0042] The infrared image super-resolution reconstruction model is trained based on a sample infrared track image and a label infrared track image corresponding to the sample infrared track image, and the resolution of the label infrared track image is higher than that of the sample infrared track image.
[0043] During training of the infrared image super-resolution reconstruction model, a public infrared image dataset (KAISTThermal, FLIR ADAS) is first used for pre-training, and then a sample infrared track image is used for transfer learning to ensure that the model has both generalization and scene adaptability. During the transfer learning process, the high-resolution infrared track image is used as a label, and bicubic downsampling is used to downsample the high-resolution infrared track image to generate a low-resolution sample infrared track image to construct a training sample label pair. During the transfer training, an L1 loss function and a structural similarity loss (SSIM Loss) are used for joint optimization to improve the image reconstruction accuracy and structural preservation capability.
[0044] In this embodiment, the infrared image super-resolution reconstruction model is used to realize super-resolution reconstruction of the infrared track image.
[0045] In some embodiments, as shown in Figure 3 The infrared image super-resolution reconstruction model includes: a shallow feature extraction module configured to extract features from the infrared track image to obtain shallow features.
[0046] a deep feature extraction module configured to extract features based on the shallow features to obtain deep features, and combine the shallow features and the deep features to obtain combined features.
[0047] an image reconstruction module configured to upsample the combined features to obtain the to-be-detected track image after super-resolution reconstruction. Specifically, the image reconstruction module is an upsampling module.
[0048] The infrared track image is a low-resolution image collected by the infrared image collection module, and after feature extraction by the shallow feature extraction module and the deep feature extraction module and merging of the features, the merged feature map is subjected to pixel rearrangement by the image reconstruction module (PixelShuffle), thereby obtaining a high-resolution track image to be detected.
[0049] Specifically, the shallow feature extraction module includes a first convolution module, which extracts shallow features from the low-resolution infrared track image.
[0050] The deep feature extraction module includes a plurality of cascaded second convolution modules, a large kernel attention module, and a third convolution module.
[0051] The first convolution module is configured to input the extracted shallow features into the plurality of cascaded second convolution modules for feature extraction in sequence, thereby obtaining first intermediate features.
[0052] The plurality of cascaded second convolution modules input the first intermediate features into the large kernel attention module, thereby obtaining second intermediate features output by the large kernel attention module.
[0053] The third convolution module extracts features from the second intermediate features, thereby obtaining the deep features.
[0054] Further, the first convolution module, the second convolution module, and the third convolution module are all edge-oriented convolution blocks (ECB). Figure 4 As shown in , a multi-branch structure is used during training, and the multi-branch network is fused into a single-branch network during inference, thereby improving the inference speed while ensuring the accuracy of the inference results.
[0055] Specifically, the infrared image super-resolution reconstruction model performs the super-resolution reconstruction process as follows: Figure 3 As shown in , the shallow feature extraction module is composed of an ECB module ECB-1, and the given input low-resolution image H , where W are the height and width of the low-resolution infrared track image LR, respectively, R represents the real number field. The shallow feature extraction module ECB-1 is represented as ECB (·), which is used to extract shallow features . The shallow feature extraction process is represented as follows: .
[0056] This represents shallow features (i.e., shallow feature maps).
[0057] Then, the shallow features are passed to the deep feature extraction module to extract deeper and more abstract high-level features. This process is expressed as follows: .
[0058] in, This indicates the deep feature extraction module. This represents deep features (i.e., deep feature maps).
[0059] This deep feature extraction module includes multiple cascaded ECBs (ECB-i) and a large kernel attention module (REPLKA), with intermediate features of each ECB-i... Extracted step by step: ; ; .
[0060] in Indicates ECB-i, This indicates a large-core attention module. This refers to the ECB following the large kernel attention module. n This indicates the number of ECBs preceding the large kernel attention module in the deep feature extraction module.
[0061] Subsequently shallow features and deep features After merging, the images are sent to the image reconstruction module to complete super-resolution reconstruction. This process can be described as follows: .
[0062] in, This represents a reconstructed, high-resolution image of the track to be detected. This represents the image reconstruction module.
[0063] like Figure 4 As shown in (a), ECB uses a multi-branch structure during training, such as... Figure 4 As shown in (b), during inference, the multi-branch network is merged into a single-branch network to improve the inference speed while ensuring the accuracy of the inference results.
[0064] Specifically, the ECB module consists of four types of operators: convolution operators (Conv-1×1 and Conv-3×3), horizontal Sobel filters, vertical Sobel filters, and Laplacian filters. These four types of operators form five parallel feature extraction branches.
[0065] The first feature extraction branch: a common 3x3 convolution Conv-3x3 is used to ensure basic performance, and the common convolution is denoted as: .
[0066] wherein, , , and denote the output feature, input feature, weight and bias of Conv-3x3 respectively.
[0067] The second feature extraction branch: an expansion-compression convolution combination, that is, a combination of Conv-1x1 and Conv-3x3. First, a 1x1 convolution Conv-1x1 is used as an expansion convolution to expand the number of channels by two to improve the expression ability, and then a 3x3 convolution Conv-3x3 is used as a compression convolution to restore the number of channels. The wider feature can significantly improve the expression ability and help to obtain better performance on the super-resolution reconstruction task. Denote the {weight, bias} of the 1x1 expansion convolution as K e , B e denote the {weight, bias} of the 3x3 compression convolution as K s , B s and the expansion-compression convolution is shown in the following formula: .
[0068] The third and fourth feature extraction branches: convolution-combination with scaling Sobel filter. In the prior art, edge information has been proved to be helpful for the super-resolution reconstruction task. The ECB incorporates the extraction of the first derivative into the design, and since it is difficult for the model to automatically learn the sharp edge filter, the ECB chooses to use a predefined edge filter and learn the scaling factor of each filter. Specifically, the input feature is first processed by a common 1x1 convolution Conv-1x1, and then two scaled Sobel filters are used to extract the gradient of the intermediate feature. Denote the horizontal and vertical Sobel filters as D x and D y respectively, D x and D y and the formula is as follows: ; .
[0069] Each channel of the intermediate feature is first processed by a Sobel filter and then scaled by a channel scaling factor, the edge information in the horizontal direction and the edge information in the vertical direction is extracted as shown in the following equation: ; .
[0070] wherein, K x , B x} are {weight, bias} of the 1x1 convolution of the horizontal branch (i.e., the third feature extraction branch), K y , B y} are {weight, bias} of the 1x1 convolution of the vertical branch (i.e., the third feature extraction branch). and are the scaling parameter and bias of the horizontal branch, and are the scaling parameter and bias of the vertical branch, the edge information extracted by Sobel filters in the horizontal and vertical directions are directly added to obtain the combined edge information (first-order edge information) as shown in the following equation: .
[0071] Fifth feature extraction branch: Convolution-combined with scaled Laplacian filter. In addition to the first derivative, ECB also extracts the second-order spatial derivative using a laplacian filter. The input feature is first passed through a normal 1x1 convolution Conv-1x1, and then through a laplacian filter D lap to extract the second-order spatial derivative, D lap as shown in the following equation: .
[0072] to extract the scaled second-order edge information as shown in the following equation: .
[0073] wherein, K l , B l} represent {weight, bias} of the 1x1 convolution, S lap represents the scaling factor of D lap ,B lap representing D lap bias.
[0074] the output of the ECB F consists of four parts: .
[0075] The combined feature map is then input into a nonlinear activation layer PReLU.
[0076] In some embodiments, the large kernel attention module comprises a depthwise convolution set layer, a depthwise dilated convolution layer, a target convolution layer, and a multiplier, the depthwise convolution set layer is connected to the depthwise dilated convolution layer, the depthwise dilated convolution layer is connected to the target convolution layer, and the target convolution layer is connected to the multiplier.
[0077] The first intermediate feature is sequentially subjected to feature extraction by the depthwise convolution set layer, the depthwise dilated convolution layer, and the target convolution layer to form a third intermediate feature.
[0078] The multiplier is configured to multiply the first intermediate feature and the third intermediate feature to obtain the second intermediate feature.
[0079] The depthwise convolution set layer comprises an adder and a plurality of first depthwise convolution layers arranged side by side in the model training stage, each of the plurality of first depthwise convolution layers is connected to the adder, and the adder is connected to the depthwise dilated convolution layer; each of the first depthwise convolution layers is configured to extract features from the first intermediate feature, and the adder is configured to add the features extracted by the first depthwise convolution layers and input the added features into the depthwise dilated convolution layer.
[0080] The depthwise convolution set layer comprises a second depthwise convolution layer in the model inference stage, and the second depthwise convolution layer is connected to the depthwise dilated convolution layer.
[0081] The weights of the second depthwise convolution layer are the sum of the weights of the first depthwise convolution layers, and the bias of the second depthwise convolution layer is the sum of the biases of the first depthwise convolution layers.
[0082] Specifically, as shown in (a) of the large kernel attention module in the model training stage comprises an adder and a plurality of first depthwise convolution layers arranged side by side, Figure 5 four 5x5 first depthwise convolution layers (DW) are shown, namely DW1-5x5, DW2-5x5, DW3-5x5, and DW4-5x5. Based on this, the process of the large kernel attention module in the training stage can be represented as: Figure 5 ; ; ; .
[0083] in, X This represents the input feature (i.e., the first intermediate feature). Y This represents the sum of the features extracted by each of the first-depth convolutional layers. Z 1 represents the features extracted from the deep hole convolutional layer DWD-5×5. Z 2 represents the feature extracted by the target convolutional layer Conv-1×1, i.e., the third intermediate feature. Z express Z 2 and X The feature resulting from the multiplication is the second intermediate feature.
[0084] , , and This represents a depthwise convolution with four 5×5 kernels, used to improve the model's expressive power. This represents a depthwise dilated convolution with a kernel size of 5×5 and a dilation rate of 3. This represents a convolution with a kernel size of 1×1. This represents element-wise multiplication in the characteristic matrix.
[0085] like Figure 5 As shown in (b), the deep convolutional ensemble layer includes a second deep convolutional layer DW-5×5 during the model inference stage.
[0086] During inference, the training process can be used... , , and The trained biases and weights are summed, and the summed biases and weights are used as the biases and weights in the DW-5×5 convolution during inference, as shown in the following formula: ; .
[0087] in, , , and They represent , , and The weight, , , and They represent , , and bias, and respectively represent the bias and weight in the DW-5x5 convolution during reasoning.
[0088] In this embodiment, Figure 5 The large kernel attention module with the structure is essentially a REparameterizable Large Kernel Attention (REPLKA) based on the large kernel reparameterization strategy. The large kernel attention module is designed for the super-resolution processing needs of infrared images. Since infrared track images generally have low resolution and few details, the advantages of the large kernel structure in capturing long-distance context information are fully utilized to make up for the lack of texture information in infrared track images. At the same time, through the structure reparameterization technology, the unity between the high expression ability in the training stage and the lightweight and high efficiency in the reasoning stage is realized. In the training stage, the REPLKA enhances the perception ability of the model to low-contrast and fine-grained targets in infrared images by introducing a multi-branch large kernel attention structure. In the reasoning stage, the complex structure is reparameterized into an equivalent single-branch form, effectively reducing the computational amount and model complexity, and improving the reasoning efficiency on edge devices. Therefore, the large kernel attention improves the feature extraction ability of the model without increasing the computational consumption, and improves the feature modeling capability in infrared image processing.
[0089] In some embodiments, before the target track image is super-resolution reconstructed to obtain a to-be-detected track image with a higher resolution than the target track image, the method further includes: removing noise of the target track image by using median filtering and bilateral filtering; and improving image contrast of the target track image after noise removal based on histogram equalization.
[0090] In this embodiment, the target track image is denoised and the contrast is improved to improve the quality of the low-resolution target track image, so as to obtain a higher-quality high-resolution to-be-detected track image after super-resolution reconstruction, and further to make the track anomaly detection more accurate.
[0091] The track anomaly detection device provided by the present application will be described below. The track anomaly detection device described below can be correspondingly referred to the track anomaly detection method described above.
[0092] The track anomaly detection device of the embodiment of the present application, as shown in Figure 6 , includes: The real-time light intensity acquisition unit 610 is configured to acquire the real-time light intensity of the external environment light.
[0093] The target track image determination module 620 is configured to: acquire a visible light track image based on visible light collection when the real-time light intensity is greater than a light intensity threshold, and determine the visible light track image as a target track image; acquire an infrared track image based on infrared light collection when the real-time light intensity is less than or equal to the light intensity threshold, and determine the infrared track image as the target track image.
[0094] The super-resolution reconstruction unit 630 is configured to perform super-resolution reconstruction on the target track image to obtain a to-be-detected track image with a higher resolution than the target track image.
[0095] The anomaly detection unit 640 is configured to perform anomaly detection on a track based on the to-be-detected track image.
[0096] In some embodiments, the target track image is the visible light track image; and the super-resolution reconstruction unit 630 is specifically configured to input the visible light track image into a visible light image super-resolution reconstruction model to obtain a to-be-detected track image with a higher resolution than the visible light track image output by the visible light image super-resolution reconstruction model.
[0097] The visible light image super-resolution reconstruction model is trained based on a sample visible light track image and a label visible light track image corresponding to the sample visible light track image, and the resolution of the label visible light track image is greater than the resolution of the sample visible light track image.
[0098] In some embodiments, the target track image is the infrared track image; and the super-resolution reconstruction unit 630 is specifically configured to input the infrared track image into an infrared image super-resolution reconstruction model to obtain a to-be-detected track image with a higher resolution than the infrared track image output by the infrared image super-resolution reconstruction model.
[0099] The infrared image super-resolution reconstruction model is trained based on a sample infrared track image and a label infrared track image corresponding to the sample infrared track image, and the resolution of the label infrared track image is greater than the resolution of the sample infrared track image.
[0100] In some embodiments, the infrared image super-resolution reconstruction model includes: The shallow feature extraction module is configured to perform feature extraction on the infrared track image to obtain a shallow feature.
[0101] The deep feature extraction module is configured to perform feature extraction based on the shallow feature to obtain a deep feature, and combine the shallow feature and the deep feature to obtain a combined feature.
[0102] The image reconstruction module is configured to perform up-sampling on the combined feature to obtain the to-be-detected track image after super-resolution reconstruction.
[0103] In some embodiments, the shallow feature extraction module comprises a first convolution module, and the deep feature extraction module comprises a plurality of cascaded second convolution modules, a kernel attention module and a third convolution module.
[0104] The first convolution module is configured to input the extracted shallow feature into the plurality of cascaded second convolution modules for feature extraction in sequence to obtain a first intermediate feature.
[0105] The plurality of cascaded second convolution modules input the first intermediate feature into the kernel attention module to obtain a second intermediate feature output by the kernel attention module.
[0106] The third convolution module performs feature extraction on the second intermediate feature to obtain the deep feature.
[0107] In some embodiments, the first convolution module, the second convolution module and the third convolution module are all edge-oriented convolution blocks.
[0108] In some embodiments, the kernel attention module comprises a depth convolution set layer, a depth hole convolution layer, a target convolution layer and a multiplier, the depth convolution set layer is connected to the depth hole convolution layer, the depth hole convolution layer is connected to the target convolution layer, and the target convolution layer is connected to the multiplier. The first intermediate feature is sequentially subjected to feature extraction by the depth convolution set layer, the depth hole convolution layer and the target convolution layer to form a third intermediate feature. The multiplier is configured to multiply the first intermediate feature and the third intermediate feature to obtain the second intermediate feature. In the model training stage, the depth convolution set layer comprises an adder and a plurality of first depth convolution layers arranged in parallel, each of the plurality of first depth convolution layers is connected to the adder, and the adder is connected to the depth hole convolution layer; each first depth convolution layer is configured to perform feature extraction on the first intermediate feature, and the adder inputs the features extracted by the first depth convolution layers into the depth hole convolution layer after adding the features. In the model inference stage, the depth convolution set layer comprises a second depth convolution layer, and the second depth convolution layer is connected to the depth hole convolution layer. The weight of the second depth convolution layer is the sum of the weights of the first depth convolution layers, and the bias of the second depth convolution layer is the sum of the biases of the first depth convolution layers.
[0109] Figure 7 An example of an entity structure diagram of an electronic device is shown in FIG. 1. Figure 7As shown, the electronic device can include a processor 710, a communications interface 720, a memory 730, and a communications bus 740, wherein the processor 710, the communications interface 720, and the memory 730 complete mutual communication through the communications bus 740. The processor 710 can invoke a logic instruction in the memory 730 to execute an orbit anomaly detection method, which includes: Obtaining a real-time light intensity of external ambient light.
[0110] In a case where the real-time light intensity is greater than a light intensity threshold, obtaining a visible light orbit image collected based on visible light, determining that the visible light orbit image is a target orbit image, in a case where the real-time light intensity is less than or equal to the light intensity threshold, obtaining an infrared orbit image collected based on infrared light, and determining that the infrared orbit image is a target orbit image.
[0111] Performing super-resolution reconstruction on the target orbit image to obtain a to-be-detected orbit image with a higher resolution than the target orbit image.
[0112] Performing orbit anomaly detection based on the to-be-detected orbit image.
[0113] In addition, the logic instruction in the memory 730 described above can be implemented in the form of a software function unit and sold or used as an independent product, which can be stored in a computer-readable storage medium. Based on such understanding, the technical solutions of the present application or parts of the present application that essentially contribute to the prior art or the parts of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a plurality of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method described in the various embodiments of the present application. The aforementioned storage medium includes a U disk, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk or an optical disk, and various media that can store program codes.
[0114] On the other hand, the present application also provides a computer program product, which includes a computer program, the computer program can be stored on a non-transitory computer readable storage medium, and the computer program is executed by a processor, so that the computer can execute the orbit anomaly detection method provided by the above-mentioned methods, which includes: Obtaining a real-time light intensity of external ambient light.
[0115] In a case that the real-time light intensity is greater than a light intensity threshold, a visible light track image based on visible light collection is acquired, and the visible light track image is determined as a target track image; in a case that the real-time light intensity is less than or equal to the light intensity threshold, an infrared track image based on infrared light collection is acquired, and the infrared track image is determined as the target track image.
[0116] The target track image is super-resolution reconstructed to obtain a to-be-detected track image with a higher resolution than the target track image.
[0117] An anomaly of a track is detected based on the to-be-detected track image.
[0118] In another aspect, the present application further provides a non-transitory computer readable storage medium, which stores a computer program, and the computer program is executed by a processor to implement the track anomaly detection method provided by the above method, and the method comprises: A real-time light intensity of an external environment light is acquired.
[0119] In a case that the real-time light intensity is greater than a light intensity threshold, a visible light track image based on visible light collection is acquired, and the visible light track image is determined as a target track image; in a case that the real-time light intensity is less than or equal to the light intensity threshold, an infrared track image based on infrared light collection is acquired, and the infrared track image is determined as the target track image.
[0120] The target track image is super-resolution reconstructed to obtain a to-be-detected track image with a higher resolution than the target track image.
[0121] An anomaly of a track is detected based on the to-be-detected track image.
[0122] The device embodiments described above are merely illustrative, wherein the units described as separate components can or can not be physically separate, and the components displayed as units can or can not be physical units, i.e., can be located in one place, or can be distributed on multiple network units. According to actual needs, part or all of the modules can be selected to achieve the purpose of the present embodiment scheme. Those skilled in the art can understand and implement without creative labor.
[0123] Those skilled in the art can clearly understand the technical solutions of the various embodiments from the above description of the embodiments, and the various embodiments can be implemented by means of software with the necessary general hardware platforms, and of course, can also be implemented by hardware. Based on such understanding, the above technical solutions, essentially or in other words, the part of the prior art that makes a contribution, can be embodied in the form of a software product, which can be stored in a computer readable storage medium, such as a ROM / RAM, a magnetic disk, an optical disk, and the like, and includes a number of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0124] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present application, and not to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that: it can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacement for some technical features therein; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A method for detecting track anomalies, characterized in that, include: To obtain the real-time light intensity of the external ambient light; When the real-time light intensity is greater than the light intensity threshold, a visible light trajectory image based on visible light acquisition is acquired, and the visible light trajectory image is determined as the target trajectory image. When the real-time light intensity is less than or equal to the light intensity threshold, an infrared trajectory image based on infrared light acquisition is acquired, and the infrared trajectory image is determined as the target trajectory image. Super-resolution reconstruction is performed on the target orbit image to obtain a target orbit image with a higher resolution than the target orbit image; Anomaly detection is performed on the track based on the image of the track to be detected.
2. The track anomaly detection method according to claim 1, characterized in that, The target orbit image is the visible light orbit image; Super-resolution reconstruction is performed on the target orbit image to obtain a target orbit image with a higher resolution than the target orbit image, including: The visible light orbit image is input into the visible light image super-resolution reconstruction model to obtain a detectable orbit image with a resolution higher than that of the visible light orbit image; The visible light image super-resolution reconstruction model is trained based on sample visible light orbit images and corresponding label visible light orbit images, wherein the resolution of the label visible light orbit images is greater than that of the sample visible light orbit images.
3. The track anomaly detection method according to claim 1, characterized in that, The target orbit image is the infrared orbit image; Super-resolution reconstruction is performed on the target orbit image to obtain a target orbit image with a higher resolution than the target orbit image, including: The infrared orbit image is input into the infrared image super-resolution reconstruction model to obtain a target orbit image with a resolution higher than that of the infrared orbit image; The infrared image super-resolution reconstruction model is trained based on the sample infrared orbit image and the corresponding label infrared orbit image, wherein the resolution of the label infrared orbit image is greater than that of the sample infrared orbit image.
4. The track anomaly detection method according to claim 3, characterized in that, Infrared image super-resolution reconstruction models include: The shallow feature extraction module is used to extract features from infrared orbit images to obtain shallow features; The deep feature extraction module is used to extract features based on the shallow features to obtain deep features, and to merge the shallow features and the deep features to obtain merged features; The image reconstruction module is used to upsample the merged features to obtain the super-reconstructed image of the track to be detected.
5. The track anomaly detection method according to claim 4, characterized in that, The shallow feature extraction module includes a first convolution module, and the deep feature extraction module includes: multiple cascaded second convolution modules, a large kernel attention module, and a third convolution module; The first convolutional module is used to input the extracted shallow features into multiple cascaded second convolutional modules for sequential feature extraction to obtain the first intermediate features; Multiple cascaded second convolutional modules input the first intermediate feature into the large kernel attention module to obtain the second intermediate feature output by the large kernel attention module; The third convolutional module extracts features from the second intermediate features to obtain the deep features.
6. The track anomaly detection method according to claim 5, characterized in that, The first, second, and third convolutional modules are all edge-oriented convolutional blocks.
7. The track anomaly detection method according to claim 5, characterized in that, The large kernel attention module includes: a deep convolutional ensemble layer, a deep dilated convolutional layer, a target convolutional layer, and a multiplier. The deep convolutional ensemble layer is connected to the deep dilated convolutional layer, the deep dilated convolutional layer is connected to the target convolutional layer, and the target convolutional layer is connected to the multiplier. The first intermediate feature is sequentially processed through a deep convolutional ensemble layer, a deep dilated convolutional layer, and a target convolutional layer to extract features and form a third intermediate feature; The multiplier is used to multiply the first intermediate feature and the third intermediate feature to obtain the second intermediate feature; The deep convolutional ensemble layer includes, during the model training phase, an adder and multiple first deep convolutional layers arranged in parallel, each of the multiple first deep convolutional layers being connected to the adder, and the adder being connected to the deep dilated convolutional layer; each first deep convolutional layer is used to extract features from a first intermediate feature, and the adder adds the features extracted by each first deep convolutional layer and inputs the sum into the deep dilated convolutional layer; The deep convolutional ensemble layer includes, during the model inference phase, a second deep convolutional layer, which is connected to the deep dilated convolutional layer. The weights of the second depthwise convolutional layer are the sum of the weights of each of the first depthwise convolutional layers, and the biases of the second depthwise convolutional layer are the sum of the biases of each of the first depthwise convolutional layers.
8. A track anomaly detection device, characterized in that, include: The real-time light intensity acquisition unit is used to acquire the real-time light intensity of the external ambient light. The target orbit image determination module is used to acquire a visible light orbit image based on visible light acquisition when the real-time light intensity is greater than the light intensity threshold, and determine the visible light orbit image as the target orbit image; and to acquire an infrared orbit image based on infrared light acquisition when the real-time light intensity is less than or equal to the light intensity threshold, and determine the infrared orbit image as the target orbit image. The super-resolution reconstruction unit is used to perform super-resolution reconstruction on the target orbit image to obtain a detectable orbit image with a resolution higher than that of the target orbit image; An anomaly detection unit is used to detect anomalies in the track based on the image of the track to be detected.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the orbital anomaly detection method as described in any one of claims 1 to 7.
10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the orbital anomaly detection method as described in any one of claims 1 to 7.