Unsupervised track anomaly detection method and system
Through the unsupervised deep learning method, combined with the region segmentation of key components of the track and the self-generated feature aggregation network, the multivariate Gaussian fitting function and anomaly fraction measurement function are used to achieve high accuracy and efficiency of track anomaly detection, solving the problems of limited detection content and low robustness in the existing technology.
Patent Information
- Application Number
- CN202510048442.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-13
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2045-01-13
AI Technical Summary
The prior art has problems in orbit abnormality detection, low robustness and the need for manual labeling of data in the detection of orbital anomalies. Especially when facing complex and changing scenarios, it is difficult to achieve high-precision and efficient detection.
Unsupervised deep learning method is adopted to determine the track anomaly image and position the abnormal pixels through the area segmentation module of the key components of the track, the self-generated feature aggregation network, the multivariate Gaussian fitting function and the internal and external equalization anomaly fraction measurement function.
It realizes accurate detection and positioning of unknown foreign objects on the track, improves detection accuracy, reduces labor costs, and solves the problem of low detection accuracy caused by small samples.
Smart Images

Figure CN120047726A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of track supervision based on computer vision, and particularly to an unsupervised track anomaly detection method and system. Background Art
[0002] Tracks are an indispensable part of rail transit line equipment, guiding and carrying the operation of trains, and are very important for the safe operation of trains. Semi-open tracks are constantly subjected to natural erosion and loads throughout the year, and key components of the tracks are prone to various types of anomalies. In addition, static foreign objects with random positions, different sizes and shapes often appear on tracks in the natural open-air environment. Such anomalies with random occurrence probabilities, unfixed positions and unknown categories pose potential threats to the stable operation of trains. With the continuous extension of the rail transit operation scope, the increasing accumulation of operation years and the continuous rise of labor costs, the limitations of the traditional manual inspection method for the above-mentioned anomaly detection content are becoming increasingly prominent, which is not only time-consuming and laborious, but also has poor detection accuracy. Therefore, in the face of the on-site requirements of track inspection with more comprehensive detection content, higher detection accuracy and greater intelligence, it is very necessary to study intelligent and accurate detection of unknown track anomalies.
[0003] Currently, supervised anomaly detection methods based on video images focus on detecting abnormal defects of key track components of known categories, and have good application effects in the field, but it is obvious that the inspection content is limited because the detectable categories are restricted by training labels. Unsupervised traditional machine vision processing methods mainly rely on the fixity and regularity of the positions of key track components for research. However, this dependence results in relatively low robustness when facing complex and changeable scenarios in actual applications. In addition, there are many types of these methods, and for different image scene requirements, the selection and combination of methods need to be flexibly adjusted, which undoubtedly increases the difficulty in actual applications. Summary of the Invention
[0004] The purpose of the present invention is to provide an unsupervised track anomaly detection method and system to solve at least one of the technical problems existing in the above background art.
[0005] To achieve the above purpose, the present invention adopts the following technical solutions:
[0006] In a first aspect, the present invention provides an unsupervised track anomaly detection method, including:
[0007] Obtain an overhead view track image to be detected;
[0008] The obtained overhead-view track image to be detected is processed using a pre-trained detection model to obtain a detection result. The pre-trained detection model includes a key component region segmentation unit, a feature aggregation network, a fitting network, and a calculation unit. The key component region segmentation unit is used for pixel-level segmentation and background simplification of the track image to generate three key component pseudo-images of the rail, tie, and track slab. The feature aggregation network is used to extract and aggregate multi-scale features in the pseudo-images to obtain a high-dimensional feature vector matrix. The fitting unit is used to fit the high-dimensional feature vector matrix using a multivariate Gaussian fitting function to obtain the optimal parameters of the normal image. The calculation unit is used to calculate the difference between the optimal parameters of the normal image and the high-dimensional feature vector matrix of the test image using an internal and external balance anomaly score metric function, calculate the anomaly score, and achieve anomaly image discrimination and anomaly pixel localization.
[0009] As a further limitation of the first aspect of the present invention, training the region segmentation unit of the track key components includes: constructing a track key component segmentation data set, with specific categories including three categories of rail, tie, and track slab; using a semantic segmentation algorithm and training with the data set to achieve pixel segmentation of the three track key components, obtaining the mask Mask of the track image and the color palette corresponding to each of the three key components component , where component is the track key component; using the mask pixel matching formula to obtain the pseudo-images of the rail, tie, and track slab respectively.
[0010] As a further limitation of the first aspect of the present invention, when training the detection model, the ResNet18 network is used as the backbone network; the first layer Layer 1 , the second layer Layer 2 and the third layer Layer 3 feature maps are used as the basic feature maps; using the transposed convolution upsampling method to unify the resolutions of the three feature maps and performing a splicing operation to obtain a preliminary vector matrix Block 1 ; based on the weighted self-generated feature mechanism, using the above Layer 1 , Layer 2 and Layer 3 feature maps to generate the fourth layer feature map Layer 4 , and performing a splicing operation with the preliminary vector matrix Block 1 to obtain a high-dimensional feature vector matrix Block 2 ; selecting the first 500 layers of the channels of Block 2 as the final feature vector matrix Block 3 .
[0011] As a further limitation of the first aspect of the present invention, the empowered self-generated feature mechanism includes: selecting the upsampled Layer 2 , Layer 3 The first 64 layers of the feature map channels, and assigning a weight value of 0.4, and adding it to the weight value of 0.2 assigned to the selected Layer 1 feature map to obtain the fourth layer feature map Layer 4 .
[0012] As a further limitation of the first aspect of the present invention, the multivariate Gaussian fitting function includes: taking n normal training images as a group, and dividing the input image into a pixel block grid, and the high-dimensional feature vector matrix of the normal pixel block at (i, j) x ij is the feature vector value at (i, j) of the normal image; using the multivariate Gaussian distribution N(μ ij , Σ ij ) to fit the high-dimensional feature vector matrix to obtain the optimal parameters μ ij and Σ ij .
[0013] As a further limitation of the first aspect of the present invention, the internal and external balanced anomaly score metric function includes: designing an external square root metric function D 1 , calculating the difference score between the test image and the optimal parameters, where x ij ’ is the feature vector value at (i, j) of the test image;
[0014]
[0015] Using the internal Mahalanobis distance metric function D 2 , calculating the difference score inside the test image;
[0016]
[0017] The internal and external balanced anomaly score metric function D is the summation calculation of the two weighted anomaly scores to obtain the final anomaly score: D = 0.8×D 1 +0.2×D 2 .
[0018] In the second aspect, the present invention provides an unsupervised track anomaly detection system, including:
[0019] An acquisition module for acquiring an overhead view track image to be detected;
[0020] The detection module is used to process the acquired top-down perspective track image by using a pre-trained detection model to obtain a detection result. The pre-trained detection model includes a key component area segmentation unit, a feature aggregation network, a fitting network, and a calculation unit. The key component area segmentation unit is used to perform pixel-level segmentation and background simplification on the track image to generate three key component pseudo-images of the rail, sleeper, and track slab. The feature aggregation network is used to extract and aggregate multi-scale features in the pseudo-image to obtain a high-dimensional feature vector matrix. The fitting unit is used to fit the high-dimensional feature vector matrix by using a multivariate Gaussian fitting function to obtain the optimal parameters of the normal image. The calculation unit is used to calculate the difference between the optimal parameters of the normal image and the high-dimensional feature vector matrix of the test image by using an internal and external balanced anomaly score metric function, calculate the anomaly score, and realize anomaly image discrimination and anomaly pixel positioning.
[0021] In a third aspect, the present invention provides a non-transitory computer-readable storage medium for storing computer instructions, which when executed by a processor, implement the unsupervised track anomaly detection method as described in the first aspect.
[0022] In a fourth aspect, the present invention provides a computer device, including a memory and a processor, the processor and the memory communicate with each other, the memory stores program instructions executable by the processor, and the processor calls the program instructions to execute the unsupervised track anomaly detection method as described in the first aspect.
[0023] In a fifth aspect, the present invention provides an electronic device, including: a processor, a memory, and a computer program; wherein, the processor is connected to the memory, the computer program is stored in the memory, and when the electronic device runs, the processor executes the computer program stored in the memory so that the electronic device executes the instructions for implementing the unsupervised track anomaly detection method as described in the first aspect.
[0024] Advantages of the present invention: By utilizing the unique spatial position distribution law of the track scene, before detecting anomalies in the entire track image, the track key component area segmentation module is used to reduce the interference of complex background factors on the detection of unknown foreign objects, preparing for subsequent anomaly detection; an unsupervised deep learning method is used to construct a self-generated feature aggregation network, which extracts and aggregates a high-dimensional feature vector matrix with semantic information and local information, and realizes optimal parameter fitting through a multivariate Gaussian fitting function. Then, an anomaly score metric function with internal and external balance is used to accurately realize the discrimination of track anomaly images and the positioning of anomaly pixels; the detection accuracy is high, and it does not require manually labeled data for training. Only by using normal track data can the accurate detection and positioning of unknown foreign objects on the track be realized, which well supplements the track inspection content, greatly reduces the labor cost, and solves the problem of low anomaly detection accuracy caused by few track anomaly samples.
[0025] The advantages of the additional aspects of the present invention will be more clearly given in the following description part, or can be understood through the practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for the description of the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0027] Figure 1 Flowchart of the unsupervised track anomaly detection method according to the embodiment of the present invention
[0028] Figure 2 Effect diagram of the detection of the unsupervised track anomaly detection method according to the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0029] The embodiments of the present invention will be described in detail below. The examples of the embodiments are shown in the drawings, where the same or similar reference numerals represent the same or similar elements or elements with the same or similar functions throughout. The embodiments described below with reference to the drawings are exemplary only for explaining the present invention and should not be construed as limiting the present invention.
[0030] Those skilled in the art of the present technology can understand that unless otherwise defined, all terms (including technical terms and scientific terms) used herein have the same meaning as the general understanding of those of ordinary skill in the art in the field to which the present invention belongs.
[0031] It should also be understood that terms such as those defined in a general dictionary should be understood as having a meaning consistent with their meaning in the context of the prior art, and will not be interpreted in an idealized or overly formal sense unless defined as here.
[0032] Those skilled in the art of this technology can understand that unless specifically stated otherwise, the singular forms "a", "an", "the", and "said" used herein may also include the plural forms. It should be further understood that the term "comprising" used in the description of the present invention means the presence of the described features, integers, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, and / or their groups.
[0033] In the description of this specification, the description with reference to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples. Without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.
[0034] To facilitate the understanding of the present invention, the following will further explain the present invention with specific embodiments in conjunction with the accompanying drawings, and the specific embodiments do not constitute a limitation to the embodiments of the present invention.
[0035] Those skilled in the art should understand that the drawings are only schematic diagrams of the embodiments, and the components in the drawings are not necessarily essential for implementing the present invention.
[0036] An unsupervised track anomaly detection method provided by the present invention uses an on-board camera to photograph the track to obtain a track image from a top-down perspective. The background of the track image is simplified through a track key component area segmentation module to generate a track key component pseudo-image. Then, a self-generated feature aggregation network and a multivariate Gaussian fitting function are used to achieve feature parameterization, obtaining the optimal parameters of the normal image and the high-dimensional feature vector matrix of the test image. Finally, an internal and external balanced anomaly score metric function is used to achieve high-precision detection of unknown foreign objects on the track.
[0037] Embodiment 1
[0038] In this Embodiment 1, first, an unsupervised track anomaly detection system is provided, including: an acquisition module for acquiring an overhead view track image to be detected; a detection module for processing the acquired overhead view track image to be detected by using a pre-trained detection model to obtain a detection result; wherein, the pre-trained detection model includes a key component region segmentation unit, a feature aggregation network, a fitting network, and a calculation unit; the key component region segmentation unit is used for pixel-level segmentation and background simplification of the track image to generate three key component pseudo-images of the rail, the sleeper, and the track slab; the feature aggregation network is used for extracting and aggregating multi-scale features in the pseudo-images to obtain a high-dimensional feature vector matrix; the fitting unit is used for fitting the high-dimensional feature vector matrix by using a multivariate Gaussian fitting function to obtain the optimal parameters of a normal image; the calculation unit is used for adopting an internal and external balance anomaly score metric function to calculate the difference between the optimal parameters of the normal image and the high-dimensional feature vector matrix of the test image, calculate the anomaly score, and realize anomaly image discrimination and anomaly pixel positioning.
[0039] In this embodiment, the above system is used to implement an unsupervised track anomaly detection method, including: using the acquisition module to acquire an overhead view track image to be detected; for example, using an unmanned aerial vehicle (UAV)-mounted track image acquisition device to acquire an overhead view track image. Using the detection module to process the acquired overhead view track image to be detected by using a pre-trained detection model to obtain a detection result; wherein, the pre-trained detection model includes a key component region segmentation unit, a feature aggregation network, a fitting network, and a calculation unit; the key component region segmentation unit is used for pixel-level segmentation and background simplification of the track image to generate three key component pseudo-images of the rail, the sleeper, and the track slab; the feature aggregation network is used for extracting and aggregating multi-scale features in the pseudo-images to obtain a high-dimensional feature vector matrix; the fitting unit is used for fitting the high-dimensional feature vector matrix by using a multivariate Gaussian fitting function to obtain the optimal parameters of a normal image; the calculation unit is used for adopting an internal and external balance anomaly score metric function to calculate the difference between the optimal parameters of the normal image and the high-dimensional feature vector matrix of the test image, calculate the anomaly score, and realize anomaly image discrimination and anomaly pixel positioning.
[0040] In this embodiment, training the detection model specifically includes the following steps:
[0041] The UAV-mounted track image acquisition device acquires an overhead view track image;
[0042] The track key component area segmentation module (key component area segmentation unit) is used to achieve pixel-level segmentation and background simplification, generating pseudo-images of three key components: the rail, the sleeper, and the track slab. The self-generated feature aggregation network is used to extract and aggregate multi-scale features in the pseudo-images, obtaining a high-dimensional feature vector matrix. The fitting unit uses a multivariate Gaussian fitting function to fit the high-dimensional feature vector matrix to obtain the optimal parameters of the normal image. An abnormal score metric function with internal and external balance is designed to calculate the difference between the optimal parameters of the normal image and the high-dimensional feature vector matrix of the test image, and calculate the abnormal score to achieve abnormal image discrimination and abnormal pixel positioning.
[0043] The track key component area segmentation module includes:
[0044] Construct a track key component segmentation dataset, with specific categories including three categories: the rail, the sleeper, and the track slab;
[0045] Use a semantic segmentation algorithm to achieve pixel segmentation of the three track key components, obtaining the mask Mask of the track image and the color palette corresponding to each of the three key components component , where component is the track key component, rail is the rail, sleeper is the sleeper, and railplate is the track slab;
[0046]
[0047] Use the mask pixel matching formula to obtain the pseudo-images of the rail, the sleeper, and the track slab respectively. Where Image(i,j) is the pixel value of the track image at (i,j), and Image component (i,j) is the pixel value of the key component pseudo-image at (i,j).
[0048]
[0049] The self-generated feature aggregation network includes:
[0050] Use the ResNet18 network as the backbone network;
[0051] Select the first layer Layer 1 , the second layer Layer 2 and the third layer Layer 3 feature maps as the basic feature maps, and the resolutions of the feature maps are 32×32, 16×16, and 8×8 respectively;
[0052] Use the transposed convolution upsampling method to unify the resolutions of the three feature maps to 32×32, and perform a concatenation operation to obtain the preliminary vector matrix Block 1 , where TC is the transposed convolution operation and torch.cat is the concatenation operation;
[0053] Block 1 = torch.cat[Layer 1 , TC(Layer 2 ), TC(Layer 3 )]
[0054] Design a weighted self-generated feature mechanism. Using the above Layer 1 , Layer 2 and Layer 3 feature maps to generate the fourth-layer feature map Layer 4 , and perform a concatenation operation with the preliminary vector matrix Block 1 to obtain a high-dimensional feature vector matrix Block 2 ;
[0055] Block 2 = torch.cat(Block 1 , Layer 4 )
[0056] To reduce the computational load, select the first 500 layers of the channels of Block 2 as the final feature vector matrix Block 3 .
[0057] The weighted self-generated feature mechanism includes:
[0058] Select the first 64 layers of the channels of the upsampled Layer 2 , Layer 3 feature maps and assign a weight value of 0.4;
[0059] Select the Layer 1 feature map and assign a weight value of 0.2, and add it to the above two items to obtain the fourth-layer feature map Layer 4 .
[0060] Layer 4 = 0.2 × Layer 1 + 0.4 × TC(Layer 2 )(64)+ 0.4 × TC(Layer 3 )(64)
[0061] The multivariate Gaussian fitting function includes:
[0062] Take n normal training images as a group, and divide the input image into a pixel block grid. The high-dimensional feature vector matrix xij of the normal pixel block at (i, j) is the feature vector value at (i, j) of the normal image;
[0063] Using the multivariate Gaussian distribution N(μ ij , Σ ij ) to fit the high-dimensional feature vector matrix to obtain the optimal parameters μ ij and Σ ij .
[0064] The outlier score metric function for internal and external balance includes:
[0065] Design the external square root metric function D 1 , calculate the difference score between the test image and the optimal parameters, where x ij ’ is the feature vector value at the test image (i, j);
[0066]
[0067] Use the internal Mahalanobis distance metric function D 2 to calculate the difference score inside the test image;
[0068]
[0069] The outlier score metric function D for internal and external balance is the sum calculation of the two weighted outlier scores to obtain the final outlier score.
[0070] D = 0.8×D 1 + 0.2×D 2
[0071] Example 2
[0072] In this Example 2, an unsupervised rail anomaly detection method is provided. The rail is photographed by a mounted camera to obtain a top-down view of the rail image. The background of the rail image is simplified by the rail key component area segmentation module to generate a pseudo-image of the rail key components. Then, the self-generated feature aggregation network and the multivariate Gaussian fitting function are used to realize feature parameterization, obtaining the optimal parameters of the normal image and the high-dimensional feature vector matrix of the test image. Finally, the outlier score metric function for internal and external balance is used to realize high-precision detection of unknown foreign objects on the rail.
[0073] As Figure 1 shown, the unsupervised rail anomaly detection method includes the following steps:
[0074] Step 1, the mounted rail image acquisition device acquires a top-down view of the rail image, and the image basically includes key components such as rails, sleepers, and track slabs;
[0075] Step 2: Use the track key component area segmentation module to achieve pixel-level segmentation and background simplification, generating three pseudo-images of the key components: the rail, the sleeper, and the track slab. Each pseudo-image only includes one key component, and the background pixel value is 0;
[0076] The specific steps of Step 2 are as follows:
[0077] Step 2.1: Construct a track key component segmentation dataset with specific categories including the rail, the sleeper, and the track slab, and use the Labelme software for pixel-level annotation;
[0078] Step 2.2: Use a semantic segmentation algorithm to train and validate the track key component segmentation dataset to achieve pixel segmentation of the three track key components, obtaining the mask Mask of the track image and the corresponding color palette palette for each of the three key components component , where component is the track key component, rail is the rail, sleeper is the sleeper, and railplate is the track slab;
[0079]
[0080] Step 2.3: Use the mask pixel matching formula, with the foreground being a single key component and the background pixel value being 0, to obtain the pseudo-images of the rail, the sleeper, and the track slab respectively. Among them, Image(i,j) is the pixel value of the track image at (i,j), and Image component (i,j) is the pixel value of the key component pseudo-image at (i,j);
[0081]
[0082] Step 3: Use the self-generated feature aggregation network to extract and aggregate multi-scale features in the pseudo-images and perform parameterization to obtain a high-dimensional feature vector matrix;
[0083] The specific steps of Step 3 are as follows:
[0084] Step 3.1: Use the ResNet18 network pre-trained on the ImageNet dataset as the backbone network;
[0085] Step 3.2: Select the feature maps of the first layer Layer 1 , the second layer Layer 2 and the third layer Layer 3 as the base feature maps. The resolutions of the feature maps are 32×32, 16×16, and 8×8 respectively, and the number of channels are 64, 128, and 256 respectively;
[0086] Step 3.3, using the transposed convolution upsampling method, unify the resolutions of the three feature maps to 32×32, and perform a concatenation operation to obtain the preliminary vector matrix Block 1 , where TC is the transposed convolution operation and torch.cat is the concatenation operation;
[0087] Block 1 = torch.cat[Layer 1 , TC(Layer 2 ), TC(Layer 3 )]
[0088] Step 3.4, design a weighted self-generated feature mechanism, select the first 64 layers of the channels of the upsampled Layer 2 and Layer 3 feature maps, and assign a weight of 0.4;
[0089] Step 3.5, select the Layer 1 feature map and assign it a weight of 0.2, add it to the above two items to obtain the fourth layer feature map Layer 4 ;
[0090] Layer 4 = 0.2×Layer 1 + 0.4×TC(Layer 2 )(64)+ 0.4×TC(Layer 3 )(64)
[0091] Step 3.6, the fourth layer feature map Layer 4 , and perform a concatenation operation with the preliminary vector matrix Block 1 to obtain the high-dimensional feature vector matrix Block 2 ;
[0092] Block 2 = torch.cat(Block 1 , Layer 4 )
[0093] Step 3.7, to reduce the computational amount, select the first 500 layers of the channels of Block 2 as the final feature vector matrix Block 3 ;
[0094] Step 4, use the multivariate Gaussian fitting function to fit the high-dimensional feature vector matrix to obtain the optimal parameters of the normal image;
[0095] The specific steps of the said Step 4 include the following steps:
[0096] Step 4.1: Take n normal training images as a group, and divide the input image into a grid of pixel blocks. The high-dimensional feature vector matrix of the normal pixel block located at (i, j) x ij is the feature vector value at (i, j) of the normal image;
[0097] Step 4.2: Use the multivariate Gaussian distribution N(μ ij , Σ ij ) to fit the high-dimensional feature vector matrix to obtain the optimal parameters μ ij and Σ ij ;
[0098]
[0099] Step 5: Design an abnormal score metric function with internal and external balance, calculate the difference between the optimal parameters of the normal image and the high-dimensional feature vector matrix of the test image, and calculate the abnormal score to achieve abnormal image discrimination and abnormal pixel localization;
[0100] The specific steps of Step 5 are as follows:
[0101] Step 5.1: Design an external square root metric function D 1 , focusing on the absolute difference between the test image and the optimal parameters, and calculate the difference score between the test image and the optimal parameters, where x ij ’ is the feature vector value at (i, j) of the test image;
[0102]
[0103] Step 5.2: Use the internal Mahalanobis distance metric function D 2 , focusing on the correlation between the internal pixels of the test image, and calculate the difference score of the internal pixel distribution of the test image;
[0104]
[0105] Step 5.3: The abnormal score metric function D with internal and external balance is the summation calculation of the two weighted abnormal scores to obtain the final abnormal score, which is used as the basis for the final abnormal image discrimination and abnormal pixel localization.
[0106] D = 0.8×D 1 + 0.2×D 2
[0107] Example 3
[0108] This Example 3 provides an unsupervised track anomaly detection method, including the following steps:
[0109] Step 1: The on-board orbital image acquisition device acquires orbital images from an aerial perspective. The images basically include key components such as steel rails, sleepers, and track slabs;
[0110] Step 2: Use the orbital key component area segmentation module to achieve pixel-level segmentation and background simplification, generating three pseudo-images of key components: steel rails, sleepers, and track slabs. Each pseudo-image only includes one key component, and the background pixel value is 0;
[0111] The specific steps of Step 2 are as follows:
[0112] Step 2.1: Construct an orbital key component segmentation dataset, with specific categories including three categories: steel rails, sleepers, and track slabs, and use the Labelme software for pixel-level annotation;
[0113] Step 2.2: Use a semantic segmentation algorithm to train and validate the orbital key component segmentation dataset to obtain the optimal weights, achieve pixel segmentation of the three orbital key components, and obtain the mask Mask of the orbital image and the corresponding color palette palette for each of the three key components component , where component is the orbital key component, rail is the steel rail, sleeper is the sleeper, and railplate is the track slab;
[0114]
[0115] Step 2.3: Use the mask pixel matching formula, with the foreground being a single key component and the background pixel value being 0, to obtain pseudo-images of steel rails, sleepers, and track slabs respectively. Among them, Image(i,j) is the pixel value of the orbital image at (i,j), and Image component (i,j) is the pixel value of the key component pseudo-image at (i,j);
[0116]
[0117] Step 3: Use the self-generated feature aggregation network to extract and aggregate multi-scale features in the pseudo-images and perform parameterization to obtain a high-dimensional feature vector matrix;
[0118] The specific steps of Step 3 are as follows:
[0119] Step 3.1: Use the ResNet18 network pre-trained on the ImageNet dataset as the backbone network. This network consists of 5 stages. Stage 0 is the convolutional layer and pooling layer, and stages 1 to 4 are all residual blocks with different numbers of channels;
[0120] Step 3.2: Select the first layer Layer of the backbone network 1 and the second layer Layer 2and the third layer Layer 3 The feature maps are basic feature maps, with resolutions of 32×32, 16×16, and 8×8 respectively, and the number of channels are 64, 128, and 256 respectively;
[0121] Step 3.3, using the transposed convolution upsampling method, unify the resolutions of the three feature maps to 32×32. For Layer 2 Use a 17×17 convolutional kernel with a stride of 1 for Layer 3 Use an 18×18 convolutional kernel with a stride of 2 to generate Layer 2 ’ and Layer 3 ’ respectively, and perform a concatenation operation to obtain the preliminary vector matrix Block 1 , where TC is the transposed convolution operation and torch.cat is the concatenation operation;
[0122] Layer 2 ′ = TC(Layer 2 , k = [17, 17], s = [1, 1], p = [0, 0])
[0123] Layer 3 ′ = TC(Layer 3 , k = [18, 18], s = [2, 2], p = [0, 0])
[0124] Block 1 = torch.cat[Layer 1 , Layer 2 ′, Layer 3 ′]
[0125] Step 3.4, design a weighted self-generated feature mechanism, select the first 64 layers of the channels of the feature maps of Layer 2 ’ and Layer 3 ’, and assign a weight of 0.4;
[0126] Step 3.5, select the feature map of Layer 1 and assign it a weight of 0.2, add it to the above two items to obtain the fourth layer feature map Layer 4 , with the number of channels being 64;
[0127] Layer 4 = 0.2×Layer 1 + 0.4×Layer 2 ′(64) + 0.4×Layer 3 ′(64)
[0128] Step 3.6, the fourth layer feature map Layer4 , and perform a splicing operation with the preliminary vector matrix Block 1 to obtain a high-dimensional feature vector matrix Block 2 , with the number of channels being 512;
[0129] Block 2 = torch.cat(Block 1 , Layer 4 )
[0130] Step 3.7, to reduce the computational amount, select the first 500 layers of the channels of Block 2 as the final feature vector matrix Block 3 ;
[0131] Step 4, use the multivariate Gaussian fitting function to fit the high-dimensional feature vector matrix to obtain the optimal parameters of the normal image;
[0132] The specific steps of the said Step 4 include the following steps:
[0133] Step 4.1, take n normal training images as a group, and divide the input image into a pixel block grid. The high-dimensional feature vector matrix x ij of the normal pixel block at (i, j) is the feature vector value at (i, j) of the normal image;
[0134] Step 4.2, use the multivariate Gaussian distribution N(μ ij , Σ ij ) to fit the high-dimensional feature vector matrix to obtain the optimal parameters μ ij and Σ ij , f z (x) is the probability density function of the multivariate Gaussian distribution. Z represents a random vector composed of n random variables. The mean vector of the random vector is μ, the covariance matrix of the random vector is ∑, and det is to take the determinant;
[0135]
[0136]
[0137] Step 5, design an abnormal score measurement function with internal and external balance, calculate the difference between the optimal parameters of the normal image and the high-dimensional feature vector matrix of the test image, and calculate the abnormal score to achieve abnormal image discrimination and abnormal pixel positioning;
[0138] The specific steps of the said Step 5 include the following steps:
[0139] Step 5.1, design an external square root measurement function D 1, focus on the absolute difference between the test image and the optimal parameters, and calculate the difference score between the test image and the optimal parameters, where x ij ’ is the feature vector value at the position (i, j) of the test image;
[0140]
[0141] Step 5.2, use the internal Mahalanobis distance metric function D 2 , focus on the correlation within the test image, and calculate the difference score of the pixel distribution within the test image;
[0142]
[0143] Step 5.3, the internal and external balanced anomaly score metric function D is the summation calculation of the two weighted anomaly scores, obtaining the final anomaly score, which is used as the basis for the final anomaly image discrimination and anomaly pixel localization.
[0144] D = 0.8×D 1 + 0.2×D 2
[0145] As Figure 2 shown, the unsupervised track anomaly detection algorithm proposed in this embodiment can achieve a detection accuracy of 97.82%. The innovation of this method is mainly reflected in the spatial distribution law of the specific track scenario. By constructing and using four steps: the track key component area segmentation module, self-generated feature aggregation, multivariate Gaussian distribution function, and internal and external balanced anomaly score metric function, using the deep learning-based method, it reduces the interference of complex backgrounds on anomaly detection, significantly improves the accuracy of track anomaly detection, supplements the content of track unknown foreign object inspection, and provides strong technical support for the inspection and maintenance of track status.
[0146] Example 4
[0147] Embodiment 4 provides a non-transitory computer-readable storage medium for storing computer instructions. When the computer instructions are executed by a processor, the unsupervised track anomaly detection method described above is implemented. The method includes: obtaining an overhead view track image to be detected; using a pre-trained detection model to process the obtained overhead view track image to be detected to obtain a detection result. Among them, the pre-trained detection model includes a key component region segmentation unit, a feature aggregation network, a fitting network, and a calculation unit. The key component region segmentation unit is used for pixel-level segmentation and background simplification of the track image to generate three key component pseudo-images of the rail, the sleeper, and the track slab. The feature aggregation network is used for extracting and aggregating multi-scale features in the pseudo-images to obtain a high-dimensional feature vector matrix. The fitting unit is used for using a multivariate Gaussian fitting function to fit the high-dimensional feature vector matrix to obtain the optimal parameters of the normal image. The calculation unit is used for adopting an internal and external balanced anomaly score metric function to calculate the difference between the optimal parameters of the normal image and the high-dimensional feature vector matrix of the test image, calculate the anomaly score, and implement anomaly image discrimination and anomaly pixel localization.
[0148] Embodiment 5
[0149] Embodiment 5 provides a computer device, including a memory and a processor. The processor and the memory communicate with each other. The memory stores program instructions executable by the processor. The processor calls the program instructions to execute the unsupervised track anomaly detection method described above. The method includes: obtaining an overhead view track image to be detected; using a pre-trained detection model to process the obtained overhead view track image to be detected to obtain a detection result. Among them, the pre-trained detection model includes a key component region segmentation unit, a feature aggregation network, a fitting network, and a calculation unit. The key component region segmentation unit is used for pixel-level segmentation and background simplification of the track image to generate three key component pseudo-images of the rail, the sleeper, and the track slab. The feature aggregation network is used for extracting and aggregating multi-scale features in the pseudo-images to obtain a high-dimensional feature vector matrix. The fitting unit is used for using a multivariate Gaussian fitting function to fit the high-dimensional feature vector matrix to obtain the optimal parameters of the normal image. The calculation unit is used for adopting an internal and external balanced anomaly score metric function to calculate the difference between the optimal parameters of the normal image and the high-dimensional feature vector matrix of the test image, calculate the anomaly score, and implement anomaly image discrimination and anomaly pixel localization.
[0150] Embodiment 6
[0151] Embodiment 6 of the present invention provides an electronic device, including: a processor, a memory, and a computer program; wherein, the processor is connected to the memory, the computer program is stored in the memory, and when the electronic device runs, the processor executes the computer program stored in the memory, so that the electronic device executes instructions for implementing the unsupervised track anomaly detection method described above. The method includes: obtaining a track image from a top-down perspective to be detected; processing the obtained track image from a top-down perspective to be detected by using a pre-trained detection model to obtain a detection result; wherein, the pre-trained detection model includes a key component area segmentation unit, a feature aggregation network, a fitting network, and a calculation unit; the key component area segmentation unit is used for pixel-level segmentation and background simplification of the track image to generate three key component pseudo-images of the rail, the sleeper, and the track slab; the feature aggregation network is used for extracting and aggregating multi-scale features in the pseudo-images to obtain a high-dimensional feature vector matrix; the fitting unit is used for fitting the high-dimensional feature vector matrix by using a multivariate Gaussian fitting function to obtain the optimal parameters of the normal image; the calculation unit is used for adopting an internal and external balance anomaly score metric function to calculate the difference between the optimal parameters of the normal image and the high-dimensional feature vector matrix of the test image, calculate the anomaly score, and realize anomaly image discrimination and anomaly pixel positioning.
[0152] Those skilled in the art should understand that the embodiments of the present invention may be provided as a method, a system, or a computer program product. Therefore, the present invention may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memory, CD-ROM, optical memory, etc.) containing computer-usable program code.
[0153] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, and the combination of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in one Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0154] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing apparatus to function in a particular manner, such that the instructions stored in the computer-readable memory produce an article of manufacture including instruction means that implement the function specified in one or more of the blocks and / or steps of the flowchart. Figure 1 one or more of the steps and / or blocks Figure 1 specified in the flowchart.
[0155] These computer program instructions may also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process, thereby providing steps for implementing the function specified in one or more of the steps and / or blocks of the flowchart. Figure 1 one or more of the steps and / or blocks Figure 1 specified in the flowchart.
[0156] Although the specific embodiments of the present invention have been described in conjunction with the accompanying drawings, such description is not intended to limit the scope of the present invention. Those skilled in the art should understand that various modifications and variations can be made without departing from the spirit and scope of the present invention, which is defined by the appended claims.
Claims
1. An unsupervised track anomaly detection method, characterized in that: include: Acquire a track image of a top-down perspective to be detected; The acquired track image of the top-down perspective to be detected is processed by using a pre-trained detection model to obtain a detection result; wherein the pre-trained detection model includes a key component region segmentation unit, a feature aggregation network, a fitting network and a calculation unit; the key component region segmentation unit is used to perform pixel-level segmentation and background simplification of the track image to generate pseudo images of three key components: rails, sleepers and track plates; the feature aggregation network is used to extract and aggregate multi-scale features in the pseudo image to obtain a high-dimensional feature vector matrix; the fitting unit is used to fit the high-dimensional feature vector matrix using a multivariate Gaussian fitting function to obtain the optimal parameters of the normal image; the calculation unit is used to use an internally and externally balanced anomaly score measurement function to calculate the difference between the optimal parameters of the normal image and the high-dimensional feature vector matrix of the test image, calculate the anomaly score, and realize abnormal image discrimination and abnormal pixel positioning.
2. The unsupervised track anomaly detection method according to claim 1, characterized in that: Training the regional segmentation unit of key track components includes: constructing a key track component segmentation dataset, including three categories: rails, sleepers and track plates; using a semantic segmentation algorithm and training the dataset to achieve pixel segmentation of the three key track components, and obtaining the mask of the track image and the color palette corresponding to the three key components. component , component is the key component of the track; using the mask pixel matching formula, the pseudo images of rails, sleepers and track plates are obtained respectively.
3. The unsupervised track anomaly detection method according to claim 1, characterized in that: The detection model is trained using the ResNet18 network as the backbone network; the first layer Layer1, the second layer Layer2 and the third layer Layer3 feature maps of the backbone network are selected as basic feature maps; the resolutions of the three feature maps are unified using the deconvolution upsampling method, and a splicing operation is performed to obtain a preliminary vector matrix Block1; based on the weighted self-generated feature mechanism, the fourth layer feature map Layer4 is generated using the above-mentioned Layer1, Layer2 and Layer3 feature maps, and a splicing operation is performed with the preliminary vector matrix Block1 to obtain a high-dimensional feature vector matrix Block2; the first 500 layers of the channel of Block2 are selected as the final feature vector matrix Block3.
4. The unsupervised track anomaly detection method according to claim 3, characterized in that: The weighted self-generated feature mechanism includes: selecting the first 64 layers of the upsampled Layer2 and Layer3 feature map channels and assigning a weight of 0.4, and adding it to the weight of 0.2 assigned to the Layer1 feature map to obtain the fourth layer feature map Layer4.
5. The unsupervised track anomaly detection method according to claim 1, characterized in that: Multivariate Gaussian fitting function, including: taking n normal training images as a group, and dividing the input image into a pixel block grid, a high-dimensional feature vector matrix of the normal pixel block at (i, j) x ij is the eigenvector value at the normal image (i, j); using the multivariate Gaussian distribution N(μ ij ,Σ ij ) for high-dimensional eigenvector matrices Fitting is performed to obtain the optimal parameter μ ij and Σ ij .
6. The unsupervised track anomaly detection method according to claim 1, characterized in that: The anomaly score measurement function of internal and external balance includes: designing an external square root measurement function D1, calculating the difference score between the test image and the optimal parameter, where x ij ' is the feature vector value at the test image (i, j); Using the internal Mahalanobis distance metric function D2, calculate the difference score inside the test image; The internal and external balanced anomaly score measurement function D is the sum of the two weighted anomaly scores, and the final anomaly score is obtained: D = 0.8 × D1 + 0.2 × D2.
7. An unsupervised track anomaly detection system, characterized in that: include: An acquisition module, used for acquiring a track image to be detected from a bird's-eye view; The detection module is used to process the acquired track image of the top view to be detected by using a pre-trained detection model to obtain a detection result; wherein the pre-trained detection model includes a key component region segmentation unit, a feature aggregation network, a fitting network and a calculation unit; the key component region segmentation unit is used to perform pixel-level segmentation and background simplification of the track image to generate pseudo images of three key components: rails, sleepers and track plates; the feature aggregation network is used to extract and aggregate multi-scale features in the pseudo image to obtain a high-dimensional feature vector matrix; the fitting unit is used to fit the high-dimensional feature vector matrix using a multivariate Gaussian fitting function to obtain the optimal parameters of the normal image; the calculation unit is used to use an internally and externally balanced anomaly score measurement function to calculate the difference between the optimal parameters of the normal image and the high-dimensional feature vector matrix of the test image, calculate the anomaly score, and realize abnormal image discrimination and abnormal pixel positioning.
8. A non-transitory computer-readable storage medium, characterized in that: The non-transitory computer-readable storage medium is used to store computer instructions, and when the computer instructions are executed by the processor, the unsupervised track anomaly detection method as described in any one of claims 1-6 is implemented.
9. A computer device, characterized in that: It includes a memory and a processor, the processor and the memory communicate with each other, the memory stores program instructions that can be executed by the processor, and the processor calls the program instructions to execute the unsupervised track anomaly detection method as described in any one of claims 1 to 6.
10. An electronic device, characterized in that: include: A processor, a memory and a computer program; wherein the processor is connected to the memory, the computer program is stored in the memory, and when the electronic device is running, the processor executes the computer program stored in the memory so that the electronic device executes instructions for implementing the unsupervised track anomaly detection method as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Hyperspectral anomaly detection method based on S1 / 2 norm low-rank representation model
CN112560975A
Mobile phone glass screen unbalanced surface defect classification method based on visual inspection system
CN115457323A
Image processing method and device, intelligent equipment, storage medium and product
CN116977248A
Unsupervised anomaly detection method based on multi-view semantic distance and constrained image reconstruction
CN118072310A
Image anomaly identification method based on nested residual self-encoding model
CN118115448A