A rail foreign matter detection method, device, equipment and medium

By using a foreground extraction model for feature extraction and background subtraction in track detection, the problem of low accuracy in track foreign object detection in existing technologies is solved, and real-time accurate detection under different weather conditions is achieved.

CN115984378BActive Publication Date: 2025-11-11ZHEJIANG DAHUA TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211666936.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-22
Publication Date
2025-11-11
Estimated Expiration
2042-12-22

AI Technical Summary

Technical Problem

Existing technologies for detecting foreign objects on railway tracks have low accuracy, and are prone to missed detection, especially under adverse weather conditions, making real-time and accurate detection impossible.

Method used

A foreground extraction model is used to extract features from the track image and the background image. The probability of each pixel belonging to the foreground target and the background image is obtained by background subtraction to determine whether there is foreign object intrusion on the track.

Benefits of technology

It improves the accuracy of foreign object detection on the track, avoids missed detections due to changes in weather and environment, and achieves real-time and accurate foreign object detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115984378B_ABST
    Figure CN115984378B_ABST
Patent Text Reader

Abstract

This application provides a method, apparatus, device, and medium for detecting foreign objects on a track, addressing the problem of low accuracy in prior art. In this method, a track image and a background image are input into a foreground extraction model. Based on the foreground extraction model, the probability of pixels in the first image belonging to the foreground target and the background image is determined, thereby determining the location of the foreground target in the first image. The location of the foreground target is then used to determine whether a foreign object intrusion exists on the track to be detected. This avoids the situation where weather conditions prevent the detection of foreign objects, leading to missed detections and improving the accuracy of track foreign object detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of image processing technology, and in particular to a method, apparatus, device, and medium for detecting foreign objects in orbit. Background Technology

[0002] In recent years, my country's railway construction and development have progressed rapidly, and passenger and freight transport volumes have continued to increase, greatly facilitating people's daily travel. However, railway tracks are often exposed to the elements, making them susceptible to severe weather and natural disasters. This can lead to rockfalls, mudslides, and other debris intrusions into the track area, affecting train safety. Therefore, early warning of foreign object intrusion into railway tracks is an indispensable part of railway safety transportation. Currently, my country's monitoring technology for foreign object intrusion into railway tracks relies on regular manual inspections. This method is not only extremely labor-intensive but also unable to effectively provide real-time early warnings of foreign objects intruding into the tracks.

[0003] Therefore, the following scheme is proposed in related technologies: Acquire sample images without falling rocks and perform background preprocessing to establish an initial background image. Use the established background model to determine the detection area, extract the red, green, and blue (RGB) color channel features and gradient amplitude features from the image, and based on background subtraction, subtract the gray value of the corresponding pixel in the current background image from the gray value of each pixel in the image to be detected at the current moment. Detect the transformation of the RGB color channel features and gradient amplitude features in the subtracted image, perform threshold segmentation to obtain preliminary detection results, and then introduce the preliminary detection results into the Hue Saturation Value (HSV) color space for analysis to output the final falling rock detection result. Although this method can achieve real-time monitoring, directly using the gray values ​​of pixels for background subtraction is easily affected by weather conditions. In cloudy or poor lighting conditions, the difference in gray values ​​between pixels in two images may be too small, which may prevent the detection of falling rocks, causing the system to miss detections and fail to accurately detect foreign objects in the track. Summary of the Invention

[0004] This application provides a method, apparatus, device, and medium for detecting foreign objects on tracks, in order to solve the problem of low accuracy in detecting foreign objects on tracks in the prior art.

[0005] In a first aspect, embodiments of this application provide a method for detecting foreign objects on a track, the method comprising:

[0006] For the track to be detected for foreign objects, acquire a background image and an image of the track to be detected. The background image includes images of the track where no foreign objects are present.

[0007] The orbit image and background image are input into the foreground extraction model. Based on the foreground extraction model, features are extracted from the background image and the orbit image to obtain feature maps of the background image and the orbit image. Background subtraction is performed on the feature maps of the background image and the orbit image to obtain a first feature map and a second feature map. The first feature map represents the probability that each pixel in the orbit image belongs to the foreground target, and the second feature map represents the probability that each pixel in the orbit image belongs to the background image.

[0008] Based on the probability that each pixel in the orbital image belongs to the foreground target and the probability that it belongs to the background image, the foreground target and its location are determined.

[0009] If the location of the foreground target is within the area where the orbit is located, then it is determined that there is a foreign object in the orbit.

[0010] Secondly, embodiments of this application also provide a track foreign object detection device, the device comprising:

[0011] The acquisition module is used to acquire a background image and an image of the track to be detected for the track to be detected. The background image includes an image of the track where no foreign objects are present.

[0012] The processing module is used to input the orbit image and background image into the foreground extraction model. Based on the foreground extraction model, features are extracted from the background image and the orbit image to obtain feature maps of the background image and the orbit image. Background subtraction is performed on the feature maps of the background image and the orbit image to obtain a first feature map and a second feature map. The first feature map represents the probability that each pixel in the orbit image belongs to the foreground target, and the second feature map represents the probability that each pixel in the orbit image belongs to the background image.

[0013] The determination module is used to determine the foreground target and its location based on the probability that each pixel in the orbital image belongs to the foreground target and the probability that it belongs to the background image.

[0014] The detection module is used to determine the presence of foreign objects in the track if the location of the foreground target is within the area where the track is located.

[0015] Thirdly, embodiments of this application also provide an electronic device, which includes at least a processor and a memory, wherein the processor is used to execute a computer program stored in the memory to implement the steps of the orbital foreign object detection method as described in any of the above claims.

[0016] Fourthly, embodiments of this application also provide a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of the orbital foreign object detection method as described in any of the preceding claims.

[0017] In this embodiment, for the track to be detected for foreign object detection, a background image and a track image to be detected are acquired. The background image includes an image of the track without foreign objects. The track image and the background image are input into a foreground extraction model. Based on the foreground extraction model, feature extraction is performed on the background image and the track image to obtain feature maps of the background image and the track image. Background subtraction is performed on the feature maps of the background image and the track image to obtain a first feature map and a second feature map. The first feature map represents the probability that each pixel in the track image belongs to the foreground target, and the second feature map represents the probability that each pixel in the track image belongs to the background image. Based on the probability that each pixel in the track image belongs to the foreground target and the probability that it belongs to the background image, the foreground target and its location are determined. If the location of the foreground target is within the area where the track is located, it is determined that there is a foreign object in the track. In this application, the orbital image and background image are input into the foreground extraction model. Based on the foreground extraction model, the probability of a pixel in the first image belonging to the foreground target and the background image is determined, thereby determining the location of the foreground target in the first image. Based on the location of the foreground target and the area where the orbit is located, it is determined whether there is foreign object intrusion on the track to be detected. This avoids the situation where the system fails to detect foreign objects due to the influence of weather conditions when directly using the background subtraction method, thus improving the accuracy of orbital foreign object detection. Attached Figure Description

[0018] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0019] Figure 1 A schematic diagram of a foreign object detection process on a track, provided for some embodiments of this application;

[0020] Figure 2 A schematic diagram of the structure of a foreground extraction model provided for some embodiments of this application;

[0021] Figure 3 A schematic diagram of the area where the track is located is provided for some embodiments of this application;

[0022] Figure 4 A schematic diagram of track foreign object detection results provided for some embodiments of this application;

[0023] Figure 5 A schematic flowchart of a foreign object detection process is provided for some embodiments of this application;

[0024] Figure 6 A schematic diagram of the structure of a foreign object detection device for rail tracks is provided for some embodiments of this application;

[0025] Figure 7 This is another schematic diagram of the structure of a terminal device provided for some embodiments of this application. Detailed Implementation

[0026] To make the objectives and implementation methods of this application clearer, the exemplary implementation methods of this application will be clearly and completely described below with reference to the accompanying drawings of the exemplary embodiments of this application. Obviously, the exemplary embodiments described are only some embodiments of this application, and not all embodiments.

[0027] It should be noted that the brief descriptions of terms in this application are only for the convenience of understanding the embodiments described below, and are not intended to limit the embodiments of this application. Unless otherwise stated, these terms should be understood in their ordinary and common meaning.

[0028] The terms "first," "second," "third," etc., used in the specification, claims, and accompanying drawings of this application are used to distinguish similar or related objects or entities, and do not necessarily imply a specific order or sequence, unless otherwise specified. It should be understood that such terms are interchangeable where appropriate.

[0029] The terms “comprising” and “having”, and any variations thereof, are intended to cover but not exclude inclusion, for example, a product or device that includes a range of components is not necessarily limited to all of the components that are clearly listed, but may include other components that are not clearly listed or that are inherent to such product or device.

[0030] The term "module" refers to any known or subsequently developed hardware, software, firmware, artificial intelligence, fuzzy logic, or combination of hardware and / or software code that is capable of performing the functions associated with that element.

[0031] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of this application.

[0032] For ease of explanation, the above description has been provided in conjunction with specific embodiments. However, the above exemplary discussion is not intended to be exhaustive or to limit the embodiments to the specific forms disclosed above. Various modifications and variations can be obtained based on the above teachings. The selection and description of the above embodiments are for the purpose of better explaining the principles and practical applications, thereby enabling those skilled in the art to better utilize the described embodiments and various different variations of embodiments suitable for specific use considerations.

[0033] This application provides a method, apparatus, device, and medium for detecting foreign objects on a track. The method involves acquiring a background image and an image of the track to be detected for the track to be detected. The background image includes images of the track without foreign objects. The track image and background image are input into a foreground extraction model. Based on the foreground extraction model, feature extraction is performed on the background image and track image to obtain feature maps of the background image and track image. Background subtraction is performed on the feature maps of the background image and track image to obtain a first feature map and a second feature map. The first feature map represents the probability that each pixel in the track image belongs to a foreground target, and the second feature map represents the probability that each pixel in the track image belongs to the background image. Based on the probabilities of each pixel in the track image belonging to the foreground target and the probability of belonging to the background image, the foreground target and its location are determined. If the location of the foreground target is within the area where the track is located, it is determined that a foreign object exists within the track. In this application, the orbital image and background image are input into the foreground extraction model. Based on the foreground extraction model, the probability of a pixel in the first image belonging to the foreground target and the background image is determined, thereby determining the position information of the foreground target in the first image. Based on the position information of the foreground target, it is determined whether there is a foreign object intrusion on the track to be detected. This avoids the situation where the system fails to detect foreign objects due to the influence of weather conditions when directly using the background subtraction method, thus improving the accuracy of track foreign object detection.

[0034] Example 1:

[0035] Figure 1 A schematic diagram of a foreign object detection process for orbits, provided for some embodiments of this application, includes:

[0036] S101: For the track to be detected for foreign objects, acquire a background image and an image of the track to be detected. The background image includes an image of the track where no foreign objects are present.

[0037] The foreign object detection method for tracks provided in this application is applied to electronic devices, such as personal computers (PCs), servers, and image acquisition devices like cameras. If the electronic device is not an image acquisition device, it can acquire background and track images from an image acquisition device.

[0038] For example, the background image can be a pre-acquired, unprocessed original historical image of the track without foreign objects, or the background image can be an image obtained by processing the original historical image including the track without foreign objects. For example, the background image can be an image of the original historical image including the track without foreign objects, cropped and retaining the track portion.

[0039] As another example, the track image to be detected can be an unprocessed raw real-time image containing the track, or the track image to be detected can be an image obtained by processing the raw real-time image containing the track. For example, the track image can be an image that is cropped from the raw real-time image containing the track and retains the track portion.

[0040] As another example, the track image to be detected can also be a single real-time image decoded from a video stream acquired in real time.

[0041] The background image and track image include the track region to be detected. Optionally, in addition to the track image to be detected, the background image and track image may also include non-track regions, such as mountains, sky, buildings, roads, etc.

[0042] Typically, the installation location and angle of image acquisition devices are fixed, and their field of view can also be fixed. Furthermore, the image acquisition device can capture images of the track to be detected. In other words, in this embodiment, the background image and track image capture the same image range and are images captured for the same track to be detected. Therefore, the presence of foreign objects on the track to be detected can be determined based on the background image and track image.

[0043] S102: Input the orbit image and background image into the foreground extraction model. Based on the foreground extraction model, perform feature extraction on the background image and orbit image to obtain feature maps of the background image and orbit image. Perform background subtraction on the feature maps of the background image and orbit image to obtain a first feature map and a second feature map. The first feature map represents the probability that each pixel in the orbit image belongs to the foreground target, and the second feature map represents the probability that each pixel in the orbit image belongs to the background image.

[0044] The foreground extraction model is a two-ended input model for detecting foreign object intrusions on railway tracks. The foreground extraction model has two inputs and one output, where the inputs are the track image to be detected and the background image, respectively, and the output is the preliminary result of foreign object detection, namely the first feature map and the second feature map.

[0045] The foreground extraction model is a pre-trained model that includes feature extraction and background subtraction functions. Based on the feature extraction function, features are extracted from the background and orbit images, resulting in feature maps for the background and orbit images. Then, based on the background subtraction function, background subtraction is performed on these feature maps to obtain a first and a second feature map that distinguishes the foreground object from the background. The foreground extraction model in this application is a deep learning-based model that exhibits good adaptability to changes in lighting and weather conditions.

[0046] The pixel value of each pixel in the first feature map represents the probability that the corresponding pixel in the orbit image belongs to the foreground target, and the pixel value of each pixel in the second feature map represents the probability that the corresponding pixel in the orbit image belongs to the background image. The pixel values ​​in both the first and second feature maps range from 0 to 1.

[0047] S103: Determine the foreground target and its location based on the probability that each pixel in the orbital image belongs to the foreground target and the probability that it belongs to the background image.

[0048] For each pixel in the orbital image, its classification as a foreground object can be determined based on the probability that it belongs to the foreground and the probability that it belongs to the background. In one possible implementation, if the probability of a pixel belonging to the foreground is greater than the probability of it belonging to the background, then the pixel is classified as a foreground object. In another possible implementation, the classification can be determined by whether the probability of a pixel belonging to the foreground is greater than a foreground threshold; if so, the pixel is classified as a foreground object. These two possible implementations are merely examples, and actual implementations may include, but are not limited to, these two possible implementations.

[0049] Based on all pixels belonging to the foreground target, determine the foreground target and its location. The foreground target may include one or more, and each foreground target may be located at a different location.

[0050] S104: If the location of the foreground target is within the area where the track is located, then it is determined that there is a foreign object in the track.

[0051] Based on the location of the foreground target and the area where the orbit is located, determine whether the location of the foreground target overlaps with the area where the orbit is located.

[0052] If the location of the foreground target overlaps with the area where the track is located, then the foreground target can be considered to be within the area where the track is located, thus confirming the presence of a foreign object on the track. In this case, the electronic device can issue an alarm. Alarm methods include, but are not limited to, text alerts and voice alerts.

[0053] If the location of the foreground target does not overlap with the area where the track is located, it can be assumed that the location of the foreground target is not within the area where the track is located, thus determining that there is no foreign object in the track. At this time, the electronic device can end the processing of the current frame image, i.e., the track image, and can continue to detect whether there is a foreign object in the track in the subsequently acquired images. For example, the subsequent images can be used as the track images to be detected, and the process of detecting foreign objects in the track can be returned to S101 to continue.

[0054] In this application, by inputting the track image and background image into the foreground extraction model, the probability of a pixel in the first image belonging to the foreground target and the background image is determined based on the foreground extraction model, thereby determining the location of the foreground target in the first image. Based on the location of the foreground target, it is determined whether there is a foreign object intrusion on the track to be detected. This avoids the situation where the system fails to detect foreign objects due to the influence of weather conditions when directly using the background subtraction method, thus improving the accuracy of track foreign object detection.

[0055] Example 2:

[0056] To further improve the accuracy of foreign object detection on the track, based on the above embodiments, in this embodiment, the track image and background image are input into a foreground extraction model. Based on the foreground extraction model, feature extraction is performed on the background image and track image to obtain feature maps of the background image and track image. Background subtraction is then performed on the feature maps of the background image and track image to obtain a first feature map and a second feature map, including:

[0057] The background image and the track image are input into the foreground extraction model. Based on multiple first convolutional units in the foreground extraction model corresponding to the background image, multiple feature extractions are performed on the background image to obtain the feature map of the background image. Based on multiple second convolutional units in the foreground extraction model corresponding to the track image, multiple feature extractions are performed on the track image to obtain the feature map of the track image.

[0058] Based on the foreground extraction model, the feature maps of the background image and the orbit image are subjected to multiple background subtraction and upsampling. Then, through the third convolutional unit of the foreground extraction model, each pixel in the image obtained after the last background subtraction and upsampling is classified to obtain the first feature map and the second feature map.

[0059] The foreground extraction model contains multiple first convolutional units and multiple second convolutional units. The structure of each first convolutional unit can be the same or different, and the structure of each second convolutional unit can also be the same or different. However, the structures of first and second convolutional units within the same layer are identical. Each first and second convolutional unit consists of a convolutional kernel, a batch normalization (BN) layer, a rectified linear unit (ReLU) activation function, a pooling layer, and an output channel. Each pass through a first convolutional unit performs a downsampling (feature extraction) of the background image; each pass through a second convolutional unit performs a downsampling (feature extraction) of the track image. The input and output image sizes, convolutional kernel size, number of output channels, and downsampling factor of each first and second convolutional unit can be the same or different; no specific restrictions are imposed here.

[0060] The number of multiple first convolutional units is the same as the number of times feature extraction is performed on the background image, that is, feature extraction is performed on the background image once every time a first convolutional unit is passed; the number of multiple second convolutional units is the same as the number of times feature extraction is performed on the track image, that is, feature extraction is performed on the track image once every time a second convolutional unit is passed; the number of first convolutional units and second convolutional units is the same, that is, the number of times feature extraction is performed on the background image and the track image is the same.

[0061] Based on the background subtraction function in the foreground extraction model, multiple background subtractions are performed on the feature maps of the background image and the track image. The feature maps after each background subtraction are upsampled and concatenated with the feature maps obtained from the previous layer's background subtraction along the channel dimension. Feature extraction is then performed through a third convolutional unit. The feature map obtained after the final background subtraction and upsampling is input into the last third convolutional unit, a two-channel convolutional module used to classify each pixel in the feature map, distinguishing the probability of each pixel belonging to the foreground target or the background image. Therefore, the output of the last third convolutional unit includes a first feature map and a second feature map. The pixel value of a pixel in the first feature map represents the probability that the pixel in the track image belongs to the foreground target, and the pixel value of a pixel in the second feature map represents the probability that the pixel in the track image belongs to the background image.

[0062] Specifically, based on the foreground extraction model, multiple background subtraction and upsampling are performed on the feature maps of the background image and the orbit image. Then, the third convolutional unit of the foreground extraction model is used to classify each pixel in the image obtained after the final background subtraction and upsampling, resulting in a first feature map and a second feature map, including:

[0063] In the foreground extraction model, for each background difference and upsampling, the feature map of the background image and the feature map of the orbit image are subtracted, and the subtracted image is upsampled to obtain the first image. The feature map output by the first convolutional unit of the previous layer is subtracted from the feature map output by the second convolutional unit of the previous layer to obtain the second image. The first image and the second image are stacked in the channel dimension, and the stacked image is input into the third convolutional unit to obtain the difference feature map.

[0064] Based on the difference feature map, multiple background subtraction and upsampling are performed in the foreground extraction model. After the last background subtraction and upsampling, each pixel in the obtained difference feature map is classified by the third convolutional unit of the foreground extraction model to obtain the first feature map and the second feature map.

[0065] In the foreground extraction model, for each background difference and upsampling in the background difference function, the feature map of the background image and the feature map of the track image are subtracted, and the subtracted feature map is upsampled to determine the first image after upsampling. The feature map of the background image output by the first convolutional unit of the previous layer and the feature map of the track image output by the second convolutional unit of the previous layer are subtracted to determine the second image after subtraction. The first image and the second image are stacked in the channel dimension, and the stacked image is input into the third convolutional unit to obtain the difference feature map.

[0066] For each background difference feature map, an upsampling operation is performed to determine the third image after upsampling. The feature map of the background image output from the first convolutional unit above this difference feature map is subtracted from the feature map of the orbital image output from the second convolutional unit above, resulting in the fourth image. The third and fourth images are then stacked along the channel dimension, and the stacked image is input into the third convolutional unit to obtain the second difference feature map. This process is repeated until the last third convolutional unit of the foreground extraction model is used to obtain the first and second feature maps.

[0067] In the foreground extraction model, after the last background difference and upsampling, the third difference feature map is determined. The third difference feature map is then passed through the last third convolutional unit to classify each pixel in the third difference feature map, that is, to distinguish the probability of each pixel belonging to the foreground target and the probability of the background image, thus obtaining the first feature map and the second feature map.

[0068] Let's illustrate with a specific example. Figure 2The diagram illustrates the structure of the foreground extraction model. "CBR" represents the convolutional module, consisting of convolution (Conv) operations, batch normalization (BN) layers, and a ReLU activation function. "Subtract" is the subtraction operation used to subtract two feature maps. "Up-sample" is the upsampling operation used to increase the size of the feature maps. "+" is the stacking operation used to fuse two feature maps along the channel dimension. The foreground extraction model has two inputs: a background image and a track image. The convolutional modules in the same layer corresponding to the background and track images have the same structure. Specifically, the CBR... 1-1 Structure and CBR 2-1 The structures are the same, CBR 1-2 Structure and CBR 2-2 The structures are the same, CBR 1-3 Structure and CBR 2-3 The structures are the same, CBR 1-4 Structure and CBR 2-4 The structures are the same, CBR 1-5 Structure and CBR 2-5 The structures are identical. The background image and track image at both inputs of the foreground extraction model are both 960×640×3 pixels, CBR. 1-1 and CBR 2-1 Each layer consists of a convolutional kernel with 32 output channels, a 3×3 size, and a stride of 1, a batch normalization (BN) layer, a ReLU activation function, and a pooling layer with a stride of 2. (Background image x) 1_0 After CBR 1-1 The first background feature map x of the background image is then obtained. 1_1 orbital image x 2_0 After CBR 2-1 The first orbit feature map x of the orbit image is then obtained. 2_1 x 1_1 and x 2_1 The dimensions are 480×320×32.

[0069] CBR 1-2 and CBR 2-2 The number of channels is 64, and other parameters are the same as CBR. 1-1 CBR 2-1 Similarly, the first background feature map is processed by CBR. 1-2 The second background feature map x of the background image is then obtained. 1_2 The first orbital feature map was processed by CBR 2-2 The second orbit feature map x of the orbit image is then obtained. 2_2 x 1_2 and x 2_2 The dimensions are 240×160×64.

[0070] CBR 1-3 and CBR 2-3 The number of channels is 128, and other parameters are the same as CBR. 1-1 CBR 2-1 Similarly, the second background feature map is processed by CBR. 1-3 The third background feature map x of the background image is then obtained. 1_3 The second orbital feature map is processed by CBR 2-3 The third orbital feature map x of the orbital image was then obtained. 2_3 x 1_3 and x 2_3 The dimensions are 120×80×128.

[0071] CBR 1-4 and CBR 2-4 The number of channels is 256, and other parameters are the same as CBR. 1-1 CBR 2-1 Similarly, the third background feature map is processed by CBR. 1-4 The fourth background feature map x of the background image is then obtained. 1_4 The third orbital feature map was processed by CBR 2-4 The fourth orbital feature map x of the orbital image was then obtained. 2_4 x 1_4 and x 2_4 The dimensions are 60×40×256.

[0072] CBR 1-5 and CBR 2-5 The number of channels is 512, and other parameters are the same as CBR. 1-1 CBR 2-1 Similarly, the fourth background feature map is processed by CBR. 1-5 The fifth background feature map x of the background image is then obtained. 1_5 The fourth orbital feature map was processed by CBR 2-5 The fifth orbital feature map x of the orbital image was then obtained. 2_5 x 1_5 and x 2_5 The dimensions are 30×20×512.

[0073] Based on the background subtraction function in the foreground extraction model, the fifth background feature map and the fifth orbital feature map are subtracted to obtain a first subtracted feature map with a size of 30×20×512. Then, this first subtracted feature map is upsampled to obtain a first upsampled feature map with a size of 60×40×512. The fourth background feature map and the fourth orbital feature map are subtracted to obtain a second subtracted feature map with a size of 60×40×256. The first upsampled feature map and the second subtracted feature map are superimposed along the channel dimension to obtain a first superimposed feature map x3 with a size of 60×40×768. The superimposed feature map is then fed into a third convolutional unit CBR with a kernel size of 1×1 and 256 channels. 3-1 Within the module, a differential feature map with a size of 60×40×256 can be obtained.

[0074] For the difference feature map, it is upsampled to obtain a second upsampled feature map with a size of 120×80×256. The third background feature map and the third track feature map are subtracted to obtain a third subtracted feature map with a size of 120×80×128. The second upsampled feature map and the third subtracted feature map are superimposed along the channel dimension to obtain a second superimposed feature map x4 with a size of 120×80×384. The stacked second superimposed feature map is fed into a third convolutional unit CBR with a kernel size of 1×1 and 128 channels. 3-2 Within the module, a second differential feature map with a size of 120×80×128 can be obtained.

[0075] For the second difference feature map, it is upsampled to determine a third upsampled feature map with a size of 240×160×128. The second background feature map and the second track feature map are subtracted to obtain a fourth subtracted feature map with a size of 240×160×64. The third upsampled feature map and the fourth subtracted feature map are superimposed along the channel dimension to obtain a third superimposed feature map x5 with a size of 240×160×192. The stacked third superimposed feature map is fed into a third convolutional unit CBR with a kernel size of 1×1 and 64 channels. 3-3 Within the module, a third difference feature map with a size of 240×160×64 can be obtained.

[0076] For the third difference feature map, it is upsampled to determine a fourth upsampled feature map with a size of 480×320×64. The first background feature map and the first track feature map are subtracted to obtain a fifth subtracted feature map with a size of 480×320×32. The fourth upsampled feature map and the fifth subtracted feature map are superimposed along the channel dimension to obtain a fourth superimposed feature map x6 with a size of 480×320×96. The stacked fourth superimposed feature map is then fed into a third convolutional unit CBR with a kernel size of 1×1 and 32 channels. 3-4 Within the module, a fourth difference feature map with a size of 480×320×32 can be obtained.

[0077] For the fourth difference feature map, it is upsampled to determine a fifth upsampled feature map x7 with a size of 960×640×32; the fifth upsampled feature map is then fed into the third convolutional unit CBR with a kernel size of 1×1 and 2 channels. 3-5 The module can obtain a feature map with a size of 960×640×2. Each channel outputs one image, resulting in two feature maps: the first feature map and the second feature map, both with a size of 960×640.

[0078] By using a deep learning-based foreground feature extraction model to perform background subtraction on the background and orbit images, the system avoids the situation where the direct background subtraction method is susceptible to weather conditions and may fail to detect foreign objects, thus improving the accuracy of orbit object detection.

[0079] Example 4:

[0080] To further improve the accuracy of foreign object detection on the track, based on the above embodiments, this application embodiment determines the foreground target and its location according to the probability that each pixel in the track image belongs to the foreground target and the probability that it belongs to the background image, including:

[0081] Based on the probability that each pixel in the orbital image belongs to the foreground target and the probability that it belongs to the background image, the pixel value of the pixel that belongs to the foreground target with a greater probability than the probability that it belongs to the background image is updated to the first pixel value, and the pixel value of the pixel that belongs to the foreground target with a probability that is not greater than the probability that it belongs to the background image is updated to the second pixel value.

[0082] Based on the updated pixel values, determine the binary image of the orbital image;

[0083] In a binary image, determine the foreground target and its location.

[0084] Since the foreground extraction model outputs two-channel feature images—a first feature map and a second feature map—and these two feature maps are used to distinguish between pixels in the track image belonging to the foreground and background, the probability of a pixel belonging to the foreground can be determined based on the pixel values ​​in the first and second feature maps. Specifically, the pixel values ​​in the first feature map represent the probability that the pixel in the track image belongs to the foreground, and the pixel values ​​in the second feature map represent the probability that the pixel in the track image belongs to the background.

[0085] For each pixel in the orbital image, based on the probability that the pixel belongs to the foreground target and the probability that it belongs to the background image (i.e., based on the pixel values ​​of the first feature map and the second feature map), the pixel values ​​of pixels whose probability of belonging to the foreground target is greater than their probability of belonging to the background image are updated to the first pixel value; that is, pixels whose pixel values ​​in the first feature map are greater than the corresponding pixel values ​​in the second feature map are updated to the first pixel value. The pixel values ​​of pixels whose probability of belonging to the foreground target is not greater than their probability of belonging to the background image are updated to the second pixel value; that is, pixels whose pixel values ​​in the first feature map are not greater than the corresponding pixel values ​​in the second feature map are updated to the second pixel value.

[0086] The process of updating the pixel value of a pixel whose probability of belonging to the foreground object is greater than its probability of belonging to the background image to the first pixel value, and updating the pixel value of a pixel whose probability of belonging to the foreground object is no greater than its probability of belonging to the background image to the second pixel value, satisfies the following formula:

[0087] Where (x,y) represents the position coordinates of the pixel, f1(x,y) represents the probability that the pixel with position coordinates (x,y) in the first feature map belongs to the foreground target, f2(x,y) represents the probability that the pixel with position coordinates (x,y) in the second feature map belongs to the background image, a represents the first pixel value, b represents the second pixel value, and f(x,y) represents the updated pixel value of the pixel with position coordinates (x,y).

[0088] In this context, the range of values ​​for x is related to the length of the track image, and the range of values ​​for y is related to the width of the track image. For example, if the size of the track image is 960×640, then the range of values ​​for x is 0≤x<960, and the range of values ​​for y is 0≤y<640. The probability that a pixel at coordinates (x,y) in the first feature image belongs to the foreground target and the probability that it belongs to the background image ranges from 0 to 1, and the sum of the probabilities of the same pixel belonging to the foreground target and the probability that it belongs to the background image is 1. The first pixel value is greater than the second pixel value. The specific values ​​of the first and second pixel values ​​are not restricted; for example, the first pixel value can be 255, and the second pixel value can be 0.

[0089] Based on the updated pixel values, a binary image of the orbital image is determined. In this binary image, each pixel has either a first pixel value or a second pixel value. Pixels with the first pixel value belong to the foreground object, while pixels with the second pixel value belong to the background image.

[0090] In a binary image, the foreground target is determined based on the pixels that belong to the foreground target in the binary image, and the location of the foreground target is determined based on the position coordinates of the pixels that belong to the foreground target.

[0091] The system determines a binary image based on the pixels in the orbital image, and then uses the binary image to determine the foreground target and its location, further improving the accuracy of foreign object detection on the orbit.

[0092] Example 5:

[0093] To further improve the accuracy of foreign object detection on the track, based on the above embodiments, in this application embodiment, determining the foreground target and its location in the binary image includes:

[0094] The binary image is filtered to detect false positives and to identify the foreground target.

[0095] For the foreground target, determine the maximum bounding rectangle of the foreground target, and define the location of the foreground target as the maximum bounding rectangle.

[0096] Erosion and dilation operations are performed on the binary image to more clearly distinguish the foreground target from the background image. The eroded and dilated binary image is then filtered to remove the background portion, revealing the foreground target in the orbital image. The orbital image may or may not contain any foreground target.

[0097] For each foreground object, determine its maximum bounding rectangle. If there are multiple foreground objects in the binary image, the maximum bounding rectangle of each foreground object can be determined.

[0098] For example, for a foreground target, the contour region of the foreground target can be calculated, the maximum bounding rectangle of the foreground target can be calculated using the result of the contour region, and the foreground target can be marked in the orbital image using the maximum bounding rectangle. The position of the maximum bounding rectangle in the orbital image or the marked position is determined as the location of the foreground target.

[0099] In another example, for the foreground target, since the pixels with the first pixel value in the binary image are determined to belong to the foreground target, and the pixels with the second pixel value are determined to belong to the background image, the maximum bounding rectangle of the first pixel value is determined based on the pixel value of each pixel and the distribution of each pixel. In other words, the maximum bounding rectangle of the foreground target is determined, and the maximum bounding rectangle is determined as the location of the foreground target.

[0100] Filtering out false positives in binary images can reduce interference from pixels belonging to the background image, further improving the accuracy of foreign object detection on the track.

[0101] Example 6:

[0102] To further improve the accuracy of foreign object detection on the track, based on the above embodiments, the method in this application embodiment further includes:

[0103] For the track, Hough line detection is used to detect the track in the background image and the track image, the track is fitted to a straight line, and the area contained by the straight line of the track is determined as the area where the track is located.

[0104] For the track to be detected in the background image and track image, Hough line detection is used to detect the track, fitting the track to be detected as a straight line, and defining the region contained within the straight line as the region where the track is located. For example... Figure 3 As shown, the straight line in the orbital image that fits the orbit to the orbital plane is... Figure 3 The white straight line in the middle, Figure 3 The area enclosed by the white line is the region where the track is located.

[0105] The input foreground extraction model can be the orbital image and the background image containing the track to be detected, or it can be an image with the track to be detected annotated.

[0106] In addition, the area where the track is located can be defined as the warning area in the track foreign object detection scenario. When the foreground target appears in the warning area, that is, when the maximum bounding rectangle of the foreground target overlaps with the warning area, it is determined that there is a foreign object in the track area to be detected.

[0107] For example, Figure 4 A schematic diagram of the foreign object detection results on the track is shown. Figure 4 The trajectory is fitted to a straight line, which is... Figure 4 The white line in the middle defines the area contained within the track line as the track area; in Figure 4 The largest bounding rectangle of the foreground target is marked in the image, which is the... Figure 4 The black rectangle in the middle. According to... Figure 4 By identifying the location of the track area and the maximum bounding rectangle of the foreground target marked in the image, it can be determined that the maximum bounding rectangle of the foreground target is within the track area, thus confirming the presence of a foreign object within the track area to be detected.

[0108] By fitting the track to be detected to a straight line through Hough line detection, the location of the track can be determined more accurately, further improving the accuracy of track foreign object detection.

[0109] Figure 5 This is a flowchart illustrating the foreign object detection process on a track. The background image and the track image are input into a foreground extraction model. Based on this model, a preliminary detection result is output, namely, a first feature map and a second feature map used to distinguish the foreground target from the background image. For this preliminary detection result, erosion and dilation operations in morphological operations are used to filter out false foreground targets, and the location of the foreground target is determined. Based on the location of the foreground target and the region where the track to be detected is located, the track foreign object detection result is determined. Specifically, the track foreign object detection result includes: if the location of the foreground target is within the region where the track is located, then a foreign object is determined to exist in the track; if the location of the foreground target is not within the region where the track is located, then no foreign object is determined to exist in the track.

[0110] Example 7:

[0111] Based on the same technical concept and the above embodiments, this application provides a track foreign object detection device. Figure 6 A schematic diagram of a track foreign object detection device is provided for some embodiments of this application, such as... Figure 6 As shown, the device includes:

[0112] The acquisition module 601 is used to acquire a background image and an image of the track to be detected for the track to be detected. The background image includes an image of the track where no foreign object exists.

[0113] The processing module 602 is used to input the orbit image and the background image into the foreground extraction model, and to extract features from the background image and the orbit image based on the foreground extraction model to obtain feature maps of the background image and the orbit image. The feature maps of the background image and the orbit image are then subjected to background subtraction to obtain a first feature map and a second feature map. The first feature map represents the probability that each pixel in the orbit image belongs to the foreground target, and the second feature map represents the probability that each pixel in the orbit image belongs to the background image.

[0114] The determination module 603 is used to determine the foreground target and its location based on the probability that each pixel in the orbit image belongs to the foreground target and the probability that it belongs to the background image.

[0115] The detection module 604 is used to determine the presence of foreign objects in the track if the location of the foreground target is within the area where the track is located.

[0116] In one possible implementation, the processing module 602 is specifically used to input the background image and the track image into the foreground extraction model, perform multiple feature extractions on the background image based on multiple first convolutional units corresponding to the background image in the foreground extraction model to obtain a feature map of the background image, and perform multiple feature extractions on the track image based on multiple second convolutional units corresponding to the track image in the foreground extraction model to obtain a feature map of the track image; based on the foreground extraction model, perform multiple background subtraction and upsampling on the feature maps of the background image and the track image, and classify each pixel in the image obtained after the last background subtraction and upsampling through the third convolutional unit of the foreground extraction model to obtain a first feature map and a second feature map.

[0117] In one possible implementation, the processing module 602 is specifically configured to, in the foreground extraction model, subtract the feature map of the background image and the feature map of the orbit image for each background difference and upsampling, upsample the subtracted image to obtain a first image, subtract the feature map output by the first convolutional unit of the previous layer from the feature map output by the second convolutional unit of the previous layer to obtain a second image, stack the first image and the second image in the channel dimension, and input the stacked image into the third convolutional unit to obtain a difference feature map; based on the difference feature map, perform multiple background differences and upsampling in the foreground extraction model, and after the last background difference and upsampling, classify each pixel in the obtained difference feature map through the third convolutional unit of the foreground extraction model to obtain a first feature map and a second feature map.

[0118] In one possible implementation, the determining module 603 is specifically configured to, based on the probability that each pixel in the orbit image belongs to the foreground target and the probability that it belongs to the background image, update the pixel values ​​of pixels whose probability of belonging to the foreground target is greater than that of belonging to the background image to a first pixel value, and update the pixel values ​​of pixels whose probability of belonging to the foreground target is not greater than that of belonging to the background image to a second pixel value; determine a binary image of the orbit image based on the updated pixel values; and determine the foreground target and its location in the binary image.

[0119] In one possible implementation, the process of updating the pixel value of a pixel whose probability of belonging to the foreground object is greater than its probability of belonging to the background image to a first pixel value, and updating the pixel value of a pixel whose probability of belonging to the foreground object is no greater than its probability of belonging to the background image to a second pixel value, satisfies the following formula: Where (x,y) represents the position coordinates of the pixel, f1(x,y) represents the probability that the pixel with position coordinates (x,y) in the first feature map belongs to the foreground target, f2(x,y) represents the probability that the pixel with position coordinates (x,y) in the second feature map belongs to the background image, a represents the first pixel value, b represents the second pixel value, and f(x,y) represents the updated pixel value of the pixel with position coordinates (x,y).

[0120] In one possible implementation, the determining module 603 is specifically used to perform filtering and false detection processing on the binary image to determine the foreground target; for the foreground target, the maximum bounding rectangle of the foreground target is determined, and the maximum bounding rectangle is determined as the location of the foreground target.

[0121] In one possible implementation, the acquisition module 601 is further configured to detect the track in the background image and the track image based on Hough line detection, fit the track to a straight line, and determine the area contained by the straight line of the track as the area where the track is located.

[0122] Example 8:

[0123] Based on the same technical concept, this application also provides an electronic device. Figure 7 This application provides a schematic diagram of an electronic device structure, such as... Figure 7 As shown, it includes: processor 701, communication interface 702, memory 703 and communication bus 704, wherein processor 701, communication interface 702 and memory 703 communicate with each other through communication bus 704.

[0124] The memory 703 stores a computer program. When the program is executed by the processor 701, the processor 701 performs the following steps:

[0125] For the track to be detected for foreign objects, acquire a background image and an image of the track to be detected. The background image includes images of the track where no foreign objects are present.

[0126] The orbit image and background image are input into the foreground extraction model. Based on the foreground extraction model, features are extracted from the background image and the orbit image to obtain feature maps of the background image and the orbit image. Background subtraction is performed on the feature maps of the background image and the orbit image to obtain a first feature map and a second feature map. The first feature map represents the probability that each pixel in the orbit image belongs to the foreground target, and the second feature map represents the probability that each pixel in the orbit image belongs to the background image.

[0127] Based on the probability that each pixel in the orbital image belongs to the foreground target and the probability that it belongs to the background image, the foreground target and its location are determined.

[0128] If the location of the foreground target is within the area where the orbit is located, then it is determined that there is a foreign object in the orbit.

[0129] In one possible implementation, the processor 701 is specifically configured to input the background image and the track image into the foreground extraction model, perform multiple feature extractions on the background image based on multiple first convolutional units corresponding to the background image in the foreground extraction model to obtain a feature map of the background image, and perform multiple feature extractions on the track image based on multiple second convolutional units corresponding to the track image in the foreground extraction model to obtain a feature map of the track image; based on the foreground extraction model, perform multiple background subtraction and upsampling on the feature maps of the background image and the track image, and classify each pixel in the image obtained after the last background subtraction and upsampling through the third convolutional unit of the foreground extraction model to obtain a first feature map and a second feature map.

[0130] In one possible implementation, the processor 701 is specifically configured to, in the foreground extraction model, subtract the feature map of the background image and the feature map of the orbit image for each background difference and upsampling, upsample the subtracted image to obtain a first image, subtract the feature map output by the first convolutional unit of the previous layer from the feature map output by the second convolutional unit of the previous layer to obtain a second image, stack the first image and the second image in the channel dimension, and input the stacked image into the third convolutional unit to obtain a difference feature map; based on the difference feature map, perform multiple background differences and upsampling in the foreground extraction model, and after the last background difference and upsampling, classify each pixel in the obtained difference feature map through the third convolutional unit of the foreground extraction model to obtain a first feature map and a second feature map.

[0131] In one possible implementation, the processor 701 is specifically configured to, based on the probability that each pixel in the orbit image belongs to the foreground target and the probability that it belongs to the background image, update the pixel value of pixels whose probability of belonging to the foreground target is greater than that of belonging to the background image to a first pixel value, and update the pixel value of pixels whose probability of belonging to the foreground target is not greater than that of belonging to the background image to a second pixel value; determine a binary image of the orbit image based on the updated pixel values; and determine the foreground target and its location in the binary image.

[0132] In one possible implementation, the process of updating the pixel value of a pixel whose probability of belonging to the foreground object is greater than its probability of belonging to the background image to a first pixel value, and updating the pixel value of a pixel whose probability of belonging to the foreground object is no greater than its probability of belonging to the background image to a second pixel value, satisfies the following formula: Where (x,y) represents the position coordinates of the pixel, f1(x,y) represents the probability that the pixel with position coordinates (x,y) in the first feature map belongs to the foreground target, f2(x,y) represents the probability that the pixel with position coordinates (x,y) in the second feature map belongs to the background image, a represents the first pixel value, b represents the second pixel value, and f(x,y) represents the updated pixel value of the pixel with position coordinates (x,y).

[0133] In one possible implementation, the processor 701 is specifically configured to perform filtering and false detection processing on the binary image to determine the foreground target; for the foreground target, determine the maximum bounding rectangle of the foreground target, and determine the location of the foreground target as the maximum bounding rectangle.

[0134] In one possible implementation, the processor 701 is specifically configured to detect the track in the background image and the track image based on Hough line detection, fit the track to a straight line, and determine the region contained by the straight line of the track as the region where the track is located.

[0135] The communication bus mentioned in the above electronic devices can be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus, etc. This communication bus can be divided into address bus, data bus, control bus, etc. For ease of illustration, only one thick line is used to represent it in the diagram, but this does not mean that there is only one bus or one type of bus.

[0136] The communication interface 702 is used for communication between the above-mentioned electronic device and other devices.

[0137] The memory may include RAM (Random Access Memory) or NVM (Non-Volatile Memory), such as at least one disk storage device. Optionally, the memory may also be at least one storage device located remotely from the aforementioned processor.

[0138] The processors mentioned above can be general-purpose processors, including central processing units, network processors (NPs), etc.; they can also be DSPs (Digital Signal Processors), application-specific integrated circuits, field-programmable gate arrays or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc.

[0139] Example 9:

[0140] Based on the same technical concept, embodiments of this application provide a computer-readable storage medium storing a computer program executable by an electronic device. When the program is run on the electronic device, the electronic device performs the following steps:

[0141] For the track to be detected for foreign objects, acquire a background image and an image of the track to be detected. The background image includes images of the track where no foreign objects are present.

[0142] The orbit image and background image are input into the foreground extraction model. Based on the foreground extraction model, features are extracted from the background image and the orbit image to obtain feature maps of the background image and the orbit image. Background subtraction is performed on the feature maps of the background image and the orbit image to obtain a first feature map and a second feature map. The first feature map represents the probability that each pixel in the orbit image belongs to the foreground target, and the second feature map represents the probability that each pixel in the orbit image belongs to the background image.

[0143] Based on the probability that each pixel in the orbital image belongs to the foreground target and the probability that it belongs to the background image, the foreground target and its location are determined.

[0144] If the location of the foreground target is within the area where the orbit is located, then it is determined that there is a foreign object in the orbit.

[0145] In one possible implementation, the orbit image and background image are input into a foreground extraction model. Based on the foreground extraction model, feature extraction is performed on the background image and orbit image to obtain feature maps of the background image and orbit image. Background subtraction is then performed on the feature maps of the background image and orbit image to obtain a first feature map and a second feature map, including:

[0146] The background image and the track image are input into the foreground extraction model. Based on multiple first convolutional units in the foreground extraction model corresponding to the background image, multiple feature extractions are performed on the background image to obtain the feature map of the background image. Based on multiple second convolutional units in the foreground extraction model corresponding to the track image, multiple feature extractions are performed on the track image to obtain the feature map of the track image.

[0147] Based on the foreground extraction model, the feature maps of the background image and the orbit image are subjected to multiple background subtraction and upsampling. Then, through the third convolutional unit of the foreground extraction model, each pixel in the image obtained after the last background subtraction and upsampling is classified to obtain the first feature map and the second feature map.

[0148] In one possible implementation, based on a foreground extraction model, multiple background subtraction and upsampling are performed on the feature maps of the background image and the orbit image. Then, the third convolutional unit of the foreground extraction model is used to classify each pixel in the image obtained after the final background subtraction and upsampling, resulting in a first feature map and a second feature map, including:

[0149] In the foreground extraction model, for each background difference and upsampling, the feature map of the background image and the feature map of the orbit image are subtracted, and the subtracted image is upsampled to obtain the first image. The feature map output by the first convolutional unit of the previous layer is subtracted from the feature map output by the second convolutional unit of the previous layer to obtain the second image. The first image and the second image are stacked in the channel dimension, and the stacked image is input into the third convolutional unit to obtain the difference feature map.

[0150] Based on the difference feature map, multiple background subtraction and upsampling are performed in the foreground extraction model. After the last background subtraction and upsampling, each pixel in the obtained difference feature map is classified by the third convolutional unit of the foreground extraction model to obtain the first feature map and the second feature map.

[0151] In one possible implementation, determining the foreground target and its location based on the probability that each pixel in the orbital image belongs to the foreground target and the probability that it belongs to the background image includes:

[0152] Based on the probability that each pixel in the orbital image belongs to the foreground target and the probability that it belongs to the background image, the pixel value of the pixel that belongs to the foreground target with a greater probability than the probability that it belongs to the background image is updated to the first pixel value, and the pixel value of the pixel that belongs to the foreground target with a probability that is not greater than the probability that it belongs to the background image is updated to the second pixel value.

[0153] Based on the updated pixel values, determine the binary image of the orbital image;

[0154] In a binary image, determine the foreground target and its location.

[0155] In one possible implementation, the process of updating the pixel value of a pixel whose probability of belonging to the foreground object is greater than its probability of belonging to the background image to a first pixel value, and updating the pixel value of a pixel whose probability of belonging to the foreground object is no greater than its probability of belonging to the background image to a second pixel value, satisfies the following formula:

[0156] Where (x,y) represents the position coordinates of the pixel, f1(x,y) represents the probability that the pixel with position coordinates (x,y) in the first feature map belongs to the foreground target, f2(x,y) represents the probability that the pixel with position coordinates (x,y) in the second feature map belongs to the background image, a represents the first pixel value, b represents the second pixel value, and f(x,y) represents the updated pixel value of the pixel with position coordinates (x,y).

[0157] In one possible implementation, determining the foreground target and its location in a binary image includes:

[0158] The binary image is filtered to detect false positives and to identify the foreground target.

[0159] For the foreground target, determine the maximum bounding rectangle of the foreground target, and define the location of the foreground target as the maximum bounding rectangle.

[0160] In one possible implementation, it also includes:

[0161] For the track, Hough line detection is used to detect the track in the background image and the track image, the track is fitted to a straight line, and the area contained by the straight line of the track is determined as the area where the track is located.

[0162] The aforementioned computer-readable storage medium can be any available medium or data storage device that can be accessed by the processor in an electronic device, including but not limited to magnetic storage such as floppy disks, hard disks, magnetic tapes, MO (magneto-optical disks), optical storage such as CDs, DVDs, BDs, HVDs, etc., and semiconductor storage such as ROMs, EPROMs, EEPROMs, NAND flash (non-volatile memory), SSDs (solid-state drives), etc.

[0163] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0164] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to this application. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0165] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0166] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0167] Obviously, those skilled in the art can make various modifications and variations to this application without departing from the spirit and scope of this application. Therefore, if such modifications and variations fall within the scope of the claims of this application and their equivalents, this application also intends to include such modifications and variations.

Claims

1. A method for detecting foreign objects on a track, characterized in that, The method includes: For the track to be detected for foreign objects, a background image and an image of the track to be detected are acquired, wherein the background image includes an image of the track where no foreign objects are present; The background image and the track image are input into the foreground extraction model. Based on multiple first convolutional units in the foreground extraction model corresponding to the background image, multiple feature extractions are performed on the background image to obtain the feature map of the background image. Based on multiple second convolutional units in the foreground extraction model corresponding to the track image, multiple feature extractions are performed on the track image to obtain the feature map of the track image. In the foreground extraction model, for each background difference and upsampling, the feature map of the background image and the feature map of the orbit image are subtracted, the subtracted image is upsampled to obtain a first image, and the feature map output by the first convolutional unit of the previous layer is subtracted from the feature map output by the second convolutional unit of the previous layer to obtain a second image. The first image and the second image are stacked in the channel dimension, and the stacked image is input into the third convolutional unit to obtain the difference feature map. Based on the difference feature map, multiple background subtraction and upsampling are performed in the foreground extraction model. After the last background subtraction and upsampling, each pixel in the obtained difference feature map is classified by the third convolutional unit of the foreground extraction model to obtain a first feature map and a second feature map. The first feature map represents the probability that each pixel in the track image belongs to the foreground target, and the second feature map represents the probability that each pixel in the track image belongs to the background image. Based on the probability that each pixel in the orbit image belongs to the foreground target and the probability that it belongs to the background image, the foreground target and its location are determined. If the location of the foreground target is within the area where the track is located, then it is determined that there is a foreign object in the track.

2. The method according to claim 1, characterized in that, Determining the foreground target and its location based on the probability that each pixel in the orbital image belongs to the foreground target and the probability that it belongs to the background image, including: Based on the probability that each pixel in the orbit image belongs to the foreground target and the probability that it belongs to the background image, the pixel value of the pixel that has a greater probability of belonging to the foreground target than the probability of belonging to the background image is updated to the first pixel value, and the pixel value of the pixel that has a less than or equal probability of belonging to the foreground target than the probability of belonging to the background image is updated to the second pixel value. Based on the updated pixel values, determine the binary image of the orbital image; In the binary image, the foreground target in the binary image and the location of the foreground target are determined.

3. The method according to claim 2, characterized in that, The process of updating the pixel values ​​of pixels whose probability of belonging to the foreground target is greater than that of belonging to the background image to a first pixel value, and updating the pixel values ​​of pixels whose probability of belonging to the foreground target is not greater than that of belonging to the background image to a second pixel value, satisfies the following formula: ,in, Represents the position coordinates of a pixel. The position coordinates in the first feature map are: The probability that a pixel belongs to the foreground object. The position coordinates in the second feature map are: The probability that the pixel corresponding to the given pixel belongs to the background image. This represents the value of the first pixel. This represents the second pixel value. The position coordinates are The updated pixel values ​​of the pixels.

4. The method according to claim 2, characterized in that, In the binary image, determining the foreground target and its location in the binary image includes: The binary image is filtered for false positives to determine the foreground target; For the foreground target, determine the maximum bounding rectangle of the foreground target, and define the maximum bounding rectangle as the location of the foreground target.

5. The method according to claim 1, characterized in that, The method further includes: For the track, the track in the background image and the track image is detected by Hough line detection, the track is fitted to a straight line, and the area contained by the straight line of the track is determined as the area where the track is located.

6. A track foreign object detection device, characterized in that, The device includes: The acquisition module is used to acquire a background image and an image of the track to be detected for the track to be detected, wherein the background image includes an image of the track where no foreign object exists; The processing module is used to input the background image and the track image into a foreground extraction model. Based on multiple first convolutional units corresponding to the background image in the foreground extraction model, it performs multiple feature extractions on the background image to obtain a feature map of the background image. Similarly, based on multiple second convolutional units corresponding to the track image in the foreground extraction model, it performs multiple feature extractions on the track image to obtain a feature map of the track image. In the foreground extraction model, for each background differencing and upsampling step, it subtracts the feature map of the background image from the feature map of the track image, upsamples the subtracted image to obtain a first image, and then compares the feature map output by the previous first convolutional unit with... The feature maps output by the second convolutional unit in the previous layer are subtracted to obtain the second image. The first image and the second image are then stacked along the channel dimension. The stacked image is input into the third convolutional unit to obtain a difference feature map. Based on the difference feature map, multiple background subtraction and upsampling are performed in the foreground extraction model. After the last background subtraction and upsampling, each pixel in the obtained difference feature map is classified by the third convolutional unit of the foreground extraction model to obtain a first feature map and a second feature map. The first feature map represents the probability that each pixel in the track image belongs to the foreground target, and the second feature map represents the probability that each pixel in the track image belongs to the background image. The determination module is used to determine the foreground target and the location of the foreground target based on the probability that each pixel in the orbit image belongs to the foreground target and the probability that it belongs to the background image. The detection module is used to determine that there is a foreign object in the track if the location of the foreground target is within the area where the track is located.

7. An electronic device, characterized in that, The electronic device includes at least a processor and a memory, wherein the processor is used to execute a computer program stored in the memory to implement the steps of the orbital foreign object detection method as described in any one of claims 1-5.

8. A computer storage medium, characterized in that, It stores a computer program executable by an electronic device, which, when run on the electronic device, causes the electronic device to perform the steps of a foreign object detection method according to any one of claims 1-5.

Citation Information

Patent Citations

  • Image segmentation method and device, electronic equipment and readable storage medium

    CN111178211A

  • Underwater fish target detection method and device

    CN112528782A