A risk grading unmanned aerial vehicle railway foreign object intrusion open detection method

By using a railway foreign object intrusion detection method based on boundary-constrained MCP and open/closed set fusion, we have achieved multi-risk level region division and foreign object intrusion detection in UAV railway inspection images. This solves the problems of risk differentiation and category uncertainty in traditional methods and improves railway operation safety.

CN119693831BActive Publication Date: 2025-11-28BEIJING UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411862663.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-17
Publication Date
2025-11-28
Estimated Expiration
2044-12-17

AI Technical Summary

Technical Problem

Traditional methods for detecting foreign object intrusion in railways cannot distinguish the risk of foreign object intrusion at different locations, and cannot effectively address the diverse and uncertain nature of foreign object types in the real world, resulting in an inability to achieve effective risk perception and control.

Method used

A multi-risk level region segmentation method based on boundary-constrained MCP for UAV railway inspection images is adopted, combined with the FOCS-MID railway foreign object intrusion detection method architecture of open and closed set fusion. Through multi-level risk region segmentation and cross-modal feature fusion, the detection of both conventional and unconventional foreign objects is achieved.

Benefits of technology

It enables differentiated risk monitoring and control of foreign objects on railway lines, effectively detects routine foreign objects and responds to undefined categories of foreign objects, thereby improving the safety of railway operations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119693831B_ABST
    Figure CN119693831B_ABST
Patent Text Reader

Abstract

The application provides a risk grading unmanned aerial vehicle railway foreign matter invasion open detection method. The method determines the risk level of the invading foreign matter according to the transverse distance between the invading foreign matter and the core area of the track, and performs smoothing processing on the track area segmentation mask through a boundary limited MCP method, so that the parameter configurable unmanned aerial vehicle railway image multi-level risk area division is realized. The application further provides a railway foreign matter invasion detection method framework FOCS-MID under the open and closed set fusion condition. On the basis of the invasion detection of the conventional category foreign matter under the closed set condition, the secondary detection of the conventional and unconventional category foreign matter under the open set condition is simultaneously performed, so that the risk perception ability of the railway line foreign matter is comprehensively improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to a railway foreign matter intrusion detection method for unmanned aerial vehicle railway inspection, belonging to the field of rail transit operation safety and security. BACKGROUND

[0002] In recent years, unmanned aerial vehicle technology has developed rapidly and is widely used in industries such as railway, power, and photovoltaic. Unmanned aerial vehicles can perform aerial reconnaissance, logistics distribution, and inspection and monitoring. Edge computing, as a supplement to cloud computing, deploys data processing and analysis capabilities at the network edge, significantly reducing latency and improving response speed. The combination of unmanned aerial vehicle technology and edge computing enables unmanned aerial vehicles to process complex and multi-source data in real time, make autonomous decisions, and play a key role in real-time inspection and monitoring of railways. SUMMARY

[0003] Foreign matters existing within a certain distance range of the railway line and its surroundings may pose a serious threat to the safe operation of trains. In severe cases, it may cause train derailment and vehicle damage, and even endanger passengers' lives. Early detection and warning of intrusion foreign matters are critical to ensuring the safe operation of the railway system. Unmanned aerial vehicles have good maneuverability, diverse perspectives, and wide field of view, and can play an important role in railway foreign matter intrusion detection. Rapid and normalized railway foreign matter intrusion detection based on unmanned aerial vehicles has excellent application value in real-time monitoring and control of railway line foreign matter risks. However, traditional railway foreign matter intrusion detection methods do not distinguish the risks posed by intrusion foreign matters appearing in different locations, but simply treat all foreign matters or potential foreign matters as risk factors at the same risk level. At the same time, in the real-world open railway scene, foreign matter categories are diverse and uncertain, and it is impossible to completely define all intrusion foreign matter categories in advance, so it is also impossible to effectively deal with intrusion foreign matters of undefined categories.

[0004] To solve the above problems, the present application proposes a railway inspection image multi-risk level region division method based on boundary-restricted MCP for unmanned aerial vehicles, which determines the risk level of intrusion foreign matters according to their lateral distance from the core area of the track, and performs smoothing processing on the track area segmentation mask through the boundary-restricted MCP method, realizing parameter-configurable multi-level risk region division of unmanned aerial vehicle railway images. The present application further proposes a railway foreign matter intrusion detection method architecture FOCS-MID under open-closed set fusion conditions. Based on the intrusion detection of regular category foreign matters under closed set conditions, the method simultaneously performs secondary detection of regular and irregular category foreign matters under open set conditions, comprehensively improving the risk perception ability of railway line foreign matters.

[0005] The application proposes a brand-new unmanned aerial vehicle railway inspection image multi-risk level area division method based on boundary-restricted MCP, and a railway foreign matter intrusion open detection method framework FOCS-MID based on the risk area division, the overall operation framework of which is shown in Figure 1 Specifically includes the following two parts, which are described below in combination with the drawings:

[0006] (1) Boundary-restricted MCP multi-level risk area division method

[0007] Aiming at the characteristics that the risks caused by foreign matters in different areas of the unmanned aerial vehicle railway inspection image are different, a boundary-restricted minimum circumscribed parallelogram (MCP) risk area division method is proposed, which realizes multi-level risk area division with boundary smoothing based on track area segmentation results, and divides the image into alert areas, warning areas and attention areas.

[0008] (2) Foreign matter intrusion open detection method framework FOCS-MID

[0009] Aiming at the problem that in a real open-world railway scene, foreign matters are diverse and uncertain, and all categories of foreign matters cannot be completely defined in advance, a multi-risk level foreign matter intrusion detection method framework (multi-risk intrusion detection method formulation under fusion of open-closed sets, FOCS-MID) based on open-closed set fusion is proposed, which realizes multi-level risk area matching foreign matter intrusion detection, and widens the railway foreign matter intrusion detection capability to the open scene.

[0010] The method is realized in the following way:

[0011] According to the distance from the track core area, the multi-level risk area of the aerial view railway image of the unmanned aerial vehicle is divided into three levels: core alert area, adjacent warning area and peripheral attention area;

[0012] The core alert area should only focus on the maximum outer boundary area through which the train runs. The maximum connected domain of the core railway area segmentation mask predicted by the depth segmentation network is represented as Where j represents the number of the maximum connected domain in the image; for the single track core area maximum connected domain parsed by the network Draw A horizontal or vertical line is used to sample the connected region, corresponding to the two cases where railway lines in the image are distributed laterally and longitudinally; this is illustrated by drawing equally spaced horizontal sampling lines along the longitudinal direction of the railway lines; for each horizontal sampling line, The midpoint of the two points that coincide with the horizontal line and are farthest apart is denoted as (x i ,y i Let d be the distance between these two points. i This allows for the longitudinal and uniform sampling of the centerline of the core orbital region. There are points, which are represented here as... Then use this Performing simple linear regression on these points yields the equation representing the central line of the railway area, and this... The center point of each point is represented as (x o ,y o ),in:

[0013]

[0014] At the same time, samples were also obtained. Maximum horizontal distance value Here, its maximum value is defined as the pixel width of the orbital core region, denoted as d:

[0015]

[0016] This The equation of the center line l of the orbital core region obtained by performing simple linear regression on the two-dimensional plane is as follows:

[0017] f(x,y)=ax+by+c=0 (3)

[0018] Translating this straight line along the positive x-axis by a distance of d / 2, we obtain the equation of the right boundary line l1 of the orbital core region in a two-dimensional plane, which is expressed as:

[0019] f1(x,y)=a(xd / 2)+by+c=0 (4)

[0020] Translating line l along the negative x-axis by a distance d / 2, we obtain the equation of line l2, the left boundary of the core region of the orbit, in the two-dimensional plane, which is expressed as:

[0021] f2(x,y)=a(x+d / 2)+by+c=0 (5)

[0022] Here, we first present the representation of the same-side and opposite-side half-plane regions of a point outside a line; In a two-dimensional plane, there is a line l with the analytical equation f(x,y)=ax+by+c=0. Now, there is a point P outside line l with coordinates (x,y)...o y o ), satisfying ax o +by o +c≠0; the line l divides the plane into two half-plane regions A and B; set A region and point P are located on the opposite side of the line l, B region and point P are located on the same side of the line l, then the A region half plane and B region half plane are represented by the following inequalities respectively:

[0023]

[0024] Both of the two half planes do not contain its boundary line l, where “·” represents the product operation, the same below; sgn(·) is the sign function, which is defined as follows:

[0025]

[0026] The track core warning area range is determined by the following plane constraints:

[0027]

[0028] Where w and h are the pixel width and height of the image; in this way, the maximum connected domain of the segmented track core area is further normalized to a region as given in equation (8) Most of the time this region is a parallelogram, two boundary lines and the upper and lower boundaries of the image together constitute a railway area with smooth boundary, which is the minimum circumscribed parallelogram MCP belonging to ;

[0029] Considering that the width of the adjacent warning area may not be completely consistent in different railway sections, here the horizontal width of the adjacent warning area on both sides of the railway line is defined as λd, and λ is uniformly set to 1, that is, the adjacent warning area and the core area have the same width; therefore, continue to move the track core warning area boundary lines l1 and l2 to the right and left by λd length respectively, to obtain the two boundary lines l3 and l4 of the adjacent warning area of the railway;

[0030] The equations of the boundary straight lines l3 and l4 of the adjacent warning area on both sides of the line are respectively:

[0031] f3(x,y)=a(x-d / 2-λd)+by+c=0 (9)

[0032] f4(x,y)=a(x+d / 2+λd)+by+c=0 (10)

[0033] The track core area Right adjacent warning area of the railway is represented by the following constraints on the two-dimensional plane:

[0034]

[0035] The track core area The left railway adjacent warning area is represented by the following constraints in two-dimensional plane:

[0036]

[0037] Thus, the adjacent area of the railway line is represented as:

[0038]

[0039] The set of all pixel points in the entire image range is defined as The peripheral attention area is calculated by the following formula:

[0040]

[0041] The above gives the method of multi-level risk area division when processing a single track line;

[0042] The UAV image risk area division: first, the multi-level division result corresponding to each track line needs to be solved, taking two track lines as an example; the area division result corresponding to track line k is set as The final image multi-level risk area division is:

[0043]

[0044] Where ∪ k (·) represents the union operation of multiple sets, ∩ represents the intersection operation of two sets, represents the complement operation of ;

[0045] In fact, from the perspective of image pixel points, risk area division is also a process of assigning three different attributes to each pixel point in the image; therefore, the region attribute of pixel point p in image is set as t, where the value of t is 0, 1 and 2, and respectively corresponds to the core warning area, adjacent warning area and peripheral attention area three attributes; it is set that there are n track lines in the image, and the region attribute of pixel point p in the area division corresponding to each track line is respectively t p,1 , t p,2 , …, t p,n , then the risk area division of the entire image can be completed by determining the final region attribute of pixel point p, and the formula equivalent to formula (15) is given as follows:

[0046]

[0047] The three stages are: the first stage mainly completes the multi-level risk area division; the second stage carries out the detection of the invasion of the conventional category of foreign matters under the closed set condition; the third stage carries out the detection of the invasion of the conventional and unconventional categories of foreign matters under the open set condition, which will be described in the following stages;

[0048] In the first stage, the unmanned aerial vehicle remote sensing railway image to be processed is first input into a semantic segmentation model encoder for image feature extraction, the extracted features are input into a model decoder for prediction to generate high-level semantic features at the pixel level, and after the track area division is completed, a track area segmentation mask is obtained

[0049]

[0050] The track area mask is traversed The largest connected domain is obtained

[0051] Then, the track area is smoothed by using the boundary-restricted MCP risk area method to obtain the final track core area; then, the boundary division of the transversely configurable adjacent area is completed, and finally, the risk area division mask image at the pixel level of the whole image is obtained

[0052] In the second stage, the unmanned aerial vehicle remote sensing railway image to be detected is input into a target detector backbone network for image multi-scale feature extraction, and different scale detection heads are defined on the image multi-scale feature maps to complete the position regression of the boundary box of the foreign matter target of different scales and the prediction of the predefined category, and finally, a relatively dense boundary box prediction is generated, and M targets are set , which represents the mth target foreign matter detected by the closed set detection model:

[0053]

[0054] In the third stage, in addition to inputting the image into the detector backbone network to generate the image multi-scale feature map, the user also needs to input the conventional foreign matter semantic vocabulary and the unconventional foreign matter semantic vocabulary into the pre-trained text encoder to generate text word embedding, and then perform multi-level cross-modal feature fusion on the visual-linguistic features, and then match the text semantics and the position area, and generate a dense target foreign matter boundary box prediction together with the multi-scale boundary box obtained by regression, and N targets are set , which represents the nth target foreign matter detected by the closed set detection model:

[0055]

[0056] After the above three stages are completed, the opening and closing set foreign object detection bounding box dense prediction results are fused and a non-maximum suppression (NMS) operation process is performed, redundant prediction targets are eliminated, and finally Q foreign objects are set, and obj q represents the nth target foreign object after non-maximum suppression:

[0057]

[0058] Finally, the target position attribute matching of the foreign object is performed on the obtained pixel-level risk area attribute mask map, and the multi-level risk matching foreign object intrusion detection work is finally completed; for each foreign object obj q , the center position pixel point Loc(obj q ) = (x q , y q ) is matched with the pixel-level risk area division mask map .

[0059]

[0060] Wherein Risk(obj q ) represents the risk level of the target foreign object obj q , and according to the different regions where the center position is located, the risk level is divided into three levels of high, medium and low.

[0061] The advantages and positive effects of the present application are:

[0062] (1) A boundary-restricted MCP multi-level risk area division method is designed, which can realize the smooth risk multi-level area division of the unmanned aerial vehicle railway inspection image boundary, thereby realizing the railway intrusion foreign object differentiated risk area matching, and realizing better railway line foreign object risk monitoring and control;

[0063] (2) A railway foreign object intrusion open detection method architecture is designed, which can not only effectively detect the foreign objects of the conventional pre-defined categories when facing the problem of multiple categories and uncertainty of foreign objects in the real open world railway scene, but also use the zero-shot detection capability of large-scale pre-training mode to cope with the foreign objects of unconventional undefined categories. BRIEF DESCRIPTION OF DRAWINGS

[0064] Figure 1 is the overall architecture of the unmanned aerial vehicle railway foreign object intrusion open detection

[0065] Figure 2 is a multi-level risk area division diagram of an unmanned aerial vehicle railway inspection image

[0066] Figure 3Same side and opposite side area representation method of straight line outer point

[0067] Figure 4 Train running track core warning area range determination

[0068] Figure 5 Train running track adjacent early warning area range determination

[0069] Figure 6 Multi-track line multi-level risk area division method

[0070] Figure 7 FOCS-MID multi-risk level railway foreign object intrusion open detection method architecture

[0071] Figure 8 Multi-level risk area division experimental results on FWA dataset

[0072] Figure 9 Multi-level risk area division experimental results on MWB dataset

[0073] Figure 10 Foreign object target detection partial results under closed set conditions

[0074] Figure 11 Foreign object target detection partial results under open and closed set conditions

[0075] Figure 12 Multi-level railway intrusion detection result visualization example DETAILED DESCRIPTION

[0076] The application will be further described in detail below in combination with the drawings and examples.

[0077] The application proposes a brand-new multi-risk level area division method for unmanned aerial vehicle railway inspection images based on boundary-restricted MCP and a railway foreign object intrusion open detection method architecture FOCS-MID. The details are as follows:

[0078] (1) Multi-risk level area division method

[0079] Generally, the closer to the core area of the track line where the train runs, the higher the risk degree of the intrusion foreign object to the train operation. Therefore, based on this idea, the multi-level risk area of the unmanned aerial vehicle overhead railway image is divided according to the distance from the track core area, which is mainly divided into three levels: core warning area, adjacent early warning area and peripheral attention area, as shown in detail in Figure 2

[0080] ​Generally speaking, the core warning zone should focus solely on the outermost boundary area traversed by trains. Therefore, theoretically, the delineation of warning zones, alert zones, and areas of concern should differ for single-track and double-track railway area images. The delineation methods for risk zones on single-track and double-track railways are provided below. Figure 2 As shown in the left and right images.

[0081] Next, we introduce a multi-level risk region segmentation method based on deep semantic segmentation networks. The orbital regions resolved using deep networks typically have relatively vague and coarse boundaries, which is not conducive to risk region segmentation based on distance metrics. Therefore, the boundaries of the final orbital core region first need to be smoothed.

[0082] First, the maximum connected component of the railway region segmentation mask obtained from the initial prediction of the deep segmentation network is represented as: Where j represents the number (or index) of the largest connected component in the image. This refers to the largest connected component in a single orbital core region parsed by the network. Draw at equal intervals along the vertical or horizontal direction. A horizontal or vertical line is used to sample the connected region, corresponding to the two cases where railway lines in the image are distributed laterally or longitudinally. The following example illustrates this: railway lines are distributed longitudinally, and equally spaced horizontal sampling lines are drawn along the longitudinal direction. For each horizontal sampling line, The midpoint of the two points that coincide with the horizontal line and are farthest apart is denoted as (x i ,y i Let d be the distance between these two points. i This allows for the longitudinal and uniform sampling of the centerline of the core orbital region. There are points, which are represented here as... Then use this Performing simple linear regression on each point yields the equation representing the central line of the railway area, and then... The center point of each point is represented as (x o ,y o ),in:

[0083]

[0084] Simultaneously obtained samples Maximum horizontal distance value Here, its maximum value is defined as the pixel width of the orbital core region, denoted as d:

[0085]

[0086] This The equation of the center line l of the orbital core region obtained by performing simple linear regression on the two-dimensional plane is as follows:

[0087] f(x,y)=ax+by+c=0 (3)

[0088] Furthermore, by translating this line along the positive x-axis by a distance of d / 2, we obtain the equation of the right boundary line l1 of the orbital core region in the two-dimensional plane, which is expressed as:

[0089] f1(x,y)=a(xd / 2)+by+c=0 (4)

[0090] Similarly, translating line l along the negative x-axis by a distance d / 2 yields the equation of line l2, the left boundary of the orbital core region, in the two-dimensional plane, which is expressed as:

[0091] f2(x,y)=a(x+d / 2)+by+c=0 (5)

[0092] This section first presents the methods for representing the same-side and opposite-side half-plane regions of a point outside a line. For example... Figure 3 As shown, there is a straight line l in a two-dimensional plane, whose analytical equation is f(x,y)=ax+by+c=0. There is a point P outside the straight line l, whose coordinates are (x,y) / (x+y) / (y+c)=0. o ,y o ), satisfying ax o +by o +c≠0. Line l divides the plane into two half-plane regions, A and B. Assuming region A is on the opposite side of line l from point P, and region B is on the same side of line l from point P, then the half-planes of region A and region B can be represented by the following inequalities:

[0093]

[0094] Neither of the two half-planes contains its boundary line l, where “·” represents the product operation, and the same applies below. sgn(·) is the sign function, specifically defined as:

[0095]

[0096] Therefore, the extent of the core warning zone is determined by the following planar constraints, such as... Figure 4 As shown:

[0097]

[0098] Where w and h are the pixel width and height of the image. Thus, the maximum connected component of the segmented orbital core region... Further normalization into a region as given by formula (8) In most cases, this region is a parallelogram. The two boundary lines, together with the upper and lower boundaries of the image, form a smoothly bordered railway region, which belongs to... the minimum circumscribed parallelogram (MCP).

[0099] Considering that the width of the adjacent warning area can not be exactly the same in different railway sections, the horizontal width of the adjacent warning area on both sides of the railway line is defined as λd, where λ is an adjustable parameter to represent the proportion of the width of the adjacent warning area relative to the width of the track core area. In this patent, λ is set to 1, i.e., the adjacent warning area and the core area have the same width. Therefore, the track core warning area boundary lines l1 and l2 are translated right and left by λd, respectively, to obtain the two boundary lines l3 and l4 of the adjacent warning area of the railway. Similarly, the equations of the adjacent warning area boundary straight lines l3 and l4 on both sides of the line are:

[0100] f3(x, y) = a(x - d / 2 - λd) + by + c = 0 (9)

[0101] f4(x, y) = a(x + d / 2 + λd) + by + c = 0 (10)

[0102] Similarly, the track core area the adjacent warning area on the right side of the railway is represented by the following constraints in the two-dimensional plane:

[0103]

[0104] Similarly, the track core area the adjacent warning area on the left side of the railway is represented by the following constraints in the two-dimensional plane:

[0105]

[0106] Thus, the adjacent area of the railway line on both sides is represented as:

[0107]

[0108] The set of all pixel points in the entire image range is defined as the peripheral attention area which is calculated by the following formula:

[0109]

[0110] The above gives the method of multi-level risk area division when processing a single track line. Further, the unmanned aerial vehicle image risk area division method with two or more track lines is given. First, the multi-level division result corresponding to each track line needs to be solved, which is illustrated by taking two track lines as an example. Assuming that the area division result corresponding to track line k is The final image multi-level risk area division is:

[0111]

[0112] Where ∪ k (·) represents the union operation of multiple sets, ∩ represents the intersection operation of two sets, represents the complement operation of . As shown in Figure 6 , the multi-level risk area division process under the double-track railway is given.

[0113] In fact, from the perspective of image pixels, risk area division is also a process of assigning three different attributes to each pixel in the image. Therefore, assume that the region attribute of pixel p in the image is t, where t takes values of 0, 1 and 2, and corresponds to the three attributes of core alert zone, adjacent early warning zone and peripheral attention zone. Assuming that there are n track lines in the image, the region attribute of pixel p in the region division corresponding to each track line is t p,1 , t p,2 , …, t p,n , then the risk area division of the whole image can be completed by determining the final region attribute of pixel p, and the formula equivalent to formula (15) can be given as follows:

[0114]

[0115] (2) Open detection method architecture of foreign matter intrusion

[0116] This study constructs the risk differentiation foreign matter intrusion detection architecture FOCS-MID based on open and closed set fusion. The method architecture is shown in Figure 7 , which is mainly divided into three stages: (1) the first stage mainly completes the multi-level risk area division; (2) the second stage performs regular category foreign matter intrusion detection under closed set conditions; (3) the third stage performs regular and irregular category foreign matter intrusion detection under open set conditions, which is described in stages as follows.

[0117] In stage one, the unmanned aerial vehicle remote sensing railway image to be processed is first input into the semantic segmentation model encoder for image feature extraction, and the extracted features are input into the model decoder to predict the pixel-level high-level semantic features, and the track area segmentation mask

[0118]

[0119] Traversing the orbital region mask By analyzing each connected component in the given data and identifying the one with the largest area, we obtain the largest connected component. Then, the orbital region is smoothed using the boundary-constrained MCP risk region method to obtain the final orbital core region; next, the lateral width of the neighboring region boundary is divided, and finally, a full-image pixel-level risk region segmentation mask map is obtained.

[0120]

[0121] In stage two, the remote sensing railway image of the UAV to be inspected is input into the backbone network of the target detector for multi-scale feature extraction. Different scale detection heads are defined on the multi-scale feature map to complete the location regression and predefined category prediction of the bounding boxes of foreign targets at different scales, ultimately generating relatively dense bounding box predictions. Assuming there are M targets, [the following steps are taken]. This represents the m-th foreign object detected by the closed-set detection model.

[0122]

[0123] In stage three, in addition to inputting the image into the detector backbone network to generate multi-scale feature maps, the user also needs to input conventional and unconventional foreign object semantic terms (or categories) into a pre-trained text encoder (usually a large-scale text-image pre-trained text encoder such as CLIP) to generate text word embeddings. Then, visual-linguistic features are fused across multiple modalities. Next, text semantics are matched with location regions, and these are combined with the regressed multi-scale bounding boxes to generate dense target-foreign object bounding box predictions. Assuming there are N targets, using... This represents the nth foreign object detected by the closed-set detection model.

[0124]

[0125] After completing the above three stages, the dense prediction results of the open / closed set foreign object detection bounding boxes are fused and subjected to non-maximum suppression (NMS) to eliminate redundant predicted targets. Assuming there are ultimately Q foreign objects, represented by obj... q This represents the nth target foreign object after nonmaximum suppression:

[0126]

[0127] Finally, the target position attribute matching of foreign matter is performed on the obtained pixel-level risk area attribute mask map, and the foreign matter intrusion detection work of multi-level risk matching is finally completed. For each foreign matter target obj q , the center position pixel point Loc(obj q ) = (x q , y q ) is matched with the pixel-level risk area division mask map .

[0128]

[0129] Wherein Risk(obj q ) represents the risk level of the target foreign matter obj q , and according to the different regions where the center position is located, the risk level is divided into three levels of high (high), medium (medium) and low (low).

[0130] In this example, the track line scene data is collected by using unmanned aerial vehicles in two different railway sections in Maanshan, and the corresponding railway / track area segmentation and closed set condition foreign matter detection data set is constructed based on the data, as shown in Table 1. The data set composed of data collected by fixed-wing and multi-rotor unmanned aerial vehicles is called FWA and MRB respectively. In FWA, the segmentation data set contains 1602 images, of which 1442 are used for training and 160 are used for testing. The data set has two classes, representing railway area and non-railway area respectively; the closed set detection data set contains 1123 images, of which 1010 are used for training and 113 are used for testing. The data set only involves one category, i.e. person category. Because the fixed-wing unmanned aerial vehicle has a relatively high flight height and a relatively fast flight speed, it is affected by the weather environment, and the collected image quality is poor and the definition is limited. At the same time, considering that the current open target detection model capability is still limited, the FWA data set is not used for open set foreign matter intrusion recognition work. In MRB, the segmentation data set contains 1255 images, of which 1004 are used for training and 251 are used for testing. The data set has two classes, representing track area and non-track area respectively; the closed set detection data set contains 770 images, of which 693 are used for training and 77 are used for testing. The data set also only involves one category, i.e. person category, and the images in the data set are used for foreign matter recognition in open scene.

[0131] Table 1 Statistics of data set constructed by images collected by unmanned aerial vehicles

[0132]

[0133] Wherein, due to the constraints of existing conditions, the number of alien invasion image data samples that can be collected is relatively small, and the number of alien categories in the collected images is very small. Therefore, in closed set detection, the conventional alien category is set as a single category, i.e. {person}, and in open scene alien detection, the unconventional alien category outside the conventional category provided to the model is set as {vehicle, stone}.

[0134] • Multi-risk level region division experimental results

[0135] Experiments were conducted on two split data sets of FWA and MRB, and multi-level track line risk region division was completed, and the results are shown in Figure 8 and Figure 9 respectively. Among them, in the FWA data set, the entire railway region is divided as the core alert area, and the experimental results show that the segmentation accuracy mIoU of the trained deep segmentation network for the railway region is 0.962. As can be seen from the figure, under the condition that the railway region presents different distribution angles and transverse widths, the proposed multi-level risk region division method based on boundary restricted MCP can achieve efficient and accurate risk region division. When part of the model segmentation results are rough due to tree or house obstruction, the proposed risk region division method can effectively cope with it, and according to the distribution characteristics of the railway line image from the unmanned aerial vehicle perspective, the through railway core region division from the image boundary to the boundary can be realized, and further the pixel-level risk degree determination of the whole image can be completed. In this experiment, the division of the entire railway region as the core alert area of the unmanned aerial vehicle image can be used for alien invasion detection tasks with railway gauge as the boundary.

[0136] At the same time, in the MRB data set, the track region is divided as the core alert area. The experimental results show that the segmentation accuracy MIoU of the track region is 0.974. More examples of rough edges or truncated sections of model segmentation results due to tree or vehicle obstruction are shown in Figure 9 . In the experimental results, examples of region division under the condition that the track line presents multiple different angles and different transverse widths are also given, and it can be seen that the region division results of the whole image show good robustness.

[0137] • Open and closed set fusion alien invasion detection experimental results

[0138] The alien invasion detection under closed set condition is carried out on the FWA closed set alien (single category alien person) data set constructed in the present application, and the alien (mainly person, vehicle, stone, etc.) invasion detection under open scene condition is carried out on the MRB data set.

[0139] (1) Closed set alien invasion detection

[0140] This section gives the potential pedestrian foreign object detection experimental results in the full image range using two target detection networks. The relevant quantitative indicators are shown in Table 2, which gives the AP-related indicator values on the FWA and MWB two closed set foreign object intrusion detection data sets. The P 50 , R 50 and mAP 50 of network 1 are as high as 0.822, 0.920 and 0.921 respectively, and the mAP value is 0.409, and the indicator values on the MRB data set are as high as 0.949, 0.895 and 0.983 respectively, and the mAP value is 0.567. This fully illustrates the feasibility of closed set detection under the current data set, and the higher prediction accuracy also shows the necessity of closed set detection for common categories of foreign objects.

[0141] At the same time, network 2 is further used to test and evaluate on the two data sets, and the results are shown in Table 2. It can be seen that the AP-related indicator values of this model on the two data sets under the closed set foreign object detection setting are relatively high, and compared with the indicator values of network 1 model, the overall has a certain improvement, but the precision improvement is not significant, which also shows that the current general target detection model can actually achieve good performance on closed set foreign object intrusion detection, and the performance difference between different algorithms is limited.

[0142] Some closed set foreign object detection image examples are given in Figure 10 , it can be seen that many small target pedestrians in the figure, including pedestrians riding electric bikes or motorcycles, most of them can be detected, but there is still a small part of the image that will exist in the case of false detection or missed detection, that is, the FP and FN cases. As shown in the red box example in the figure, the FN case shows that the pedestrian appearing in the railway core warning area cannot be detected, and the FP case mistakenly detects other unknown objects as pedestrian foreign object targets.

[0143] Table 2 AP indicator values of foreign object intrusion detection

[0144]

[0145] (2) Open set foreign object intrusion detection

[0146] Further closed set and open set foreign object intrusion detection is carried out on the MRB data set, and some recognition results are given as Figure 11As shown in columns 2 and 5 of the figure, it can be seen from all the examples that the foreign object detection model trained under closed-set conditions can only detect and identify predefined foreign object targets, namely <person>, and has a high prediction confidence and low false alarm rate. However, it basically does not respond to other categories outside the predefined categories, such as motorcycles and cars, and therefore cannot adapt to the requirements of foreign object recognition in real-world open scenes. Simultaneously, zero-shot detection was also performed using a model pre-trained on a large-scale dataset without fine-tuning on any UAV image dataset. In addition to the previously defined closed-set category <person>, <car> and <motorcycle> were added to evaluate the zero-shot generalization ability of foreign object detection in open scenes. Relevant image examples are shown in columns 3 and 6 of the figure. As can be seen from the examples, it can not only identify pedestrians in the images (as shown in some images in columns 3 and 6), but also identify cars and motorcycles on the road next to the railway (as shown in rows 1 and 2, and columns 3 and 6). This also demonstrates that open-set detectors trained on large-scale visual-language datasets possess strong zero-shot generalization capabilities, maintaining considerable recognition ability even when dealing with categories other than those specified in the closed set predefined categories. However, it also shows that pedestrians on railway tracks in some images cannot be completely detected (as shown in images 3 row 3, 1 row 2, 3 row 6, etc.). This indicates that relying solely on open-set detectors that haven't been fine-tuned on relevant datasets from a drone perspective cannot achieve high-confidence and high-precision detection of common categories of foreign objects, but can serve as an effective supplement to foreign object detection under closed-set conditions. From the above results, we can conclude that...

[0147] When the foreign object recognition results under closed and open set conditions are fused and further NMS operations are performed, richer foreign object perception and recognition results in railway scenes from the perspective of UAVs are obtained. Different confidence thresholds are required for closed set and open set detection due to their different scene settings.

[0148] (3) Foreign object intrusion detection with risk area matching

[0149] After completing the open-closed set fusion detection of potential foreign objects in the entire image, multi-level risk regions can be divided and all potential foreign objects can be matched based on their spatial location and risk attribute results, thereby completing the railway foreign object intrusion detection through risk region matching. For example... Figure 12 As shown, a possible representation of foreign object detection results based on risk area matching is presented. Detected potential foreign objects are represented by solid circles of different colors to assess the risk of foreign object intrusion in railway scenarios. Red, yellow, and green represent high, medium, and low risk, respectively. For image clarity, the outer regions of interest are not drawn with obvious colors.

[0150] In summary, it can be seen that the method of the present application can efficiently realize the foreign matter intrusion detection of the unmanned aerial vehicle railway inspection image, and expand the detection capability from the predefined category under the closed set condition to the open scene. The present application provides a general real-world open scene foreign matter intrusion detection mode for unmanned aerial vehicle-based railway line inspection, which has a wide application prospect for unmanned aerial vehicle-based railway inspection. The architecture can not only be used for foreign matter intrusion detection in the open scene of the railway, but also be used in other open scenes. The core feature of such an open scene is that the predefined target category is relatively concentrated, but other uncertain target categories cannot be completely defined in advance. Using the method of the present application, the predefined category and precision recognition, and the effective detection of undefined category strong generalization zero sample can be effectively realized.

Claims

1. A risk classification method for detecting open railway intrusion by unmanned aerial vehicles, characterized by Comprise the following two parts: (1) Multi-level risk area division method Aiming at the characteristics that the risk caused by foreign matter in different areas of the unmanned aerial vehicle railway inspection image is not the same, a boundary-constrained minimum circumscribed parallelogram risk zoning method is proposed. Based on the track area segmentation result, the multi-level risk area division with smooth boundary is realized, and the image is divided into alert area, early warning area and attention area; (2) Foreign matter intrusion open detection method architecture Aiming at the diversity and uncertainty of foreign matter in real open world railway scene, a multi-risk level foreign matter intrusion detection method architecture based on open and closed set fusion is proposed, which realizes the multi-level risk area matching foreign matter intrusion detection, and widens the railway foreign matter intrusion detection capability to the open scene; The method is divided into three stages: stage one completes the multi-level risk area division; Stage two carries out the regular category foreign matter intrusion detection under the closed set condition; Stage three carries out the regular and irregular category foreign matter intrusion detection under the open set condition, which is described as follows; In stage one, the unmanned aerial vehicle remote sensing railway image to be processed is first input into a semantic segmentation model encoder for image feature extraction, the extracted features are input into a model decoder to predict and generate pixel-level high-level semantic features, and after track area segmentation, a track area segmentation mask A is obtained rail,j : {A rail,j}=DeepSeg(image) (17) Traverse the track area mask A rail,j The largest connected domain in the area of the largest connected domain Then the track area is smoothed by the risk area method of the boundary limited MCP to obtain the final track core area; then the boundary division of the transverse width configurable adjacent area is completed, and finally the risk area division mask image graph {A core ,A adj ,A per} is obtained. In stage two, the unmanned aerial vehicle remote sensing railway image to be detected is input to the target detector backbone network for image multi-scale feature extraction, and different scale detection heads are defined on the image multi-scale feature maps to complete the position regression of the boundary box of the foreign object of different scales and the prediction of the predefined class, and finally a relatively dense boundary box prediction is generated, and M targets are set, and the mth target foreign object detected by the closed set detection model is represented as ​ In stage three, in addition to inputting the image into the detector backbone network to generate the image multi-scale feature map, the user also needs to input the conventional foreign object semantic vocabulary and the unconventional foreign object semantic vocabulary into the pre-trained text encoder to generate text word embedding, and perform multi-level cross-modal feature fusion on the visual-linguistic features, then match the text semantics and the location area, and jointly generate the dense target foreign object bounding box prediction with the regression obtained multi-scale bounding box, set N targets, and represent the nth target foreign object detected by the closed set detection model: After the above three stages are completed, the open-close set foreign object detection bounding box dense prediction results are fused and a non-maximum suppression (NMS) operation is performed, redundant prediction targets are eliminated, and finally Q foreign objects are set, and obj q represents the nth target foreign object after non-maximum suppression. Finally, the target position attribute matching of the foreign matter is performed on the obtained pixel-level risk area attribute mask map, and the foreign matter intrusion detection work of multi-level risk matching is finally completed; for each foreign matter target obj q , the center position pixel point Loc(obj q ) = (x q , y q ) is matched with the pixel-level risk area division mask map {A core , A adj , A per}. where Risk(obj q ) represents the risk level of the target foreign object obj q , and is divided into three levels of high, medium and low according to the area where the center position is located.

2. The method of claim 1, wherein The method is realized by the following way; According to the distance from the track core area, the unmanned aerial vehicle overhead railway image is divided into multi-level risk area, which is divided into: core alert area, adjacent early warning area and peripheral attention area three levels; The core warning area only focuses on the maximum outer boundary area through which the train runs; the maximum connected domain of the core railway area segmentation mask obtained by the deep segmentation network is represented as where j represents the number of the maximum connected domain in the image; the maximum connected domain of the single track core area analyzed by the network is represented as N equally spaced horizontal or vertical lines are drawn along the longitudinal or transverse direction to sample the connected domain, corresponding to the two cases of horizontal and vertical distribution of the railway line in the image; the case of longitudinal distribution of the railway line is illustrated; for each horizontal sampling line, the midpoint of the two points farthest away from the horizontal line is represented as (x i ,y i ), and the distance between the two points is represented as d i ; thus, the center line of the track core area is longitudinally uniformly sampled into N points, which are represented as {(x1,y1),...,(x N ,y N )}; then, a simple linear regression is performed using the N points to obtain the equation of the center straight line of the railway area, and the center point of the N points is represented as (x o ,y o ), wherein: At the same time, N horizontal maximum distance values {d1, d2,..., d N} of the samples are obtained, and the maximum value of the N horizontal maximum distance values is defined as the pixel width of the track core region, denoted as d: d = max {d1, d2,..., d N} (2) The equation of the track core area center straight line l obtained by performing simple linear regression on the N points in the two-dimensional plane is expressed as: f(x,y)=ax+by+c=0 (3) The right boundary straight line l1 of the track core area is obtained by translating the straight line along the positive direction of x axis by d / 2 distance, and the equation of the straight line in the two-dimensional plane is expressed as: f1(x,y)=a(x-d / 2)+by+c=0 (4) The left boundary straight line l2 of the track core area is obtained by translating the straight line along the negative direction of x axis by d / 2 distance, and the equation of the straight line in the two-dimensional plane is expressed as: f2(x,y)=a(x+d / 2)+by+c=0 (5) Here, we first present the representation of the same-side and opposite-side half-plane regions of a point outside a line; In a two-dimensional plane, there is a line l with the analytical equation f(x,y)=ax+by+c=0. Now, there is a point P outside line l with coordinates (x,y)... o ,y o ), satisfying ax o +by o +c≠0; Line l divides the plane into two half-plane regions, A and B; assuming region A and point P are on opposite sides of line l, and region B and point P are on the same side of line l, then the half-planes of region A and region B can be represented by the following inequalities: Here, the two half planes do not contain the dividing line l, where "·" represents the product operation, and the same in the following; The sign function sgn(·) is defined as: The range of the track core alert area is determined by the following plane constraint: where w and h are the width and height of the image in pixels; and the maximum connected component of the segmented track core region is further regularized to a region A as given in equation (8) core,j ; most of the time this region is a parallelogram whose two boundary lines together with the upper and lower image boundaries form a railway region with smooth boundaries, which is the minimum circumscribed parallelogram MCP whose boundaries are confined to ​ Considering that the width of the adjacent early warning area may not be completely consistent in different railway sections, the horizontal width of the adjacent early warning area on both sides of the railway line is defined as λd, and λ is uniformly set to 1, that is, the adjacent early warning area has the same width as the core area; Therefore, the boundary lines l1 and l2 of the track core alert area are respectively translated to the right and left by λd length, and the two boundary lines l3 and l4 of the adjacent early warning area of the railway are obtained; The equations of the adjacent early warning area boundary straight lines l3 and l4 on both sides of the line are respectively: f3(x,y)=a(x-d / 2-λd)+by+c=0 (9) f4(x,y)=a(x+d / 2+λd)+by+c=0 (10) The track core area A core,j The right railway approach warning area is represented by the following constraints on a two-dimensional plane: The track core area A core,j The left railway near warning area is expressed by the following constraints on the two-dimensional plane: Thus, the neighborhood area A on both sides of the railway line adj is represented as: The set of all pixel points in the entire image range is defined as I = {(x, y) | 0 < x < w, 0 < y < h}; the peripheral attention region A per This is calculated by the following formula: I = A core,j + A adj,j + A per,j (14) The above gives the method of multi-level risk area division when processing a single track line; Unmanned aerial vehicle image risk area division: first, the corresponding multi-level division result of each track line needs to be solved, taking two track lines as an example; the area division result corresponding to track line k is set as {A c,k ,A a,k ,A p,k}, and the final image multi-level risk area division is: where U k (·) denotes a union operation on multiple sets, I denotes an intersection operation on two sets, denotes a complement operation on A core ; In fact, from the perspective of image pixels, the risk area division is also a process of assigning three different attributes to each pixel in the image; therefore, the area attribute of pixel p in image I is set as t, where t takes values of 0, 1 and 2, and respectively corresponds to the three attributes of core alert area, adjacent pre-warning area and peripheral attention area; it is set that there are n track lines in the image, and the area attribute of pixel p in the area division corresponding to each track line is respectively t p,1 , t p,2 , …, t p,n , then the risk area division of the whole image can be completed by determining the final area attribute of pixel p, and the formula equivalent to formula (15) is given as follows: t p = min{t p,1 , t p,2 ,..., t p,n}, p∈I (16).

Citation Information

Patent Citations

  • Rail foreign matter detection method and system under space-based visual angle based on weak supervised learning

    CN111582084A

  • Railway intrusion foreign matter unmanned aerial vehicle detection method, device and system based on deep learning

    CN114248819A