Method and System for Identifying Anti-Bird Spikes of High-Speed Railway Catenary Based on Relative Position Perception

Through the target detection model based on relative position perception, the relative position relationship between the support device and the bird spine prevention is optimized, and the accuracy and efficiency of bird spine identification during high-speed railway contact network patrol is solved, and efficient and accurate bird spine prevention detection is achieved.

CN115511876BActive Publication Date: 2025-07-11BEIJING JIAOTONG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202211355480.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-01
Publication Date
2025-07-11
Estimated Expiration
2042-11-01

AI Technical Summary

Technical Problem

The prior art is difficult to efficiently identify bird stabs in high-speed railway contact network patrols, especially in small target detection and bad weather conditions, resulting in low detection accuracy, high calculation complexity and low manual inspection efficiency.

Method used

The object detection model based on relative position perception is adopted, and the object detection model is trained through the training set, and the relative position relationship between the support device and the bird stabbing is used, combining multi-level feature pyramids and manual annotations to optimize the loss function and improve the bird stabbing recognition accuracy.

Benefits of technology

It improves the accuracy and sensitivity of bird spike recognition, reduces the computational complexity, assists in manual inspection, and improves inspection efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115511876B_ABST
    Figure CN115511876B_ABST
Patent Text Reader

Abstract

The present invention provides a method and system for identifying anti-bird spines on high-speed railway catenaries based on relative position perception, belonging to the field of computer vision technology, including: obtaining inspection images of high-speed railway catenaries; using a pre-trained object detection model to process the obtained inspection images of high-speed railway catenaries to obtain anti-bird spine identification results; wherein, the pre-trained object detection model is obtained by training with a training set, and the training set includes multiple inspection images of high-speed railway catenaries and labels for annotating anti-bird spines in the inspection images of high-speed railway catenaries. The present invention focuses on the inspection scenario of high-speed railway catenaries, makes full use of industry background knowledge, and through fine-grained characterization and representation modeling of the relative position relationship between support devices (large targets) and anti-bird spine components (small targets), improves the recognition accuracy of the model for small anti-bird spine targets, has low computational complexity, can assist manual inspection, and improves the work efficiency of inspection personnel.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision technology, and particularly relates to a method and system for identifying anti-bird thorns on high-speed rail catenaries based on relative position perception. Background Art

[0002] The catenary is the only way for high-speed trains to obtain electrical energy. Since it is erected outdoors, it is extremely vulnerable to various foreign objects. Among them, bird damage accounts for the highest proportion and has become one of the main factors threatening the safe operation of the high-speed railway system. Anti-bird thorns are often installed on the support devices of the catenary system to prevent birds from building nests, which can effectively reduce the failure rate of power supply equipment. However, affected by bad weather, anti-bird thorns often fall off, thus losing their anti-bird function. At present, each railway administration mainly checks whether the anti-bird thorns are intact by manually verifying and inspecting images, but this method has deficiencies such as low efficiency and high labor costs.

[0003] Identifying anti-bird thorns from inspection images is a typical computer vision problem; more specifically, it is an object detection problem, that is, to determine whether the target to be detected exists in the input image; if it exists, further identify the location of the target with a bounding box. Since the shooting angle of the railway inspection vehicle is very large, the pixel ratio of the anti-bird thorn components in the obtained inspection images is very small (a typical small object detection problem), and in addition, the anti-bird thorn is a radial object with low visual recognition, so the detection task of anti-bird thorn components is more challenging than conventional object detection tasks (such as face detection, pedestrian detection, vehicle detection, etc.). Currently, most of the mainstream object detection methods in the industry are based on deep learning; according to different proposal box generation mechanisms, existing deep learning object detection methods can be roughly divided into two categories: two-stage methods and one-stage methods. Representative two-stage methods include R-CNN, Faster R-CNN, and Cascade R-CNN, etc. Representative one-stage methods include YOLOv3, SSD, and Retinanet, etc.

[0004] Due to the particularity of the application scenario, directly applying the above existing object detection technologies to anti-bird thorn identification is difficult to meet the requirements of railway inspection, and there are problems such as high computational cost and low detection accuracy. Summary of the Invention

[0005] The purpose of the present invention is to provide a method and system for identifying anti-bird thorns on high-speed rail catenaries based on relative position perception to solve at least one of the technical problems in the above background art.

[0006] To achieve the above purpose, the present invention adopts the following technical solutions:

[0007] On the one hand, the present invention provides a method for identifying anti-bird thorns on high-speed rail catenaries based on relative position perception, including:

[0008] Obtain the inspection images of the high - speed rail catenary;

[0009] Use the pre - trained object detection model to process the obtained inspection images of the high - speed rail catenary to obtain the anti - bird - spike recognition result; wherein, the pre - trained object detection model is obtained by training with a training set, and the training set includes multiple inspection images of the high - speed rail catenary and labels annotating the anti - bird - spikes in the inspection images of the high - speed rail catenary.

[0010] Preferably, training the object detection model includes: collecting inspection images of the high - speed rail catenary to obtain an original data set; manually annotating the inspection images of the high - speed rail catenary in the original data set to obtain a labeled sample set; setting the network structure of the model; setting the output anchor box style of the model; setting the loss function of the model; based on the above settings, using the labeled sample set, according to the model output and the true label, applying the stochastic gradient descent algorithm to optimize the loss function of the model to obtain the optimal model parameters.

[0011] Preferably, the network structure set for training the object detection model includes:

[0012] A backbone network for obtaining a multi - level feature pyramid of the input image, where the shallow pyramid corresponds to the primary visual features of the image and the top - level pyramid corresponds to the high - level semantic features;

[0013] A neck network that enhances the pyramid representation ability through cross - level feature crossing, so that each layer of the pyramid has both primary visual features and high - level semantic information;

[0014] A head network that predicts the anti - bird - spike components and their bounding boxes in the image according to the enhanced feature pyramid.

[0015] Preferably, manually annotating the inspection images of the high - speed rail catenary in the original data set includes: manually annotating two types of targets, namely the boom support device and the anti - bird - spike, in the inspection images, saving them as independent annotation files, which together with the inspection images form a labeled sample set; wherein, the annotation file records the target category, the width, height and center point coordinates of its bounding box.

[0016] Preferably, the loss function used by the object detection model is:

[0017]

[0018] wherein, represents the classification loss of the double targets of the support device and the anti - bird - spike, represents the regression loss of the double - target bounding box, Indicates the relative position loss between the predicted bounding box of the support device and the predicted bounding box of the anti-bird spike. α, β, and γ are hyperparameters used to adjust the proportion of each term in the loss function;

[0019] Classification loss Is defined as the cross-entropy loss between the true class label of the target and the class label predicted by the model; Regression loss Is defined as the squared loss between the true bounding box offset of the target and the predicted bounding box offset of the model;

[0020] Relative position loss Is defined as:

[0021]

[0022] Where A is the set composed of all pixels within the predicted box of the support device, B is the set composed of all pixels within the predicted box of the anti-bird spike, and |·| represents the number of elements in the set; when B is contained in A, the relative position loss is zero; when the number of elements in A is more than the number of elements in B and there is no inclusion relationship between them, the relative position loss is defined as the difference between the number of elements in the union of the two and the number of elements in A; when A is contained in B, the relative position loss is defined as a constant multiple of the difference between the number of elements of the two.

[0023] Preferably, the online prediction process of the object detection model includes: removing the predicted bounding boxes with relatively low prediction confidence through a first threshold; at the same time, when the intersection over union of two predicted bounding boxes exceeds a second threshold, the non-maximum suppression strategy is used to select the predicted bounding box with relatively low prediction confidence among the two; where

[0024] The prediction confidence is defined as:

[0025] O = IoU1 × IoU2

[0026] Where IoU1 is the intersection over union between the true bounding box of the anti-bird spike and the predicted bounding box of the anti-bird, and IoU2 is the intersection over union between the predicted bounding box of the support device and the predicted bounding box of the anti-bird spike, which are defined as follows respectively;

[0027]

[0028]

[0029] Where C represents the set composed of all pixels within the true bounding box of the anti-bird spike, and D represents the set composed of all pixels within the minimum bounding rectangle of the predicted bounding box of the support device and the predicted bounding box of the anti-bird spike.

[0030] In a second aspect, the present invention provides a high-speed rail catenary anti-bird spike recognition system based on relative position perception, including:

[0031] An acquisition module for acquiring high-speed rail catenary inspection images;

[0032] A detection module, configured to process the acquired inspection image of the high-speed rail catenary by using a pre-trained object detection model to obtain an anti-bird-spike recognition result; wherein, the pre-trained object detection model is obtained by training with a training set, and the training set includes multiple inspection images of the high-speed rail catenary and labels annotating the anti-bird spikes in the inspection images of the high-speed rail catenary.

[0033] In a third aspect, the present invention provides a non-transitory computer-readable storage medium, which is used to store computer instructions. When the computer instructions are executed by a processor, the method for identifying anti-bird spikes of a high-speed rail catenary based on relative position perception as described above is implemented.

[0034] In a fourth aspect, the present invention provides a computer program product, including a computer program, which is used to implement the method for identifying anti-bird spikes of a high-speed rail catenary based on relative position perception as described above when running on one or more processors.

[0035] In a fifth aspect, the present invention provides an electronic device, including: a processor, a memory, and a computer program; wherein, the processor is connected to the memory, the computer program is stored in the memory, and when the electronic device runs, the processor executes the computer program stored in the memory so that the electronic device executes instructions for implementing the method for identifying anti-bird spikes of a high-speed rail catenary based on relative position perception as described above.

[0036] Term explanation:

[0037] The high-speed rail catenary is a power transmission line erected above the railway line and is the power source for electric locomotives. Its working state directly affects the transportation capacity of the high-speed rail.

[0038] The support device is a device in the catenary system used to bear all the mechanical loads of the positioning device and the contact suspension and transfer them to the pillar and the foundation.

[0039] The anti-bird spike is composed of 24 steel wires and is in a scattered shape, used to prevent birds from building nests, and is usually installed on the support device of the power transmission and transformation line equipment.

[0040] A small target refers to a target to be recognized whose area accounts for no more than 0.12% in the whole picture, such as the anti-bird spike component mentioned in the present invention.

[0041] Advantages of the present invention: Focusing on the inspection scenario of high-speed railway catenaries, making full use of industry background knowledge (anti-bird spines are mostly installed on the cantilever support device), by finely depicting and representing the relative position relationship between the support device (large target) and the anti-bird spine component (small target), the recognition accuracy of the anti-bird spine small target by the model is improved, the detection sensitivity is high, the computational complexity is low, it can assist manual inspection, and the work efficiency of the inspection personnel is improved.

[0042] The advantages of the additional aspects of the present invention will be more clearly given in the following description part, or understood through the practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0044] Figure 1 It is a schematic design flow diagram of the anti-bird spine recognition method for high-speed railway catenaries based on relative position perception according to the embodiments of the present invention.

[0045] Figure 2 It is a schematic network structure diagram of the target detection model according to the embodiments of the present invention.

[0046] Figure 3 It is a comparison diagram of the target detection results according to the embodiments of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0047] The following details the embodiments of the present invention. The examples of the embodiments are shown in the drawings, where the same or similar reference numerals represent the same or similar elements or elements with the same or similar functions from beginning to end. The embodiments described below with reference to the drawings are exemplary only for explaining the present invention and should not be construed as limiting the present invention.

[0048] Those skilled in the art of the present technology can understand that, unless otherwise defined, all terms (including technical terms and scientific terms) used here have the same meaning as the general understanding of those of ordinary skill in the art in the field to which the present invention belongs.

[0049] It should also be understood that terms such as those defined in a general dictionary should be understood to have a meaning consistent with the meaning in the context of the prior art, and will not be interpreted with an idealized or overly formal meaning unless defined as here.

[0050] Those skilled in the art can understand that, unless specifically stated otherwise, the singular forms "a", "an", "the", and "said" used herein may also include the plural forms. It should be further understood that the term "including" used in the description of the present invention means the presence of the described features, integers, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, and / or their groups.

[0051] In the description of this specification, the description with reference to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.

[0052] For ease of understanding the present invention, the following will further explain the present invention with specific examples in conjunction with the accompanying drawings, and the specific examples do not constitute a limitation on the embodiments of the present invention.

[0053] Those skilled in the art should understand that the drawings are only schematic diagrams of the embodiments, and the components in the drawings are not necessarily essential for implementing the present invention.

[0054] Embodiment 1

[0055] In this Embodiment 1, a high-speed rail catenary anti-bird thorn recognition system based on relative position perception is provided, including:

[0056] An acquisition module, configured to acquire high-speed rail catenary inspection images;

[0057] A detection module, configured to use a pre-trained object detection model to process the acquired high-speed rail catenary inspection images to obtain an anti-bird thorn recognition result; wherein, the pre-trained object detection model is obtained by training with a training set, and the training set includes multiple high-speed rail catenary inspection images and labels annotating the anti-bird thorns in the high-speed rail catenary inspection images.

[0058] In this Embodiment 1, using the above system, a high-speed rail catenary anti-bird thorn recognition method based on relative position perception is realized, including:

[0059] Using the acquisition module to acquire high-speed rail catenary inspection images; for example, high-speed rail catenary inspection images can be collected by an inspection vehicle-mounted camera.

[0060] Then, using the pre-trained object detection model in the monitoring module, process the acquired inspection images of the high-speed rail catenary to obtain the anti-bird thorn recognition result; among them, the pre-trained object detection model is obtained by training with a training set, and the training set includes multiple inspection images of the high-speed rail catenary and labels annotating the anti-bird thorns in the inspection images of the high-speed rail catenary.

[0061] Among them, training the object detection model includes: collecting inspection images of the high-speed rail catenary to obtain an original data set; manually annotating the inspection images of the high-speed rail catenary in the original data set to obtain a labeled sample set; setting the network structure of the model; setting the output anchor box style of the model; setting the loss function of the model; based on the above settings, using the labeled sample set, according to the model output and the true label, applying the stochastic gradient descent algorithm to optimize the loss function of the model to obtain the optimal model parameters.

[0062] Among them, the network structure set for training the object detection model includes:

[0063] A backbone network, used to obtain a multi-level feature pyramid of the input image, where the shallow pyramid corresponds to the primary visual features of the image, and the top pyramid corresponds to the high-level semantic features;

[0064] A neck network, which improves the pyramid representation ability through cross-level feature intersection, so that each layer of the pyramid has both primary visual features and high-level semantic information;

[0065] A head network, which predicts the anti-bird thorn components and their bounding boxes in the image according to the enhanced feature pyramid.

[0066] Manually annotating the inspection images of the high-speed rail catenary in the original data set includes: manually annotating two types of targets, namely the cantilever support device and the anti-bird thorn, in the inspection images, and saving them as independent annotation files, which together with the inspection images form a labeled sample set; among them, the annotation file records the target category, the width and height of its bounding box, and the center point coordinates.

[0067] The loss function is:

[0068]

[0069] Among them, represents the classification loss of the double targets of the support device and the anti-bird thorn, represents the regression loss of the double-target bounding box, represents the relative position loss between the predicted bounding box of the support device and the predicted bounding box of the anti-bird thorn. α, β, and γ are hyperparameters used to adjust the proportion of each item in the loss function;

[0070] Classification loss Defined as the cross-entropy loss between the target true class label and the class label predicted by the model; regression loss Defined as the squared loss between the target true bounding box offset and the bounding box offset predicted by the model;

[0071] Relative position loss Defined as:

[0072]

[0073] where A is the set composed of all pixels within the prediction box of the support device, B is the set composed of all pixels within the prediction box of the anti-bird thorn, |·| represents the number of elements in the set; when B is contained in A, the relative position loss is zero; when the number of elements in A is more than that in B and there is no inclusion relationship between them, the relative position loss is defined as the difference between the number of elements in their union and the number of elements in A; when A is contained in B, the relative position loss is defined as a constant multiple of the difference between the number of their elements.

[0074] The online prediction process of the object detection model includes: removing the prediction boxes with relatively low prediction confidence through the first threshold; at the same time, when the intersection over union of two prediction boxes exceeds the second threshold, adopting the non-maximum suppression strategy to select the box with relatively low prediction confidence among them; where

[0075] The prediction confidence is defined as:

[0076] O = IoU1 × IoU2

[0077] where IoU1 is the intersection over union between the true box of the anti-bird thorn and the predicted box of the anti-bird thorn, and IoU2 is the intersection over union between the predicted box of the support device and the predicted box of the anti-bird thorn, which are defined as follows respectively;

[0078]

[0079]

[0080] where C represents the set composed of all pixels within the true box of the anti-bird thorn, and D represents the set composed of all pixels within the minimum bounding rectangle of the predicted box of the support device and the predicted box of the anti-bird thorn.

[0081] Embodiment 2

[0082] In this Embodiment 2, a method for identifying anti-bird thorns of high-speed railway catenaries based on relative position perception is provided, and this method includes the following steps:

[0083] Collect high-speed railway catenary inspection images through the inspection vehicle-mounted camera to obtain the original data set;

[0084] Manually annotate the inspection images to obtain the labeled sample set;

[0085] Train a model offline using labeled samples to obtain a relative position perception object detection model RPOD;

[0086] Use the trained RPOD model to perform online prediction on unlabeled inspection pictures and output the anti-bird-spike recognition results.

[0087] Among them, the process of obtaining the catenary inspection images includes: installing a high-speed camera between the operation console and the windshield in the inspection vehicle cab to capture the train running environment; superimposing context information such as line name, speed, mileage, and time on the video; sampling the inspection video to obtain a high-speed railway catenary inspection image dataset.

[0088] Specifically, the process of labeling the inspection images includes: manually labeling two types of targets, namely the boom support device and the anti-bird-spike, in the inspection images, saving them as independent labeling files, and together with the inspection images, forming a labeled sample set; among them, the labeling file records the target category and the width, height, and center point coordinates of its bounding box.

[0089] The training process of the RPOD model includes: setting the network structure of the RPOD model; setting the output anchor box style of the RPOD model; setting the loss function of the RPOD model; based on the above settings, according to the model output and the true label, use the stochastic gradient descent algorithm to optimize the loss function of the RPOD model to obtain the optimal model parameters.

[0090] The network structure of the RPOD model includes: a backbone network for obtaining a multi-level feature pyramid of the input image, where the shallow pyramid corresponds to the primary visual features of the image, and the top pyramid corresponds to the high-level semantic features; a neck network that enhances the representation ability of the pyramid through cross-level feature intersection, so that each layer of the pyramid has both primary visual features and high-level semantic information; a head network that predicts the anti-bird-spike components and their bounding boxes in the image based on the enhanced feature pyramid.

[0091] The output anchor box style of the RPOD model includes: representing the manually labeled bounding box described in claim 3 as a 2D feature vector composed of width and height, then performing clustering analysis on it to obtain K clusters, and finally truncating and rounding the cluster center vector to obtain K styles of anchor boxes.

[0092] The loss function of the RPOD model is defined as:

[0093]

[0094] Among them represents the classification loss of the double targets of the support device and the anti-bird-spike, represents the regression loss of the double-target bounding box, Indicates the relative position loss between the predicted bounding box of the support device and the predicted bounding box of the anti-bird spike. α, β, and γ are hyperparameters used to adjust the weights of each term in the loss function;

[0095] Classification loss Is defined as the cross-entropy loss between the true class label of the target and the class label predicted by the model; Regression loss Is defined as the squared loss between the true bounding box offset of the target and the predicted bounding box offset of the model;

[0096] Relative position loss Is defined as:

[0097]

[0098] Where A is the set composed of all pixels within the predicted box of the support device, B is the set composed of all pixels within the predicted box of the anti-bird spike, and · represents the number of elements in the set; when B is contained in A, the relative position loss is zero; when the number of elements in A is more than that in B and there is no inclusion relationship between them, the relative position loss is defined as the difference between the number of elements in their union and the number of elements in A; when A is contained in B, the relative position loss is defined as a constant multiple of the difference between the number of their elements.

[0099] The online prediction process of the RPOD model includes: eliminating the prediction bounding boxes with relatively low prediction confidence through the first threshold; at the same time, when the intersection over union of two prediction bounding boxes exceeds the second threshold, the non-maximum suppression strategy is used to select the prediction bounding box with relatively low prediction confidence between the two.

[0100] The prediction confidence of the RPOD model is defined as:

[0101] O = IoU1 × IoU2

[0102] Where IoU1 is the intersection over union between the true bounding box of the anti-bird spike and the predicted bounding box of the anti-bird, and IoU2 is the intersection over union between the predicted bounding box of the support device and the predicted bounding box of the anti-bird spike, which are defined as follows respectively;

[0103]

[0104]

[0105] Where C represents the set composed of all pixels within the true bounding box of the anti-bird spike, and D represents the set composed of all pixels within the minimum bounding rectangle of the predicted bounding box of the support device and the predicted bounding box of the anti-bird spike.

[0106] Example 3

[0107] In this Example 3, a method for identifying anti-bird devices for high-speed rail catenaries based on relative position perception is provided, and its working process is as Figure 1As shown below. First, obtain the inspection images of the high-speed railway catenary (step S1); then label the wrist arm support device and anti-bird thorn components in the inspection images of the high-speed railway catenary to construct a dataset (step S2); use the labeled data to train the RPOD model offline (step S3); finally, predict the position of the anti-bird thorn component in the image according to the trained RPOD model and output the model prediction result.

[0108] In Embodiment 3 of the present invention, the network structure of the proposed model is as Figure 2 shown. This structure includes three parts: a backbone network, a neck network, and a head network. The backbone network uses CSPDarknet53 to obtain a multi-level feature pyramid; the neck network uses FPN and PAN to enhance the representation ability of the feature pyramid; the head network uses a "dual prediction head design" to predict both the support device and the anti-bird thorn component simultaneously (the model only needs to output the prediction information of the anti-bird thorn component finally).

[0109] To verify the effectiveness of the model proposed in Embodiment 3 of the present invention, Figure 3 the experimental comparison results between the proposed model and the current mainstream object detection models are given. The comparison methods include the one-stage models YOLOv3 and YOLOv5, and the two-stage models FasterR-CNN and Cascade R-CNN. The dataset used in the experiment was taken from the upward and corresponding downward routes of the high-speed railway from Xuchang City to Zhengzhou City and from Zhengzhou City to Anyang City, a total of 4 route data and 60,060 pictures with a resolution of 2448*2048. Finally, 3,536 support devices and 1,101 anti-bird thorns were labeled. The hyperparameters of the model are set as follows: the number of clustering clusters is 9 (that is, there are 9 types of output anchor box styles of the model), the batch size is 16, the momentum size is 0.843, the learning rate is from 0.032 to 0.12, the weight decay is 0.00036, the maximum number of epochs is 300, the classification loss weight is 0.36, the regression loss weight is 0.04, and the twin loss function weight is 0.6.

[0110] The experimental results show that the performance of the RPOD model proposed in this embodiment is significantly better than other similar models. Compared with the Cascade R-CNN model with the closest results, the RPOD model proposed in this embodiment reaches an AP value of 43.22 in the anti-bird thorn dataset, achieving a 2.2% improvement. It can be seen that this embodiment has achieved a better effect on anti-bird thorn detection and has practical application value.

[0111] Embodiment 4

[0112] Embodiment 4 of the present invention provides a non-transitory computer-readable storage medium, which is used to store computer instructions. When the computer instructions are executed by a processor, the method for identifying anti-bird thorns on the high-speed railway catenary based on relative position perception is realized. The method includes:

[0113] Obtain inspection images of the high - speed rail catenary;

[0114] Use a pre - trained object detection model to process the obtained inspection images of the high - speed rail catenary to obtain an anti - bird - spike recognition result; wherein, the pre - trained object detection model is trained by a training set, and the training set includes multiple inspection images of the high - speed rail catenary and labels annotating the anti - bird - spikes in the inspection images of the high - speed rail catenary.

[0115] Embodiment 5

[0116] Embodiment 5 of the present invention provides a computer program (product), including a computer program, which when running on one or more processors, is used to implement a method for identifying anti - bird - spikes on the high - speed rail catenary based on relative position perception. The method includes:

[0117] Obtain inspection images of the high - speed rail catenary;

[0118] Use a pre - trained object detection model to process the obtained inspection images of the high - speed rail catenary to obtain an anti - bird - spike recognition result; wherein, the pre - trained object detection model is trained by a training set, and the training set includes multiple inspection images of the high - speed rail catenary and labels annotating the anti - bird - spikes in the inspection images of the high - speed rail catenary.

[0119] Embodiment 6

[0120] Embodiment 6 of the present invention provides an electronic device, including: a processor, a memory, and a computer program; wherein, the processor is connected to the memory, the computer program is stored in the memory, and when the electronic device runs, the processor executes the computer program stored in the memory so that the electronic device executes instructions for implementing a method for identifying anti - bird - spikes on the high - speed rail catenary based on relative position perception. The method includes:

[0121] Obtain inspection images of the high - speed rail catenary;

[0122] Use a pre - trained object detection model to process the obtained inspection images of the high - speed rail catenary to obtain an anti - bird - spike recognition result; wherein, the pre - trained object detection model is trained by a training set, and the training set includes multiple inspection images of the high - speed rail catenary and labels annotating the anti - bird - spikes in the inspection images of the high - speed rail catenary.

[0123] In summary, the exclusive anti-bird thorn recognition method for high-speed railway catenary safety inspection described in the embodiments of the present invention makes reasonable use of scene prior knowledge. Based on the relative position relationship between the wrist arm support device and the anti-bird thorn, the detection range of the anti-bird thorn is reduced from the entire inspection image to the surrounding area of the wrist arm support device, greatly improving the prediction accuracy of the detection model. In the embodiment, the network structure and objective function of the anti-bird thorn detection model RPOD, the backbone network and loss function used are CSPDarknet and SymbioticLoss respectively. Those skilled in the art can obviously make various modifications to the above embodiments easily. For example, replacing the backbone network with ResNet-101 or VGG19, etc., or replacing the loss function with CIoULoss, GIoULoss, etc.

[0124] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0125] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the embodiments of the present invention. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, as well as the combination of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for realizing the functions specified in Figure 1 one or more of the processes Figure 1 or multiple processes and / or blocks

[0126] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device, and the instruction device realizes the functions specified in Figure 1 one or more of the processes Figure 1 or multiple processes and / or blocks

[0127] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus, where they perform a series of operational steps to produce a computer-implemented process, so that the instructions executed on the computer or other programmable apparatus provide steps for implementing the functions specified in one process or a plurality of processes and / or blocks Figure 1 one process or a plurality of processes and / or blocks Figure 1 in one block or a plurality of blocks.

[0128] Although the specific embodiments of the present invention have been described in conjunction with the accompanying drawings, they are not intended to limit the scope of the present invention. Those skilled in the art should understand that, based on the technical solutions disclosed in the present invention, various modifications or variations that can be made by those skilled in the art without creative efforts should be covered within the scope of the present invention.

Claims

1. A method for identifying anti-bird spurs of high-speed railway catenary based on relative position perception, characterized in that, Including: Obtain inspection images of high-speed rail catenaries; Use a pre-trained object detection model to process the obtained inspection images of high-speed rail catenaries to obtain an anti-bird-spike recognition result; wherein, the pre-trained object detection model is obtained by training with a training set, and the training set includes multiple inspection images of high-speed rail catenaries and labels annotating anti-bird spikes in the inspection images of high-speed rail catenaries; The training process of the object detection model includes: collecting inspection images of high-speed rail catenaries to obtain an original data set; manually annotating the inspection images of high-speed rail catenaries in the original data set to obtain a labeled sample set; setting the network structure of the model; setting the output anchor box style of the model; setting the loss function of the model; based on the above settings, using the labeled sample set, according to the model output and the true label, applying the stochastic gradient descent algorithm to optimize the loss function of the model to obtain the optimal model parameters; The loss function used by the object detection model is: Among them, represents the classification loss of the dual targets of the support device and the anti-bird spike, represents the regression loss of the dual-target bounding box, represents the relative position loss between the predicted bounding box of the support device and the predicted bounding box of the anti-bird spike, and α, β, and γ are hyperparameters used to adjust the proportion of each item in the loss function; Classification loss Defined as the cross-entropy loss between the target true class label and the class label predicted by the model; Regression loss Defined as the squared loss between the target true bounding box offset and the bounding box offset predicted by the model; Relative position loss It is defined as: where A is a set composed of all pixels within the prediction box of the support device, B is a set composed of all pixels within the prediction box of the anti-bird spike, and |·| represents the number of elements in the set; when B is included in A, the relative position loss is zero; when the number of elements in A is more than the number of elements in B and there is no inclusion relationship between the two, the relative position loss is defined as the difference between the number of elements in the union of the two and the number of elements in A; when A is included in B, the relative position loss is defined as a constant multiple of the difference between the number of elements of the two.

2. The method for identifying anti-bird spurs of high-speed rail catenary based on relative position perception according to claim 1, characterized in that The network structure set by the object detection model includes: A backbone network for obtaining a multi-level feature pyramid of the input image, where the shallow pyramid corresponds to the primary visual features of the image and the top pyramid corresponds to the high-level semantic features; A neck network that enhances the representation ability of the pyramid through cross-level feature crossing, so that each layer of the pyramid has both primary visual features and high-level semantic information; A head network that predicts the anti-bird-spike components and their bounding boxes in the image according to the enhanced feature pyramid.

3. The method for identifying anti-bird spurs of high-speed railway catenary based on relative position perception according to claim 1, wherein Manually annotating the inspection images of high-speed rail catenaries in the original data set includes: manually annotating two types of targets, namely the cantilever support device and the anti-bird spike, in the inspection images, and saving them as independent annotation files, which together with the inspection images form a labeled sample set; wherein, the annotation file records the target category, the width, height and center point coordinates of its bounding box.

4. The method for identifying anti-bird spurs of high-speed rail catenary based on relative position perception according to claim 1, wherein The online prediction process of the object detection model includes: removing prediction boxes with relatively low prediction confidence through a first threshold; at the same time, when the intersection over union of two prediction boxes exceeds a second threshold, adopting a non-maximum suppression strategy to select the box with relatively low prediction confidence between the two; wherein, The prediction confidence is defined as: O = IoU1 × IoU2 where IoU1 is the intersection over union between the true box of the anti-bird spike and the predicted box of the anti-bird, and IoU2 is the intersection over union between the predicted box of the support device and the predicted box of the anti-bird spike, which are defined as follows respectively; where C represents a set composed of all pixels within the true box of the anti-bird spike, and D represents a set composed of all pixels within the minimum bounding rectangle of the predicted box of the support device and the predicted box of the anti-bird spike.

5. A bird-spike recognition system for high-speed railway catenary based on relative position perception, characterized in that, Including: An acquisition module for obtaining inspection images of high-speed rail catenaries; A detection module, which is used to process the acquired inspection images of the high-speed rail catenary by using a pre-trained object detection model to obtain an anti-bird-spike recognition result; wherein, the pre-trained object detection model is obtained by training with a training set, and the training set includes multiple inspection images of the high-speed rail catenary and labels annotating the anti-bird-spikes in the inspection images of the high-speed rail catenary; The training process of the object detection model includes: collecting inspection images of the high-speed rail catenary to obtain an original data set; manually annotating the inspection images of the high-speed rail catenary in the original data set to obtain a labeled sample set; setting the network structure of the model; setting the output anchor box style of the model; setting the loss function of the model; based on the above settings, using the labeled sample set, according to the model output and the true label, applying the stochastic gradient descent algorithm to optimize the loss function of the model to obtain the optimal model parameters; The loss function used by the object detection model is: Among them, represents the classification loss of the dual targets of the support device and the anti-bird spike, represents the regression loss of the dual-target bounding box, represents the relative position loss between the predicted bounding box of the support device and the predicted bounding box of the anti-bird spike, and α, β, γ are hyperparameters used to adjust the proportion of each item in the loss function; Classification loss Defined as the cross-entropy loss between the target true class label and the class label predicted by the model; regression loss Defined as the squared loss between the target true bounding box offset and the bounding box offset predicted by the model; Relative position loss It is defined as: wherein, A is a set composed of all pixels within the prediction box of the support device, B is a set composed of all pixels within the prediction box of the anti-bird-spike, and |·| represents the number of elements in the set; when B is included in A, the relative position loss is zero; when the number of elements in A is more than the number of elements in B and there is no inclusion relationship between the two, the relative position loss is defined as the difference between the number of elements in the union of the two and the number of elements in A; when A is included in B, the relative position loss is defined as a constant multiple of the difference between the number of elements of the two.

6. A non-transitory computer-readable storage medium, characterized in that, The non-transitory computer-readable storage medium is used to store computer instructions, and when the computer instructions are executed by a processor, the method for identifying anti-bird-spikes of a high-speed rail catenary based on relative position perception as described in any one of claims 1-4 is implemented.

7. A computer program product, characterized in that, It includes a computer program, and when the computer program runs on one or more processors, it is used to implement the method for identifying anti-bird-spikes of a high-speed rail catenary based on relative position perception as described in any one of claims 1-4.

8. An electronic device, characterized in that, It includes: a processor, a memory, and a computer program; wherein, the processor is connected to the memory, the computer program is stored in the memory, and when the electronic device runs, the processor executes the computer program stored in the memory so that the electronic device executes the instructions for implementing the method for identifying anti-bird-spikes of a high-speed rail catenary based on relative position perception as described in any one of claims 1-4.