A method and system for detecting and analyzing foreign objects in track switches

Through the combination of multi-camera sets and deep learning networks, high-precision detection and type analysis of foreign objects of track switches are achieved, and the problem of unrecognizing foreign objects in the prior art is solved, which improves the accuracy and safety of detection.

CN120411656BActive Publication Date: 2025-08-29CRRC HANGZHOU DIGITAL TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510905288.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-02
Publication Date
2025-08-29
Estimated Expiration
2045-07-02

AI Technical Summary

Technical Problem

The existing track switch foreign object detection technology cannot accurately identify foreign object categories, resulting in the inability to formulate targeted treatment measures, poses safety risks, and is prone to missed inspections and missed inspections in complex environments.

Method used

Multi-camera sets are used to combine deep learning networks to realize multi-stage and multi-strategic foreign object detection and type analysis through image processing and machine learning algorithms, including object detection, image acquisition, coarse classification, detailed feature extraction and multi-feature analysis. The loss weight is dynamically adjusted in combination with foreign object type-specific losses and fuzzy IoU losses, and foreign objects are identified and classified.

Benefits of technology

It realizes real-time and accurate identification of various foreign objects during the train operation, reduces the rate of false detection and missed detection, adapts to complex and changeable track environments, and provides timely safety guarantees.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120411656B_ABST
    Figure CN120411656B_ABST
Patent Text Reader

Abstract

The present invention provides a method and system for detecting and analyzing foreign objects in a track switch. A first camera group is provided at the front end of a train. The switch position is identified by a preset target detection model, and foreign objects in the switch are calibrated by a preset foreign object detection model. An image calibrated to the foreign object in the switch is captured by a second camera group as an image to be analyzed. A second boundary area where wheels roll over the foreign object is extracted from the image to be analyzed by a preset wheel rolling track detection model, and the foreign objects in the image to be analyzed are roughly classified to obtain a rough classification result. A boundary detection strategy is selected according to the rough classification result in the foreign object image to identify the first boundary area. Corresponding texture features are respectively extracted from the first boundary area and the second boundary area by a deep learning network, and the extracted texture features are mapped one-to-one to obtain a mapping relationship group. The feature change rate is then obtained by analyzing the mapping relationship group, and the specific foreign object type is obtained by analyzing the multi-feature analysis strategy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a track foreign body identification system, and in particular to a track switch foreign body detection and type analysis method and system. Background Art

[0002] In the railway transportation system, turnouts are critical connecting devices for track lines. Their proper operation plays a decisive role in the safe and efficient operation of trains. The presence of foreign objects in the turnout area can lead to serious accidents such as train derailment and signal failure, resulting in significant economic losses and safety hazards.

[0003] Currently, foreign object detection in rail switches primarily relies on manual inspections and traditional sensor detection. Manual inspections rely on regular visual inspections of the switch area, which is not only inefficient and labor-intensive, but also subject to human factors, making it difficult to detect small foreign objects or accurately detect them in adverse weather conditions. Traditional sensor detection, such as infrared sensors and pressure sensors, while capable of automated detection to a certain extent, often only detects the presence of foreign objects and cannot accurately analyze their type. Different types of foreign objects have varying degrees of impact on switch operation. For example, metal foreign objects may cause electrical shorts, while foreign objects such as paper towels and plastic bags may affect the normal rotation of the switch. Therefore, the inability to analyze the foreign object type makes it difficult to formulate targeted treatment measures, posing a significant safety risk. Furthermore, some existing image-based detection methods typically utilize only a single camera to capture images. Due to their single perspective and limited information, they are prone to missed detections and false detections in complex environments, failing to meet the high-precision and high-reliability requirements for switch foreign object detection in railway transportation. Summary of the Invention

[0004] In view of the shortcomings of the prior art, the purpose of the present invention is to provide a method and system for detecting and analyzing foreign objects in a track switch, so as to overcome the above-mentioned shortcomings in the prior art.

[0005] To achieve the above object, the present invention provides the following technical solutions:

[0006] A method for detecting and classifying foreign objects in a track switch, comprising:

[0007] In the switch foreign object identification step, a first camera group is provided at the front end of the train. The first camera group is used to shoot track video and perform frame detection. The switch position is then identified using a preset target detection model. Foreign objects in the switch are calibrated using a preset foreign object detection model, and the corresponding image in the video is captured as a foreign object map.

[0008] In an image acquisition step, a second camera group arranged behind the train wheels captures an image calibrated to the foreign object in the turnout as an image to be analyzed;

[0009] an image processing step, extracting a second boundary region of the wheel on the foreign object in the image to be analyzed using a preset wheel rolling track detection model, and roughly classifying the foreign objects in the image to be analyzed to obtain a rough classification result, wherein the rough classification result includes hard objects and soft foreign objects, and selecting a boundary detection strategy based on the rough classification result in the foreign object image to identify the boundary region as the first boundary region;

[0010] In the detail feature extraction step, corresponding texture features are extracted in the first boundary area and the second boundary area respectively through a deep learning network, and the extracted texture features are mapped one by one to obtain a mapping relationship group, and then the feature change rate is obtained by analyzing the mapping relationship group. The feature change rate includes texture contrast ratio, color contrast ratio and shape contrast ratio, and then the specific foreign body type is obtained through multi-feature analysis strategy based on the texture contrast ratio, color contrast ratio and shape contrast ratio.

[0011] Preferably, the flexible foreign objects include paper towels, plastic bags, and fabrics. When the rough classification result is a flexible foreign object, the probability distribution vector representing the probability value of each specific category in the flexible foreign object is output. The database is provided with foreign object type labels and priority boundary recognition strategies and standard sample features corresponding to each category. According to the distribution vector of the probability value of each specific category in the rough classification result, the database is queried to obtain the corresponding boundary detection strategy.

[0012] Preferably, when the probability values ​​of multiple specific categories reach the threshold, the first boundary area corresponding to each category of foreign matter is obtained according to each category identification strategy, and the foreign matter boundaries in the foreign matter image are analyzed through the image comparison strategy, and the first boundary area is determined according to the image similarity, and the probability value is adjusted according to the comparison result.

[0013] Preferably, the boundary detection strategy includes:

[0014] In a targeted image processing step, based on the coarse classification results, when the foreign object is a tissue or fabric, image contrast is enhanced and edges are preserved, using multi-scale edge detection combined with morphological operations to strengthen edges; when the foreign object is a plastic bag, reflections are suppressed through a polarization strategy, and boundaries are restored using a deep learning-based deblurring model and interpolation;

[0015] The model training and inference steps dynamically adjust the loss weights by combining foreign object type specificity and fuzzy IoU loss. The foreign object type specificity loss includes texture constraint loss for paper towels and fabrics, transparent area compensation loss for plastic bags, and reflective area suppression loss. The fuzzy IoU loss calculates and constructs fuzzy intersection based on fuzzy set theory.

[0016] Preferably, a position identification step is also included, in which the position of the foreign object is determined based on the position of the contour line of the first boundary area and the contour line of the track switch, and the deformation trend of the foreign object is analyzed based on the contour line of the second boundary area and the contour line of the first boundary area, the position of the deformed foreign object is analyzed, and a danger warning is generated according to the type of foreign object.

[0017] Preferably, the multi-feature analysis strategy includes combining the feature change rate with the probability weight of the coarse classification, and using a weighted voting mechanism to output the foreign body type. When there is a conflict between the feature change rate and the coarse classification result, the weight is dynamically adjusted according to the credibility of different features.

[0018] Preferably, a verification step is also included. A verification module is also provided on the train, and the verification module is provided behind the second camera group. The verification module includes a main camera, an auxiliary camera and a laser radar. According to the type of the foreign object, the lighting intensity and camera parameters of the auxiliary camera are adjusted, and the image data collected by the main camera and the auxiliary camera and the laser electrical signal collected by the laser radar are combined, and the data is verified based on multimodal data fusion processing. Based on the verification result, the abnormality analysis result is obtained, and compared with the foreign object type to determine whether it is consistent.

[0019] Preferably, the target detection model is provided with a small target positioning strategy, including utilizing the network architecture for feature extraction and embedding an attention mechanism module in the network. The attention mechanism module includes a dynamic channel weight module, which gives high weights to the key feature channels of the turnout dataset through the SE-Net structure, and performs cross-scale channel interaction to screen the key features of the turnout target. The spatial attention module captures the global position of the turnout through global pooling and generates an attention map to locate the target area.

[0020] Preferably, the second camera group is a high-definition camera. When the first camera group detects a turnout, the second camera is adjusted so that the second camera captures an image of the area where the turnout is located.

[0021] A track switch foreign body detection and classification analysis system, comprising:

[0022] The switch foreign object recognition module is equipped with a first camera group at the front of the train. It uses the first camera group to shoot track video and perform frame detection. It then uses a preset target detection model to identify the switch position and a preset foreign object detection model to calibrate foreign objects in the switch. The corresponding image in the video is captured as a foreign object map.

[0023] An image acquisition module captures an image calibrated to a foreign object in the turnout through a second camera set behind the train wheels as an image to be analyzed;

[0024] An image processing module extracts a second boundary region of the wheel on the foreign object in the image to be analyzed using a preset wheel rolling track detection model, and roughly classifies the foreign objects in the image to be analyzed to obtain a rough classification result, wherein the rough classification result includes hard objects and soft foreign objects. A boundary detection strategy is selected based on the rough classification result in the foreign object image to identify the boundary region as the first boundary region;

[0025] The detail feature extraction module extracts corresponding texture features in the first boundary area and the second boundary area respectively through a deep learning network, and maps the extracted texture features one by one to obtain a mapping relationship group, and then obtains the feature change rate through the mapping relationship group analysis. The feature change rate includes texture contrast ratio, color contrast ratio and shape contrast ratio, and then obtains the specific foreign body type through multi-feature analysis strategy based on the texture contrast ratio, color contrast ratio and shape contrast ratio.

[0026] The beneficial effects of the present invention are as follows: through a multi-stage, multi-strategy processing approach, full use is made of the image information before and after the foreign object is compressed, and combined with multi-feature analysis, it is possible to accurately identify various types of foreign objects and effectively reduce the false detection and missed detection rates. For foreign objects of different types and shapes, especially flexible foreign objects, through a variety of pre-set boundary recognition strategies and a rich standard sample library, good detection and classification effects can be achieved, adapting to the complex and changing track environment. Based on the efficient computing power of image processing and machine learning algorithms, foreign object detection and type analysis can be completed in real time during train operation, providing timely protection for train operation safety. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] Figure 1 It is the overall flow chart of the present invention;

[0028] Figure 2 It is an image processing flow chart of the present invention;

[0029] Figure 3 It is a module connection diagram of the present invention;

[0030] Figure 4 It is a camera group layout diagram of the present invention. DETAILED DESCRIPTION

[0031] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0032] It should be noted that when a component is referred to as being "fixed to" another component, it may be directly on the other component or there may also be a central component. When a component is considered to be "connected to" another component, it may be directly connected to the other component or there may also be a central component. When a component is considered to be "set on" another component, it may be directly set on the other component or there may also be a central component. The terms "vertical", "horizontal", "left", "right" and similar expressions used herein are for illustrative purposes only.

[0033] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one skilled in the art to which this invention pertains. The terms used in this specification of the present invention are for the purpose of describing specific embodiments only and are not intended to limit the present invention. The term "and / or" as used herein includes any and all combinations of one or more of the associated listed items.

[0034] The embodiments of the present invention are further described below in conjunction with the accompanying drawings:

[0035] like Figure 1-Figure 4 The present invention provides a method for detecting and classifying foreign objects in a track switch, comprising:

[0036] The turnout foreign object identification step involves a first camera set installed at the front of the train. This camera sets captures track video and performs frame detection. A preset object detection model is then used to identify the turnout location and calibrate foreign objects within the turnout. The corresponding image from the video is captured as a foreign object map. If a turnout is present within the current image frame, the process proceeds to the image acquisition step. If not, the detection process continues. The first camera set installed at the front of the train consists of multiple cameras at different angles, enabling comprehensive, full-angle track video. While the train is in motion, the first camera set continuously captures track video and performs frame detection at regular intervals (e.g., 1-2 frames per second). A pre-trained object detection model is used to detect the turnout location on the extracted image frames. This object detection model utilizes a deep learning-based algorithm, such as the YOLO series, trained on a large dataset containing turnout images to enable the model to learn features such as the turnout's shape and structure. The target detection model is used to identify the switch position, and the foreign objects in the switch are calibrated through the preset foreign object detection model. According to the switch position output by the target detection model, a detection ROI is generated, and the image of the ROI area is enhanced to improve the visibility of foreign objects in low light. The foreign object detection model is used to calibrate the foreign objects in the switch, and the image is captured as the foreign object map.

[0037] In the image acquisition step, a second camera group installed behind the train wheels captures an image calibrated to the foreign object in the switch as the image to be analyzed; when the foreign object is detected at the switch, an image currently containing the switch is obtained as the foreign object map. At the same time, a second camera group is installed behind the train wheels. This camera group consists of multiple high-definition cameras, and the positions and angles of the cameras are precisely calibrated. When the train wheels pass over the foreign object, the second camera group ensures that it can capture an image of the track area at the same position as the foreign object map, and the captured image is used as the image to be analyzed. By using two camera groups to capture images of the same area at the same position, the difference between the foreign object after being crushed by the wheels can be clearly compared. After comparing and analyzing the differences between the two, a more accurate analysis of the type of foreign object can be made.

[0038] The second camera group is a high-definition camera. When the first camera group detects the switch, the second camera is adjusted so that the second camera can capture images calibrated to the area where the foreign object is located in the switch. The first camera uses a lower-definition camera. During the shooting process, the track can be continuously photographed and stored. The second camera is a high-definition camera. When the first camera detects a foreign object in the switch, the second camera is started to capture the image of the foreign object in the switch at the same position as the first camera. Different cameras are used to capture foreign objects in the switch. While ensuring continuous shooting and detection, the second camera can be used to capture high-definition images of foreign objects, which facilitates the acquisition of surface texture characteristics of foreign objects.

[0039] The target detection model is equipped with a small target positioning strategy, which includes using the network architecture for feature extraction and embedding an attention mechanism module in the network. The attention mechanism module includes a dynamic channel weight module, which uses the SE-Net structure to assign high weights to the key feature channels of the turnout dataset, and performs cross-scale channel interaction to screen the key features of the turnout target. The spatial attention module uses global pooling to capture the global position of the turnout and generate an attention map to locate the target area.

[0040] A multi-layer convolutional neural network is used as the basic framework, divided into shallow, mid, and deep layers. The network gradually extracts turnout features from details to the whole image. The shallow network focuses on preserving image details, using small convolution kernels, such as a 3×3 kernel, and a large number of channels, such as 64. In the first convolution layer, a 3×3×3 kernel (assuming the input image is a three-channel RGB image) is convolved with the original image, outputting a 64-channel feature map. This captures basic features such as the subtle edges of the turnout rail and the local structure of the switch, laying the foundation for subsequent extraction of higher-level features. The mid-layer network increases the number of convolution kernels, for example, from 64 to 128, and employs pooling operations, such as 2×2 max pooling, to expand the network's image observation range while reducing the resolution of the feature map. After two layers of convolution and one layer of pooling, the feature map is reduced in size by half, and the number of channels is increased to 128. At this point, the network can gain a preliminary understanding of the combination and relative positions of the various turnout components. The deep network introduces dilated convolution and global average pooling technology. Taking the dilated convolution with a void ratio of 2 as an example, without increasing the number of network parameters, the network can "see" a wider range of image content, expanding the original 3×3 observation range to 5×5. Combined with global average pooling, the overall semantic information of the image is obtained, allowing the network to understand the position and role of the turnout in the entire track scene. A two-way feature fusion path is constructed. From the top to the bottom, the feature map with semantic information in the high-level layer is upsampled to a size similar to the low-level feature map, and then added to the high-resolution feature map of the low-level layer; from the bottom to the top, the feature map containing detailed information in the low-level layer is passed to the high-level layer through convolution and pooling operations. For example, in the second layer of FPN, the semantic feature map with a high-level size of H / 8×W / 8 and 256 channels is upsampled to H / 4×W / 4 and then added to the detail feature map of the middle layer with the same size and 128 channels to obtain a fused feature map. This achieves the complementary advantages of turnout features at different scales, giving the network a more comprehensive understanding of turnouts.

[0041] The dynamic channel weight module, based on the principles of SE-Net, automatically adjusts the weights of each channel to enhance key feature channels related to turnouts and enable channel information exchange at different scales. The SE-Net module consists of a squeeze phase, which performs a global average pooling operation on the feature maps extracted by the network. It also performs an excitation phase, which processes the compressed vectors through two fully connected layers. The vectors are first mapped to a lower-dimensional space to reduce the number of parameters and introduce nonlinear changes. The vectors are then mapped back to their original number of channels. Finally, an activation function is used to determine the attention weights corresponding to each channel, allowing the network to learn to identify feature channels related to key turnout components, such as switches and frogs. During training, when input images containing switch switches are fed, the network gradually increases the weights of channels that describe features such as switch shape and material texture. Furthermore, a channel weight transfer mechanism is established between different layers of the feature pyramid network: key channel weights learned in higher-level feature maps are fused to the corresponding channels in lower-level feature maps through a weighted summation. This allows lower-level feature maps to highlight key information related to turnouts while preserving existing details. After convolution and pooling operations, the low-level feature map feeds back the local detail information contained in it to the high-level feature map by adjusting the attention weight, helping the high-level feature map to better understand the overall structure of the turnout.

[0042] The spatial attention module integrates global image information and generates an attention map, enabling precise localization of the switch target area while effectively reducing background interference. Global average pooling is performed on the input feature map in both the horizontal and vertical directions. The feature vectors in these two directions are concatenated along the channel dimension, then processed through a 7×7×1 convolutional layer to fuse and reduce the features. The final output is an attention map of the same size as the input feature map. Each pixel value in the attention map represents the likelihood of a switch target at that location. The feature map is adjusted by multiplying the attention map with the corresponding element in the original feature map. For small switch targets, global pooling integrates information from the entire image, minimizing interference from surrounding background information due to the small size of the target. For example, when a small switch rail target is present in the image, the attention map assigns a higher weight to the rail region and a lower weight to the background region. This multiplication significantly enhances the features of the rail region and suppresses those of the background region, enabling the network to more accurately locate the switch target during subsequent object detection.

[0043] Through the close cooperation of network feature extraction and attention mechanism modules, the small target positioning strategy can significantly improve the detection accuracy and positioning accuracy of small switch targets in complex track scenarios.

[0044] In the image processing step, a preset wheel-rollover trajectory detection model is used to extract the second boundary region of the wheel-rollover foreign object in the image to be analyzed. The foreign objects in the image to be analyzed are then roughly classified, yielding a coarse classification result. This coarse classification result includes both hard and soft foreign objects. Based on the coarse classification results in the foreign object image, a boundary detection strategy is selected to identify the boundary region as the first boundary region. The Canny edge detection algorithm is used to accurately extract the second boundary region in the image to be analyzed. A semantic analysis model based on a convolutional neural network (CNN) is constructed to perform preliminary identification of the foreign object type in the image to be analyzed. During the model training phase, over 100,000 image samples containing common foreign objects such as plastic bags, tissues, and metal fragments are collected, and transfer learning techniques are used to optimize network parameters. The model extracts and analyzes image features and outputs probabilities for foreign objects belonging to different categories. For example, the probability of a foreign object being a plastic bag is 85%, while the probability of a paper towel is 15%, achieving rapid coarse classification. Based on the coarse classification results, a corresponding boundary detection strategy is selected to identify the boundary region of the foreign object image as the first boundary region. The boundary recognition strategy is selected based on the foreign object type and attributes, effectively improving the accuracy and efficiency of boundary recognition.

[0045] Soft foreign objects include paper towels, plastic bags, and fabric. When the coarse classification result indicates a soft foreign object, a probability distribution vector representing the probability of each specific category is output. A database contains foreign object type labels, priority boundary recognition strategies, and standard sample features corresponding to each category. Based on the distribution vectors of the probability values ​​for each specific category in the coarse classification result, the database is queried to obtain the corresponding boundary detection strategy. When detecting foreign objects in track switches, a convolutional neural network (CNN) model analyzes the image using the semantic analysis strategy in the image processing step to obtain a coarse classification result for the foreign objects, classifying them into two categories: hard objects and soft objects. When a foreign object is determined to be a soft foreign object, the model outputs a probability distribution vector representing the probability of each specific category. For example, if the soft foreign objects include paper towels, plastic bags, and fabric, the probability distribution vector might be [paper towel probability value, plastic bag probability value, fabric probability value], such as [0.3, 0.6, 0.1]. This vector reflects the model's judgment of the probability of the foreign object belonging to different soft foreign object categories. The generation of the probability distribution vector is based on the model's learning from a large number of annotated images during training. During training, the model continuously adjusts its parameters to minimize the error between the predicted results and the true labels. In practice, when a new image is input for analysis, the model's neural network layers gradually extract and analyze features such as color, shape, and texture. Finally, the fully connected layers output probability values ​​for each category, which are combined to form the probability distribution vector.

[0046] The database stores labels for each specific category of foreign matter, along with the preferred boundary recognition strategy and standard sample features corresponding to each label. For each specific category of flexible foreign matter, a preferred boundary recognition strategy has been developed through extensive experiments and real-world case studies. For example, a boundary recognition strategy for paper towels might include using a texture analysis method based on a gray-level co-occurrence matrix (GLCM) to detect low-contrast, irregular edge regions. A boundary recognition strategy for fabrics might focus on analyzing the unique weave texture and wrinkle morphology, using a feature extraction method based on local binary patterns (LBP) combined with morphological operations to identify boundaries.

[0047] Standard sample features are also a crucial component of the database. They contain image feature data for different flexible foreign objects under various lighting conditions and shooting angles, such as the fiber texture parameters of paper towels, the grid pattern feature vectors of plastic bags, and the weave structure feature descriptions of fabrics. These standard sample features are used in the subsequent detailed feature extraction step to compare with the extracted features of the actual foreign object, helping to determine the precise type of foreign object.

[0048] When the probability value of a certain category in the probability distribution vector is significantly greater than the probability values ​​of other categories, the following conditions are met:

[0049]

[0050] in, is the maximum probability value in the probability distribution vector, are the probability values ​​of other categories, For the preset advantage coefficient (e.g. = 1.5, obtained through machine learning optimization). At this point, the category with the highest probability value is identified as the category to which the foreign object belongs, and this category is used as an index to query the database for the corresponding priority boundary recognition strategy. For example, if the probability distribution vector is [0.1, 0.8, 0.1], the probability value of 0.8 for a plastic bag is significantly greater than the probability values ​​for paper towels and fabric. If these conditions are met, the database is searched for the corresponding boundary recognition strategy for the plastic bag. The boundary recognition strategy for plastic bags in the database may include prioritizing the detection of highly reflective areas with smooth edges, and using a combination of threshold segmentation and morphological operations to highlight the boundary outline of the plastic bag in the image.

[0051] When the probability values ​​for multiple specific categories reach the threshold, the system uses the corresponding category identification strategy to obtain the first boundary region corresponding to each category of foreign object. The image comparison strategy analyzes the foreign object boundary in the foreign object image, determines the first boundary region based on image similarity, and adjusts the probability value based on the comparison results. If the probability values ​​for multiple specific categories reach the threshold (for example, if the probability distribution vector is [0.35, 0.3, 0.25], the probability values ​​for paper towels, plastic bags, and fabric all reach the preset threshold of 0.3), the system processes the image to be analyzed according to the boundary identification strategy corresponding to each category to obtain the foreign object boundary.

[0052] By combining the similarity calculation results of multiple specific categories, the region with the highest similarity and that conforms to the overall foreign object morphology logic is selected and determined as the first boundary region. For example, if the region corresponding to the paper towel category has the highest similarity in the foreign object map and this region reasonably connects with the boundary regions determined for other categories in terms of spatial position and morphology, then this region is determined as the final first boundary region. After determining the first boundary region, the extracted features of the first boundary region (such as texture features and shape features) are compared with the features of various standard samples in the database. Combined with the previous coarse classification probability values, the probability values ​​of each category are adjusted and optimized.

[0053] Boundary detection strategies include:

[0054] Targeted image processing steps, based on the coarse classification results, enhance image contrast and preserve edges when the foreign object is a tissue or fabric. Multi-scale edge detection combined with morphological operations is used to enhance edges. When the foreign object is a plastic bag, a polarization strategy is used to suppress reflections, and a deep learning-based deblurring model and interpolation are used to repair boundaries. Paper towels and fabric foreign objects typically have low contrast, irregular shapes, and rich textures, and their boundaries are easily confused with the track background. For these foreign objects, the system uses the following processing flow:

[0055] Image enhancement preprocessing uses the adaptive histogram equalization (CLAHE) algorithm to increase image contrast by 20%-30%, enhancing the grayscale difference between the tissue / fabric and the background. For example, when detecting a white tissue on a track, this operation can make the tissue area more prominent in the image. A bilateral filtering algorithm is used to smooth the image, removing noise while preserving edge information. This algorithm adjusts pixel values ​​by calculating weights in the spatial and value domains, effectively preserving the detailed features of the tissue edge.

[0056] The Canny operator is combined with multi-scale fusion to perform edge detection, generating an initial edge map. Multi-scale analysis methods are then combined to detect edges at different scales and fuse the results, capturing multi-level edge information, from subtle textures to overall contours. For example, the wrinkled edges of a tissue can be detected at a small scale, while the overall contour of the tissue can be grasped at a large scale.

[0057] Edges are refined through dilation and erosion operations. Dilation is used to connect broken edges, and erosion is used to remove excess edge noise, ultimately resulting in a continuous and clear tissue / fabric boundary.

[0058] Foreign objects such as plastic bags are transparent, reflective, and easily deformable. Boundary detection faces interference from reflections and blurred edges due to transparency. Polarization strategies are used to suppress reflections. The following methods can be used to address this:

[0059] Polarization strategies include reflection suppression and deblurring. By leveraging the polarization properties of plastic bag reflections and adjusting the angle of the polarization filter, we can effectively reduce surface reflections and enhance the visibility of the bag's texture and edges. For example, when a plastic bag's surface has strong reflections, polarization filtering can reveal details in the reflective areas.

[0060] Retinex algorithm enhancement: The single-scale Retinex (SSR) algorithm is applied to perform brightness enhancement and color correction on the image, improving the contrast between the transparent area of ​​the plastic bag and the background, and highlighting potential boundary information.

[0061] Deep learning deblurring and boundary restoration utilizes a deblurring model based on the U-Net architecture to restore blurred boundaries caused by transparency and reflections. By training on a large number of blurred and clear image pairs, the model can predict and restore the true edge shape of the plastic bag. For invisible boundaries caused by transparency, interpolation calculations are performed using adjacent frames or surrounding environmental information to infer and repair missing boundary sections. For example, when a portion of the plastic bag blends with the background, the model completes this boundary by analyzing the texture and color characteristics of the adjacent region.

[0062] During model training and inference, the object type-specific loss and fuzzy IoU loss are combined to dynamically adjust loss weights. The object type-specific loss includes a texture-constrained loss for tissues and fabrics, a transparent area compensation loss for plastic bags, and a reflective area suppression loss. The fuzzy IoU loss is constructed by calculating fuzzy intersections based on fuzzy set theory. To better learn the boundary characteristics of different object types, the object type-specific loss and fuzzy IoU loss are combined to dynamically adjust loss weights. For tissue / fabric texture-constrained loss, a texture-constrained loss function is designed to take advantage of the rich texture characteristics of tissues and fabrics. This function extracts local binary pattern (LBP) features from the image and calculates the difference between the predicted and true boundaries in texture space, encouraging the model to focus on texture details and improving the detection accuracy of boundaries for soft and porous materials. For plastic bags, a transparent area compensation loss is introduced to address the difficulty of detecting transparent areas in plastic bags. This loss function penalizes the model's predicted transparent area boundaries, encouraging the model to more accurately identify the edges of these areas. For plastic bags, a reflective area suppression loss is designed to reduce the interference of reflections on boundary detection. By analyzing the brightness and color features of the image, it identifies reflective areas and applies an additional penalty to boundary prediction errors in these areas, guiding the model to ignore reflective interference and focus on the true boundary. Based on fuzzy set theory, boundary prediction is treated as a fuzzy event, and the fuzzy intersection and union of the predicted boundary and the true boundary are calculated. Traditional IoU loss is sensitive to slight deviations from the boundary, while fuzzy IoU loss, by introducing the concept of fuzziness, can better handle boundary uncertainty, improving the model's tolerance and robustness to imprecise boundaries.

[0063] A multi-task learning network is built to simultaneously optimize both object classification and boundary detection. By sharing feature extraction layers, the model can learn semantic information related to boundaries, improving the accuracy of boundary detection.

[0064] The weights of various loss functions are dynamically adjusted based on the difficulty of detecting different types of foreign objects during training. For example, when the boundary detection error of a plastic bag is large, the weight of the loss function related to the plastic bag is automatically increased, prompting the model to pay more attention to the boundary characteristics of this type of foreign object.

[0065] The model generates initial boundary predictions based on the input image. It then optimizes and refines the predicted boundaries by combining them with enhanced edge information obtained through targeted image processing. Non-maximum suppression (NMS) is used to remove redundant boundary predictions. The final boundary detection results are then evaluated for confidence based on the fuzzy IoU loss calculation to ensure that the output boundaries are accurate and reliable.

[0066] Through the above-mentioned targeted image processing and model training inference strategies, the rail switch foreign object boundary detection system can effectively cope with the challenges of different types of foreign objects and achieve high-precision boundary detection of common foreign objects such as paper towels, fabrics, and plastic bags.

[0067] In the detailed feature extraction step, a deep learning network is used to extract corresponding texture features in the first and second boundary regions. These extracted texture features are mapped one-to-one to form a mapping relationship group. This mapping relationship group is then analyzed to determine feature change rates, which include texture contrast ratio, color contrast ratio, and shape contrast ratio. Based on these texture contrast ratios, a multi-feature analysis strategy is then used to determine the specific foreign object type. A convolutional neural network (such as ResNet or DenseNet) is used to analyze the images of the first and second boundary regions and extract texture features. An auxiliary extraction mechanism calculates pixel-level similarity between the first and second boundary regions, generating a mapping weight matrix. For each feature point in the first boundary region, the corresponding point with the most similar texture in the second boundary region is found, forming a set of "feature point pairs." The weight matrix reflects the strength of correlation between features in different regions and suppresses background noise interference. Texture features include foreign object grain features, foreign object color features, and foreign object deformation features. The network learns and captures subtle surface textures of foreign objects, such as the fiber texture of a tissue, the wrinkles of a plastic bag, and scratches on a metal surface. At the same time, by comparing the texture features of the two sets of images, the texture changes of the foreign objects under different viewing angles are analyzed, such as whether there is texture blurring caused by reflection, whether new lines are generated due to rolling, etc. A network structure combined with spatial pyramid pooling (SPP) is used to analyze the shape changes of foreign objects from multiple scales. The system not only calculates the overall geometric indicators such as the aspect ratio and roundness of the foreign objects, but also focuses on local deformations, such as the bending angle of tree branches and the degree of stretching of plastic bags. By comparing the shape contours of the two boundary areas, it is determined whether the foreign objects undergo dynamic changes such as squeezing and twisting when the train passes. The texture features are analyzed through mapping relationship groups to obtain the feature change rate, which includes the image feature change rate.

[0068] The multi-feature analysis strategy includes combining the feature change rate with the probability weight of the coarse classification, and using a weighted voting mechanism to output the foreign body type. When there is a conflict between the feature change rate and the coarse classification result, the weight is dynamically adjusted according to the credibility of different features.

[0069] After the image processing steps are complete, the system extracts the probability values ​​for each specific category from the coarse classification results from the semantic analysis strategy. For example, if the coarse classification identifies a foreign object as a flexible foreign object, the system outputs a distribution vector containing the probabilities of categories such as paper towels, plastic bags, and fabric. For example, the probability vectors for paper towels, plastic bags, and fabric are [0.3, 0.5, 0.2].

[0070] From the first boundary area of ​​the foreign body image, the local binary pattern (LBP) and gray level co-occurrence matrix (GLCM) methods are used to extract texture features. Taking LBP as an example, binary codes are generated by comparing the grayscale relationship between the central pixel and the neighboring pixels to describe the texture details (such as the fiber structure of paper towels and the grid pattern of plastic bags). The extracted texture feature vector Texture characteristics of standard samples corresponding to foreign body types in the database To calculate similarity, the cosine similarity formula is commonly used:

[0071]

[0072] Calculated This is the texture contrast ratio. The closer its value is to 1, the more similar the texture characteristics are to the standard sample.

[0073] Compare and analyze the first boundary area of ​​the foreign body image with the second boundary area of ​​the image to be analyzed, calculate the change rate of geometric parameters such as length, width, and area, and extract the shape change characteristics such as the degree of wrinkles and edge distortion of the indentation. For example, calculate the area change rate

[0074]

[0075] is the area after compression, is the area before compression. These deformation features are compared to the contours of corresponding standard samples in the database, using cosine similarity or Euclidean distance to obtain deformation feature values. Incorporating contour features of the first boundary region (such as perimeter and concavity) further enhances the accuracy of deformation analysis.

[0076] Histogram equalization is performed on the first and second boundary regions to improve the stability of image features. The mean, variance, and skewness of the RGB channels are calculated to describe the concentration and dispersion of color distribution. The color contrast ratio is then calculated by combining these feature change rates. A larger color contrast ratio indicates a more pronounced color difference between the first and second boundary regions, which can reduce color shifts and hue shifts caused by foreign object reflections.

[0077] Voting weights are assigned to each specific type based on the probability value of the coarse classification. The higher the probability value of a category, the greater its weight is, reflecting the priority of the category in the preliminary judgment. Weights are assigned to the texture feature contrast ratio, color contrast ratio, and shape contrast ratio to highlight the importance of different features in classification. For each possible foreign body type, its voting score is calculated, and the voting scores calculated for all categories are compared. The category with the highest score is selected as the final foreign body type output. When there is an obvious conflict between the texture contrast ratio, color contrast ratio, shape contrast ratio, or coarse classification result features, the feature change rate is adjusted according to the different degrees of dependence of different foreign body features on the feature change rate. When the surface texture is obviously a paper towel, the texture feature weight ratio is increased. For plastic bags that are easily reflective, the shape feature weight ratio is increased to reduce the deviation effect of texture features and shape contrast features caused by reflective interference.

[0078] The system also includes a position identification step, determining the location of the foreign object based on the position of the contour line of the first boundary area and the contour line of the track switch. The deformation trend of the foreign object is analyzed based on the contour line of the second boundary area and the contour line of the first boundary area. The position of the deformed foreign object is analyzed, and a danger warning is generated based on the type of foreign object. After completing the foreign object boundary identification in the image processing step, the Suzuki contour tracking algorithm is used to process the first boundary area identified in the foreign object image. Starting from the edge pixels of the image, the algorithm uses an eight-neighborhood search to track the boundary pixels in a clockwise or counterclockwise direction, gradually constructing a complete contour line. During the tracking process, the algorithm records the coordinates of each pixel, ultimately forming a pixel coordinate sequence that describes the shape and extent of the foreign object before compression. To improve the accuracy of the contour line, the boundary identification results can be subjected to morphological refinement to remove edge burrs and redundant pixels, making the contour smoother and more accurate. The system pre-stores a high-precision standard track switch contour model. This model is constructed based on track design drawings and actual survey data and contains precise geometric information for each switch component (point rail, stock rail, frog, etc.). During real-time detection, the image registration technology based on feature point matching is first used to extract SIFT or SURF feature points from the collected track switch images, match them with the feature points in the standard model, calculate the image rotation, translation and scaling parameters, and align the real-time image with the standard model.

[0079] Subsequently, the Canny edge detection algorithm is used to extract the actual contour line of the track turnout. Due to interference factors such as oil stains, wear, and uneven lighting on the track surface, the image is first subjected to adaptive histogram equalization processing before Canny detection to enhance the image contrast. At the same time, according to the characteristics of the track image, the high and low thresholds of the Canny algorithm are dynamically adjusted (for example, by calculating the statistical distribution of the image gradient, the threshold is automatically set to 1.5 times and 3 times the gradient mean) to ensure that the turnout contour can be fully extracted and noise interference can be effectively suppressed. Finally, the morphological closing operation is used to connect the broken parts on the contour line, and the opening operation is used to remove small noise areas to obtain the accurate track turnout contour line. The track turnout area is divided into multiple sub-areas such as the point rail area, the switch area, the guardrail area, and the area near the switch according to its functional and structural characteristics, and a corresponding coordinate range and geometric model are established for each sub-area. The coordinates of the geometric center (center of mass) of the contour line of the first boundary area are calculated using the following formula:

[0080]

[0081] in, is the coordinate of the i-th point on the contour line, and n is the total number of contour points. By determining which subregion the centroid coordinates fall within, the approximate location of the foreign object can be quickly determined. The overlapping area between the first boundary area contour line and the track switch contour line is calculated to quantify the impact of the foreign object on track operation. Using the pixel counting method, the areas containing the two contour lines are converted into binary images. A logical AND operation is performed to obtain the binary image of the overlapping portion. The number of pixels in the overlapping area is counted and then converted to the actual area based on the actual image resolution. If the overlapping area exceeds a preset threshold (e.g., 30%), it indicates that the foreign object poses a significant threat to the normal operation of the track, and the risk level assessment needs to be increased. Using the principles of affine transformation and perspective transformation in image processing, the second boundary area contour line (after compression) is aligned with the first boundary area contour line (before compression) to accurately compare their shape and position changes. Specifically, the contour moments are used to calculate the center position and principal axis direction of the two contour lines. Then, based on the center position offset and principal axis rotation angle, an affine transformation matrix is ​​constructed to transform the second boundary area contour line to the same coordinate system as the first boundary area contour line. Calculate the displacement vector between the two contour lines, including the translation distance and rotation angle. The translation distance is calculated by calculating the Euclidean distance between the center of the second boundary region's contour line and the center of the first boundary region's contour line after the transformation; the rotation angle is determined based on the change in the main axis direction. Simultaneously, analyze the extension direction of the indentation, for example, by calculating the tangent direction of the indentation contour line, to determine the foreign object's tendency to slide or roll under train pressure. If the second boundary region's contour line translates toward the track center and rotates relative to the first boundary region's contour line, this indicates that the foreign object has moved toward a critical track area under pressure, potentially affecting the normal switching of the turnout or the train's wheel-rail contact. In this case, the risk level should be further increased. The foreign object type is flexible (such as plastic bags, fabric, etc.) and completely covers critical track areas (such as the contact area between the point rail and the stock rail, and the frog throat area). Although flexible foreign objects are soft, they can become entangled in turnout components under the influence of train airflow and wheel pressure, affecting the normal switching of the turnout and causing loss of train control. After the system triggers a high-risk warning, it also performs an emergency braking operation and notifies the ground control center for processing; if the flexible foreign object partially covers the track but does not completely block it, and is located in a non-critical area (such as the guardrail area or the edge of the track), a medium-risk warning is generated. At this time, the system sends a deceleration instruction to the train, reducing the train speed to a safe speed (such as 20km / h) to reduce the impact of foreign objects on train operation. At the same time, the dispatching system arranges inspection personnel to carry tools to the scene as soon as possible to clean up foreign objects, and strengthen monitoring of the area during subsequent operations; for small and remote foreign objects, such as small pieces of paper and small stones on the edge of the track, a low-risk warning is issued. The system records the location information, type and image data of the foreign object and incorporates it into the daily inspection plan.During the next track inspection, the inspection personnel will clean it up and conduct a key inspection of the area to ensure that foreign objects do not pose a potential threat to track operation.

[0082] like Figure 4 As shown, a verification step is also included. A verification module is also provided on the train. The verification module is provided behind the second camera group. The verification module includes a main camera, an auxiliary camera and a laser radar. According to the type of foreign matter, the lighting intensity and camera parameters of the auxiliary camera are adjusted, and the image data collected by the main camera and the auxiliary camera and the laser electrical signal collected by the laser radar are combined. The data is verified based on multimodal data fusion processing, and based on the verification result, the abnormality analysis result is obtained, and compared with the foreign matter type to determine whether it is consistent. When the target detection model preliminarily determines the type of foreign matter, the verification module will immediately automatically adjust the working parameters of the auxiliary camera according to the characteristics of the foreign matter. The auxiliary camera is provided with a time-sharing strobe function. Using a group of light sources of different colors, by precisely controlling the strobe of each light source at different time points, image information under light sources of different colors is obtained, and according to the properties of the object, image photos under different lighting effects are obtained.

[0083] For highly reflective objects: If a metal object is detected, the smooth surface of such objects easily produces reflections, resulting in an overly bright image and loss of detail. In this case, the auxiliary camera will reduce the lighting intensity to avoid strong light exposure that exacerbates the reflection problem, ensuring that the texture and contours of the metal object's surface are clearly presented in the image.

[0084] Targeting foreign objects with low color contrast: Light-colored foreign objects, such as white tissues, that are similar in color to the background are difficult to distinguish against the track. The auxiliary camera selects warm-toned images and uses different lighting colors to enhance the color difference between the foreign object and the background. It also optimizes camera parameters such as exposure time and contrast to make the white tissue stand out more clearly in the image, facilitating subsequent analysis.

[0085] For other types of foreign objects: For foreign objects of different materials and shapes, such as plastic bags and tree branches, the auxiliary camera will also adjust parameters accordingly. For example, for thin, transparent plastic bags, the focus parameters will be adjusted to ensure the edges are clear; for irregularly shaped foreign objects such as tree branches, the shooting angle and field of view will be adjusted to fully capture the shape.

[0086] Deep learning algorithms (such as ResNet+FPN) are used to analyze dual-camera images, extracting over 20 visual features such as edge contours, texture features (LBP algorithm), and color histograms. The LiDAR point cloud data is subjected to denoising (voxel filtering), and geometric features such as point cloud density, surface curvature, and normal vectors are calculated to construct a 3D spatial description model. The point cloud data is projected onto the image plane to generate a depth image with spatial information, enabling a preliminary correlation between visual and spatial data. Image and point cloud features are classified and predicted separately, and the results are integrated using a weighted voting algorithm to ultimately output a unified foreign object description vector.

[0087] The fused foreign body description vector is matched with the standard feature template in the system database:

[0088] Successful match: If the similarity is greater than 90% and is consistent with the initial judgment of the target detection model, the verification is passed, the system records the foreign object information and generates a regular alarm;

[0089] Matching failure: Exception handling is triggered when the following situations occur:

[0090] Feature similarity is less than 80%; or there are contradictions in the multimodal data (for example, visual judgment is plastic, and the lidar shows metal reflective characteristics), which conflicts with the initial judgment of the foreign body type.

[0091] At this point, the system immediately starts the secondary detection program and calls historical data for comparative analysis; if confirmation is still not possible, a high-priority alarm is sent to the train dispatching center, accompanied by detailed multimodal data, prompting manual intervention for verification.

[0092] like Figure 3 As shown, a track switch foreign body detection and classification analysis system includes:

[0093] The switch foreign object recognition module is equipped with a first camera group at the front of the train. It uses the first camera group to shoot track video and perform frame detection. It then uses a preset target detection model to identify the switch position and a preset foreign object detection model to calibrate foreign objects in the switch. The corresponding image in the video is captured as a foreign object map.

[0094] An image acquisition module captures an image calibrated to a foreign object in the turnout through a second camera set behind the train wheels as an image to be analyzed;

[0095] The image processing module uses a preset wheel rolling track detection model to extract a second boundary area on the foreign object in the image to be analyzed, and roughly classifies the foreign objects in the image to be analyzed to obtain a rough classification result. The rough classification result includes hard objects and soft foreign objects. The boundary detection strategy is selected based on the rough classification result in the foreign object image to identify the boundary area as the first boundary area;

[0096] The detail feature extraction module extracts corresponding texture features in the first boundary area and the second boundary area respectively through a deep learning network, and maps the extracted texture features one by one to obtain a mapping relationship group. The feature change rate is then obtained through analysis of the mapping relationship group. The feature change rate includes texture contrast ratio, color contrast ratio and shape contrast ratio. The specific foreign body type is then obtained through analysis of the texture contrast ratio, color contrast ratio and shape contrast ratio through a multi-feature analysis strategy.

[0097] The above are merely preferred embodiments of the present invention. The scope of protection of the present invention is not limited to the above embodiments. All technical solutions based on the principles of the present invention are within the scope of protection of the present invention. It should be noted that improvements and modifications that do not depart from the principles of the present invention are within the scope of protection of the present invention.

Claims

1. A method for detecting and analyzing foreign objects in a track switch, characterized in that: include: In the switch foreign object identification step, a first camera group is provided at the front end of the train. The first camera group is used to shoot track video and perform frame detection. The switch position is then identified using a preset target detection model. Foreign objects in the switch are calibrated using a preset foreign object detection model, and the corresponding image in the video is captured as a foreign object map. In an image acquisition step, a second camera group arranged behind the train wheels captures an image calibrated to the foreign object in the turnout as an image to be analyzed; an image processing step of extracting a second boundary region of the wheel on the foreign object in the image to be analyzed using a preset wheel rolling track detection model, performing coarse classification on the foreign objects in the image to be analyzed to obtain a coarse classification result, the coarse classification result including hard objects and soft foreign objects, and selecting a boundary detection strategy in the foreign object map based on the coarse classification result to identify the boundary region as the first boundary region; The detail feature extraction step extracts corresponding texture features in the first boundary area and the second boundary area respectively through a deep learning network, and maps the extracted texture features one by one to obtain a mapping relationship group, and then obtains the feature change rate through the mapping relationship group analysis. The feature change rate includes texture contrast ratio, color contrast ratio and shape contrast ratio, and then the specific foreign body type is obtained through multi-feature analysis strategy based on the texture contrast ratio, color contrast ratio and shape contrast ratio; the multi-feature analysis strategy includes combining the feature change rate with the probability weight of the coarse classification, and using a weighted voting mechanism to output the foreign body type. When there is a conflict between the feature change rate and the coarse classification result, the weight is dynamically adjusted according to the credibility of different features.

2. A method for detecting and analyzing foreign objects in a track switch according to claim 1, characterized in that: The flexible foreign objects include paper towels, plastic bags, and fabrics. When the rough classification result is a flexible foreign object, a probability distribution vector representing the probability value of each specific category in the flexible foreign object is output. The database is provided with foreign object type labels and priority boundary recognition strategies and standard sample features corresponding to each category. According to the distribution vector of the probability value of each specific category in the rough classification result, the database is queried to obtain the corresponding boundary detection strategy.

3. A method for detecting and analyzing foreign matter in a track switch according to claim 2, characterized in that: When the probability values ​​of multiple specific categories all reach the threshold, the first boundary area corresponding to each category of foreign objects is obtained according to each category recognition strategy, and the foreign object boundaries in the foreign object image are analyzed through the image comparison strategy. The first boundary area is determined according to the image similarity, and the probability value is adjusted according to the comparison result.

4. A method for detecting and analyzing foreign matter in a track switch according to claim 2, characterized in that: The boundary detection strategy includes: In a targeted image processing step, based on the coarse classification results, when the foreign object is a tissue or fabric, image contrast is enhanced and edges are preserved, using multi-scale edge detection combined with morphological operations to strengthen edges; when the foreign object is a plastic bag, reflections are suppressed through a polarization strategy, and boundaries are restored using a deep learning-based deblurring model and interpolation; The model training and inference steps dynamically adjust the loss weights by combining foreign object type specificity and fuzzy IoU loss. The foreign object type specificity loss includes texture constraint loss for paper towels and fabrics, transparent area compensation loss for plastic bags, and reflective area suppression loss. The fuzzy IoU loss calculates and constructs fuzzy intersection based on fuzzy set theory.

5. A method for detecting and analyzing foreign matter in a track switch according to claim 1, characterized in that: It also includes a position identification step, which determines the location of the foreign object based on the position of the contour line of the first boundary area and the contour line of the track switch, and analyzes the deformation trend of the foreign object based on the contour line of the second boundary area and the contour line of the first boundary area, analyzes the position of the deformed foreign object, and generates a danger warning based on the type of foreign object.

6. A method for detecting and analyzing foreign matter in a track switch according to claim 1, characterized in that: The method also includes a verification step. A verification module is also provided on the train, and the verification module is provided behind the second camera group. The verification module includes a main camera, an auxiliary camera and a laser radar. According to the type of the foreign object, the lighting intensity and camera parameters of the auxiliary camera are adjusted, and the image data collected by the main camera and the auxiliary camera and the laser electrical signal collected by the laser radar are combined. The data is verified based on multimodal data fusion processing, and based on the verification result, the abnormality analysis result is obtained, and compared with the foreign object type to determine whether it is consistent.

7. A method for detecting and analyzing foreign matter in a track switch according to claim 1, characterized in that: The target detection model is equipped with a small target positioning strategy, including utilizing the network architecture for feature extraction and embedding an attention mechanism module in the network. The attention mechanism module includes a dynamic channel weight module, which assigns high weights to the key feature channels of the turnout dataset through the SE-Net structure, performs cross-scale channel interaction, filters the key features of the turnout target, and uses a spatial attention module to capture the global position of the turnout through global pooling and generate an attention map to locate the target area.

8. The method for detecting and analyzing foreign matter in a track switch according to claim 1, characterized in that: The second camera group is a high-definition camera. When the first camera group detects a turnout, the second camera is adjusted so that the second camera captures an image of the area where the turnout is located.

9. A track switch foreign body detection and classification analysis system, characterized in that: include: The switch foreign object recognition module is equipped with a first camera group at the front of the train. It uses the first camera group to shoot track video and perform frame detection. It then uses a preset target detection model to identify the switch position and a preset foreign object detection model to calibrate foreign objects in the switch. The corresponding image in the video is captured as a foreign object map. An image acquisition module captures an image calibrated to a foreign object in the turnout through a second camera set behind the train wheels as an image to be analyzed; an image processing module, which extracts a second boundary region on the foreign object where the wheel has rolled over the foreign object in the image to be analyzed using a preset wheel rolling track detection model, and roughly classifies the foreign objects in the image to be analyzed to obtain a rough classification result, wherein the rough classification result includes hard objects and soft foreign objects, and selects a boundary detection strategy in the foreign object map based on the rough classification result to identify the boundary region as the first boundary region; The detail feature extraction module extracts corresponding texture features in the first boundary area and the second boundary area respectively through a deep learning network, and maps the extracted texture features one by one to obtain a mapping relationship group, and then obtains the feature change rate through the mapping relationship group analysis. The feature change rate includes texture contrast ratio, color contrast ratio and shape contrast ratio, and then the specific foreign body type is obtained through multi-feature analysis strategy based on the texture contrast ratio, color contrast ratio and shape contrast ratio; the multi-feature analysis strategy includes combining the feature change rate with the probability weight of the coarse classification, and using a weighted voting mechanism to output the foreign body type. When there is a conflict between the feature change rate and the coarse classification result, the weight is dynamically adjusted according to the credibility of different features.

Citation Information

Patent Citations

  • Rail train anti-collision identification device and method in unmanned state

    CN118918552A

  • Method and system for railway obstacle detection based on rail segmentation

    WO2020012475A1