Boundary dynamic sensing detection method for railway weak light environment

By using the detection method of the lighting correction network and multi-scale feature extraction module in the low-light environment of railways, the accuracy and accuracy of railway perimeter intrusion detection in the low-light environment is solved, and high-precision target detection and monitoring are achieved to meet the railway safety monitoring needs.

CN120071244APending Publication Date: 2025-05-30BEIJING JIAOTONG UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510123408.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-26
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

In low-light environments, the accuracy and accuracy of railway perimeter intrusion detection are low, and the prior art is difficult to effectively identify and deal with small targets, and the adaptability in complex environments is insufficient.

Method used

A boundary dynamic perception detection method for railway low-light environment is adopted, and the image to be detected is brightened in all regions, dark correction and edge details enhancement through the illumination correction network module, and the image after correction is detected through the detection module. The method includes a CBR module, a SC module, an I module and a CB module, which is used for feature extraction and lighting correction, and combines the first backbone feature extraction module, a second backbone feature extraction module and a detection head to realize multi-scale feature extraction and object detection.

Benefits of technology

It improves the target detection accuracy and monitoring visualization level in low-light environments, enhances the understanding of complex scenes, and can accurately identify the perimeter idle personnel under extremely low light conditions to meet the railway perimeter safety monitoring needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120071244A_ABST
    Figure CN120071244A_ABST
Patent Text Reader

Abstract

The invention provides a railway weak light environment boundary dynamic perception detection method comprising the following steps: obtaining a to-be-detected image of a railway perimeter in a weak light environment and inputting the to-be-detected image to a boundary dynamic perception detection network, the boundary dynamic perception detection network comprising an illumination correction network module and a detection module; performing global brightening, darkness correction and edge detail enhancement on the to-be-detected image through the illumination correction network module to obtain a corrected image; and performing target detection on the corrected image through the detection module to obtain a detection result. According to the method provided by the invention, the surrounding miscellaneous personnel can still be accurately identified under an extremely low illumination condition, so that the safety detection capability of a railway at night or in a tunnel is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of railway traffic safety, and particularly to a method for dynamically perceiving and detecting the boundary of a railway low-light environment. Background Art

[0002] In recent years, the Chinese railway network has expanded rapidly. As of November 2024, the total railway mileage in the country reached 1.6 million kilometers, of which the high-speed railway mileage reached 450,000 kilometers, bringing huge challenges to perimeter security. The risk of intrusion events is increasing continuously, threatening the safety of train operation, especially under low-light conditions such as at night or in tunnels. Traditional monitoring and detection methods often fail to work. Most traditional perimeter security strategies usually rely on a combination of manual and technical defenses, mainly using visible light cameras. However, these systems are difficult to function in low visibility environments, resulting in frequent false alarms and missed detections. The derailment of the BNSF freight train in Colorado, USA in 2023 and the tragic collision of the 42109 freight train in 2024 are examples.

[0003] Multiple safety accidents caused by nighttime railway perimeter intrusions have exposed the deficiencies in perimeter intrusion detection in the railway low-light environment. Since most of the visible light cameras used along the Chinese railway are ordinary ones, the image clarity and contrast of the monitoring videos in the low-light environment are significantly reduced, and the detail features are severely lost, affecting the accuracy and precision of the existing perimeter intrusion detection algorithms. Therefore, it is urgent to carry out research on railway perimeter intrusion detection algorithms in the low-light environment, improve the accuracy of nighttime intrusion detection in key sections, and provide guarantee for all-weather railway operation safety.

[0004] With the rapid development of image processing technologies based on deep learning and neural networks. The theory based on Retinex enhances low-light images by estimating the true reflection characteristics of the scene and defines an unsupervised training loss, thus improving the object detection performance in non-uniform and low-light environments. Other models, such as the ULRE-Net and the enhanced CycleGAN model, introduce the adaptive instance normalization (AdaIN) and the detail enhancement module. These algorithms improve object detection in low-light environments through innovative neural network architectures. However, their complex network structures cannot meet the real-time requirements of actual railway applications. In addition, they also ignore the missing features in dealing with small targets under low-light conditions and lack adaptability to complex environments. Similarly, systems using lidar technology, infrared and visible light image fusion technology, and millimeter wave radar technology have made significant progress in perimeter intrusion detection in the railway low-light environment, achieving high detection rates and improving reliability. However, these solutions are often limited in terms of cost, resource consumption, and adaptability to complex environments.

[0005] In the actual application of railways, in low-light situations such as dim light or tunnels, the low-frequency features of railway images, such as brightness and contrast, are low, making it difficult to capture the detailed features of the images, thus affecting the accurate detection of targets. Secondly, the railway line usually includes a complex background, including buildings, signal equipment, etc., which are easily confused with detection targets such as trains and personnel. Especially in low-light conditions, it is more difficult to distinguish the background from the targets. In addition, objects in the railway environment (such as trains) usually appear in high-speed motion, easily causing blurring and dynamic distortion during shooting, and it is necessary to process targets with different scales and deformations.

[0006] Therefore, how to apply existing advanced technologies to the dynamic boundary perception detection of railways in low-light environments to meet the requirements of railway perimeter security monitoring remains an unsolved problem. Summary of the Invention

[0007] The purpose of the present invention is to overcome the deficiencies of the prior art and provide a method for dynamically perceiving and detecting the boundaries of railways in low-light environments.

[0008] To achieve the above purpose, the present invention adopts the following technical solutions.

[0009] In the first aspect, the present invention provides a method for dynamically perceiving and detecting the boundaries of railways in low-light environments, including:

[0010] Obtain the image to be detected of the railway perimeter in a low-light environment and input it into the boundary dynamic perception detection network, where the boundary dynamic perception detection network includes an illumination correction network module and a detection module;

[0011] Perform global brightening, dark correction, and edge detail enhancement on the image to be detected through the illumination correction network module to obtain a corrected image;

[0012] Perform target detection on the corrected image through the detection module to obtain a detection result.

[0013] Furthermore, feature extraction is performed on the image to be detected through the CBR module and the SC module to obtain a feature extraction result g(x k ):

[0014] g(x k ) = SC(CBR 1,1 (x k ))

[0015]

[0016] where x k represents the k-th frame of the image to be detected, and g(x k) is the feature extraction result of the k-th frame image to be detected. The CBR module consists of a convolutional layer, a batch normalization layer, and a ReLU activation function. The SC module is used for image feature extraction operations. t represents the index of a stage or time step, which is used to describe the multi-stage iteration in the feature processing process. y represents the image or feature input, and x t represents the input for element-wise multiplication of the intermediate feature obtained by step-by-step processing in the t-th stage and y, and z t represents the result obtained through element-wise multiplication processing, and κ θ is the introduced parameterized operator, and s t represents the transformed input calculated through κ θ and v t represents another transformed stage output composed of y + s t ;

[0017] The feature extraction result g(x k ) is corrected for illumination through the I module to obtain the illumination-corrected feature I(x k )

[0018]

[0019] where g t represents the feature in the t-th stage, represents the parameterized mapping function, and u t is the residual term in the t-th stage. The illumination-related information is accumulated into the current feature g t through the residual u t to gradually correct the illumination, and g t+1 represents the feature in the t + 1-th stage;

[0020] The illumination-corrected feature I(x k ) is processed through the CB module to obtain the corrected image F(x k )

[0021] F(x k ) = CB 1,1 (I(g(x k )))

[0022] Furthermore, the detection module includes a first backbone feature extraction module, a second backbone feature extraction module, and a detection head;

[0023] The corrected image is subjected to object detection through the detection module to obtain a detection result, including:

[0024] The corrected image is subjected to multi-scale feature extraction through the first backbone feature extraction module to obtain multiple first backbone features;

[0025] The corrected image and the first backbone features are subjected to multi-scale feature extraction by the second backbone feature extraction module to obtain a plurality of second backbone features;

[0026] The first backbone features and the second backbone features are input into the detection head to obtain a detection result.

[0027] Further, the multi-scale feature extraction of the corrected image by the first backbone feature extraction module to obtain first backbone features includes:

[0028] Performing multi-scale feature extraction on the corrected image through a first backbone feature extraction formula to obtain first backbone features:

[0029]

[0030] P 3 ,P 4 ,P 5 =L main (x i )

[0031] where RepCSP is a multi-stage enhanced feature extraction module, which includes segmented convolution, residual connection, cross-stage connection, feature fusion and hierarchical operation, P 3 ,P 4 ,P 5 are the first backbone features of different scales respectively; L main( ·) represents the main loss function, which is used to optimize the learning process of the first backbone feature extraction module;

[0032] The multi-scale feature extraction of the corrected image and the first backbone features by the second backbone feature extraction module to obtain second backbone features includes:

[0033] Performing multi-scale feature extraction on the corrected image and the first backbone features through a second backbone feature extraction formula to obtain second backbone features:

[0034]

[0035] C 3 ,C 4 ,C 5 =L aux (P 3 ,P 4 ,P 5 ,x i )

[0036] Wherein, Laux(·) is an auxiliary loss function, which processes the first backbone features P3, P4, P5 and the corrected image x i to obtain the second backbone features C3, C4, C5 at different scales. X includes the corrected image x i and the first backbone features P3, P4, P5. CBL is a basic convolutional feature extraction module, including combined operations of convolution, batch normalization and activation functions. DS represents a downsampling operation.

[0037] Furthermore, the detection head includes a first sub-detection head and a second sub-detection head;

[0038] The step of inputting the first backbone feature and the second backbone feature into the detection head to obtain a detection result includes:

[0039] Fusing and detecting multiple first backbone features through the first sub-detection head to obtain multiple initial detection results;

[0040] Fusing and detecting the second backbone feature and the initial detection results through the second sub-detection head to obtain a final detection result.

[0041] Furthermore, the second sub-detection head includes a plurality of convolutional blocks, a pooling-convolution fusion module and a feature recombination module;

[0042] The step of fusing and detecting the second backbone feature and the initial detection results through the second sub-detection head to obtain a final detection result includes:

[0043] Convolving the second backbone feature and the initial detection results respectively through the plurality of convolutional blocks to obtain target features;

[0044] Performing max pooling and average pooling operations on the target features through the pooling-convolution fusion module to obtain key information;

[0045] Performing upsampling and downsampling operations on the key information through the feature recombination module to obtain recombined features, and obtaining a prediction result based on the recombined features.

[0046] In a second aspect, the present invention further provides a boundary dynamic perception detection device for a railway low-light environment, including:

[0047] An image acquisition module, configured to acquire an image to be detected of a railway perimeter in a low-light environment and input it into a boundary dynamic perception detection network, where the boundary dynamic perception detection network includes an illumination correction network module and a detection module;

[0048] An image correction module, configured to perform global brightening, shading correction, and edge detail enhancement on the image to be detected through the illumination correction network module, so as to obtain a corrected image;

[0049] A detection module, configured to perform target detection on the corrected image through the detection module, so as to obtain a detection result.

[0050] In a third aspect, the present invention further provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the above-mentioned method is implemented.

[0051] In a fourth aspect, the present invention further provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the above-mentioned method is implemented.

[0052] In a fifth aspect, the present invention further provides a computer program product, including a computer program, characterized in that when the computer program is executed by a processor, the above-mentioned method is implemented.

[0053] Advantages of the present invention: The boundary dynamic perception detection method for the railway low-light environment provided by the present invention performs global brightening, shading correction, and edge detail enhancement on the image to be detected through the illumination correction network module, pays attention to the intrusion of idle personnel in the detection area in real time, improves the target detection accuracy and monitoring visualization level in the track area, discovers potential risks of railway perimeter intrusion, so as to improve the detection accuracy of the subsequent detection module, and ensure that the perimeter idle personnel can still be accurately identified under extremely low light conditions, meeting the requirements of railway perimeter security monitoring.

[0054] Additional aspects and advantages of the present invention will be given in part in the following description, and these will become obvious from the following description, or can be understood through the practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0055] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0056] Figure 1 It is one of the flow diagrams of a boundary dynamic perception detection method for a railway low-light environment provided by an embodiment of the present invention;

[0057] Figure 2 It is the structural diagram of the boundary dynamic perception detection network provided by an embodiment of the present invention;

[0058] Figure 3 Schematic diagram of the lighting correction network module provided by an embodiment of the present invention;

[0059] Figure 4 Schematic diagram of global image brightening comparison provided by an embodiment of the present invention;

[0060] Figure 5 Schematic diagram of comparison between the ROI algorithm and the SCINet algorithm provided by an embodiment of the present invention;

[0061] Figure 6 Schematic diagram of the structure of the second sub-detection head provided by an embodiment of the present invention;

[0062] Figure 7 is Figure 6 Schematic diagram of the structure of the PCRC module in

[0063] Figure 8 Schematic diagram of comparison of detection results of each object detection algorithm under different low-light data provided by an embodiment of the present invention;

[0064] Figure 9 Schematic diagram of small object detection in the railway low-light scene provided by an embodiment of the present invention. Detailed implementation manners

[0065] The following details the implementation manners of the present invention. Examples of the implementation manners are shown in the accompanying drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The implementation manners described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and cannot be construed as a limitation of the present invention.

[0066] Those skilled in the art of the present technology can understand that unless specifically stated otherwise, the singular forms "a", "an" and "the" used herein may also include the plural forms. It should be further understood that the term "comprising" used in the specification of the present invention means the presence of features, integers, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or their groups. It should be understood that when we say an element is "connected" or "coupled" to another element, it can be directly connected or coupled to other elements, or there may also be intermediate elements. In addition, the "connection" or "coupling" used herein may include wireless connection or coupling. The phrase "and / or" used herein includes any and all combinations of one or more of the associated listed items.

[0067] Those skilled in the art can understand that, unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as commonly understood by those of ordinary skill in the art to which this invention belongs. It should also be understood that terms such as those defined in a general dictionary should be understood to have a meaning consistent with the meaning in the context of the prior art, and will not be interpreted with an idealized or overly formal meaning unless defined as such here.

[0068] For ease of understanding the embodiments of the present invention, the following will further explain with several specific embodiments in conjunction with the accompanying drawings, and each embodiment does not constitute a limitation on the embodiments of the present invention.

[0069] Embodiment 1

[0070] See Figures 1 to 2 , a method for dynamically perceiving and detecting the boundary of a weak light environment of a railway, comprising the following steps:

[0071] S101, obtaining an image to be detected of the railway perimeter in a weak light environment and inputting it into a boundary dynamic perception detection network (abbreviated as BDPNet), where the boundary dynamic perception detection network includes an illumination correction network module and a detection module. The boundary dynamic perception detection network is pre-trained according to training data and corresponding label information.

[0072] Specifically, video images along the railway under low light conditions, especially in key areas such as entrances, exits, tunnels, and bridges, are obtained, and basic information such as the time to which the video belongs and the monitored geographical location is extracted.

[0073] S102, performing global brightening, dark correction, and edge detail enhancement on the image to be detected through the illumination correction network module (abbreviated as SCINet module) to obtain a corrected image.

[0074] In this step, the illumination correction network module enables the image to remain clearly visible even under extremely low light conditions. This module uses deep learning technology to learn image features under different light conditions and automatically optimizes image processing parameters. The illumination-corrected image uses a weight sharing mechanism and a multi-layer network structure for edge enhancement, improving the detail information of the image and providing better input for subsequent object detection. Schematically, Figure 4 shows the effects before and after illumination correction under different light conditions. Figure 5 The leftmost side of Figure 5 shows the original image, and the upper right image in Figure 5The lower-right image in shows the result after processing by the SCINet module, further optimizing the brightness and clarity of the target area, improving the target detection effect in the image, especially in the dark area, where details are better presented.

[0075] S103. Perform target detection on the corrected image through the detection module to obtain a detection result.

[0076] The boundary dynamic perception detection method for the weak light environment of railways provided by the embodiments of the present invention performs global brightening, dark correction, and edge detail enhancement on the image to be detected through the illumination correction network module, pays attention to the intrusion of unauthorized personnel in the detection area in real time, improves the target detection accuracy and monitoring visualization level in the track area, discovers potential risks of railway perimeter intrusion, so as to improve the detection accuracy of the subsequent detection module and meet the requirements of railway perimeter security monitoring.

[0077] In some embodiments, as Figure 3 shown, the illumination correction network module includes a CBR module, an SC module, an I module, and a CB module connected in sequence.

[0078] The step of performing global brightening, dark correction, and edge detail enhancement on the image to be detected through the illumination correction network module to obtain a corrected image includes:

[0079] Extract features from the image to be detected through the CBR module and the SC module to obtain a feature extraction result g(x k ):

[0080] g(x k ) = SC(CBR 1,1 (x k ))

[0081]

[0082] where x k represents the k-th frame of the image to be detected, g(x k ) is the feature extraction result of the k-th frame of the image to be detected. The CBR module consists of a convolutional layer, a batch normalization layer, and a ReLU activation function. The SC module is used for image feature extraction operations. t represents the index of a stage or time step, used to describe the multi-stage iteration in the feature processing process. y represents an image or feature input. x t represents the input for element-wise multiplication of the intermediate feature obtained by step-by-step processing in the t-th stage and y. z t represents the result after element-wise multiplication processing. κ θ is the introduced parameterized operator, and s t represents the result obtained through κ θThe calculated transformed input, v t represents another output of the transformation stage composed of y + s t

[0083] The I module performs illumination correction on the feature extraction result g(x k ) to obtain the illumination-corrected feature I(x k )

[0084]

[0085] where g t represents the feature at the t-th stage, represents the parametric mapping function, u t is the residual term at the t-th stage, and the illumination-related information is accumulated into the current feature g t through the residual u t to gradually correct the illumination, and g t+1 represents the feature at the (t + 1)-th stage.

[0086] It should be emphasized that the illumination correction network module adopts a weight sharing mechanism, that is, the same structure and weights are used in each stage.

[0087] The CB module processes the illumination-corrected feature I(x k ) to obtain the corrected image F(x k )

[0088] F(x k ) = CB 1,1 (I(g(x k )))

[0089] It should be emphasized that after each stage, the illumination correction network module is used to correct the output of the current stage. This correction can ensure that the output of each stage gradually stabilizes as the cascading process progresses and ensure that the quality of the final output does not degrade due to too many processing steps. In addition, due to the use of weight sharing, the model can effectively learn how to process low-light images during the training process. In railway field applications, only one illumination estimation module is needed, which greatly simplifies the model and improves the inference speed.

[0090] The illumination correction network module in the boundary dynamic perception detection method for the railway low-light environment provided by the embodiments of the present invention can automatically adjust the processing parameters, dynamically optimize the brightness and contrast according to the illumination conditions of the input image, and train the model using deep learning methods to adapt to different environmental illuminations.

[0091] In some embodiments, the detection module includes a first backbone feature extraction module, a second backbone feature extraction module, and a detection head.​

[0092] Performing object detection on the corrected image through the detection module to obtain a detection result, including:

[0093] Performing multi-scale feature extraction on the corrected image through the first backbone feature extraction module to obtain multiple first backbone features.

[0094] Performing multi-scale feature extraction on the corrected image and the first backbone features through the second backbone feature extraction module to obtain multiple second backbone features.

[0095] Inputting the first backbone features and the second backbone features into the detection head to obtain a detection result.

[0096] The boundary dynamic perception detection method for the railway low-light environment provided by the embodiments of the present invention extracts features of different scales through the first backbone feature extraction module and the second backbone feature extraction module to capture feature information at different levels of the image. The first backbone feature extraction module can capture the global features of the image, and the second backbone feature extraction module focuses on capturing local detail features. The output feature maps of the two networks are merged through a feature fusion module, which assigns corresponding weights by learning the importance of different feature maps to capture tiny and complex features in the image, thereby improving the understanding and recognition ability of the low-light image content and making it more stable when processing images of different resolutions and sizes.

[0097] In some embodiments, performing multi-scale feature extraction on the corrected image through the first backbone feature extraction module to obtain first backbone features includes:

[0098] Performing multi-scale feature extraction on the corrected image through the first backbone feature extraction formula to obtain first backbone features:

[0099]

[0100] P 3 ,P 4 ,P 5 =L main (x i )

[0101] In the formula, P3, P4, and P5 respectively represent different-scale first backbone features obtained through different layers or modules in the network, and multi-scale information of the image is captured by extracting feature maps of different sizes from the input image. RepCSP is a multi-stage enhanced feature extraction module, including segmented convolution, residual connection, cross-stage connection, feature fusion, and hierarchical operations, which optimizes the feature extraction process and improves the computational efficiency and the expression ability of multi-scale features. L main(·) represents the main loss function, which is used to optimize the learning process of the first backbone feature extraction module.

[0102] As Figure 2 shown, the first backbone feature extraction module includes multiple convolutional blocks and multiple RepCSPELAN4s. Through the computational processing of the convolutional blocks and RepCSPELAN4s, first backbone features P 3 , P 4 , P 5 are obtained.

[0103] The first backbone features P 3 , P 4 , P 5 are low-level to high-level features extracted layer by layer by the first backbone feature extraction module. Through layer-by-layer downsampling and a feature pyramid, multi-scale first feature maps are output. These first feature maps capture different levels of semantic information and target features at different scales. The first feature maps at each scale are used for the subsequent detection heads, thereby achieving the ability to detect multi-scale targets. This multi-scale feature pyramid design enhances the adaptability of the model to targets of different sizes.

[0104] The multi-scale feature extraction of the corrected image and the first backbone features by the second backbone feature extraction module to obtain second backbone features includes:

[0105] Performing multi-scale feature extraction on the corrected image and the first backbone features through the second backbone feature extraction formula to obtain second backbone features:

[0106]

[0107] C 3 , C 4 , C 5 = L aux (P 3 , P 4 , P 5 , x i )

[0108] In the formula, X includes the corrected image x i and the first backbone features P3, P4, P5, C 3 , C 4 , C 5They are the second backbone features at different scales. C3, C4, and C5 are the second backbone features obtained through operations such as downsampling, convolution, batch normalization, and activation, further enhancing the network's expressive ability. Laux(·) is an auxiliary loss function used to assist in the training of the network, usually used to improve the network's multi-task learning ability or multi-scale feature fusion. US represents the upsampling operation, which is used to increase the size of the feature map and is used to restore from a low-resolution feature map to a higher-resolution image or feature map. CBL is a basic convolutional feature extraction module, including a combined operation of convolution, batch normalization, and activation function, which is usually applied in a convolutional neural network to extract features and increase non-linearity. DS represents the downsampling operation, which is used to reduce the size of the feature map to extract higher-level features. Downsampling can be performed through pooling or convolution, etc.

[0109] It should be noted that x i is F(x k ), or it can also be the image after other processing steps on F(x k ).

[0110] As Figure 2 shown, the second backbone feature extraction module includes multiple convolutional blocks, multiple RepCSPELAN4, and multiple CBLinear. Among them, the input of the first CBLinear includes the corrected image and the first backbone feature P 3 , the input of the second CBLinear includes the output of the previous CBLinear and the first backbone feature P 4 , and the input of the third CBLinear includes the output of the previous CBLinear and the first backbone feature P 5 .

[0111] The second backbone feature extraction module reduces the spatial resolution of the feature map to extract higher-level features and expand the receptive field. Among them, the downsampled feature map is passed to CBLinear, which combines convolution, batch normalization, and linear mapping operations, reducing parameters and computational complexity while retaining features. Then, the feature map output by CBLinear will be upsampled, usually through bilinear interpolation or deconvolution, to restore a spatial resolution similar to the initial feature map. Upsampling can make feature maps at different scales consistent with the original resolution for subsequent feature fusion.

[0112] In addition, during the model training process, training is performed based on the total loss L total of the first backbone feature extraction module and the second backbone feature extraction module.

[0113] L total = L main + L aux

[0114] In the formula, L total represents the total loss function, that is, the loss of the entire network during the training process. L main is the loss of the first backbone feature extraction module, and L aux is the loss of the second backbone feature extraction module.

[0115] Finally, add the first backbone feature extraction module and the second backbone feature extraction module. The second backbone feature extraction module helps the network retain more useful feature information during deep training by generating gradient information at different levels, thereby reducing the problem of information loss in deep supervision.

[0116]

[0117] Among them, Y(i,j,k): represents the value of the k-th channel in the i-th row and j-th column of the output feature map Y. The output feature map is the result obtained by extracting features from the input image (or feature map) through a convolution operation.

[0118] X(i+m,j+n,l): represents the element in the l-th channel of the (i+m)-th row and (j+n)-th column of the input feature map X. That is to say, the convolution operation extracts a small area (determined by the convolution kernel size) from the input feature map for calculation.

[0119] W(m,n,l,k): represents the weight between the l-th input channel and the k-th output channel in the m-th row and n-th column of the convolution kernel W. The convolution kernel performs a weighted calculation on a specific area in the input feature map, and the size of W determines the range of the window.

[0120] b(k): represents the bias term of the k-th channel in the output feature map. The bias term is used to adjust the result of the convolution operation, enabling the model to more flexibly learn different features.

[0121] By performing a weighted sum of the convolution kernel and the local area of the input feature map and adding the bias term, the value at a certain position of the output feature map is obtained. The convolution layer is applied to each stage of the dual-backbone network, mainly used to capture the local features of the image. By stacking convolution layers, the network can gradually extract deep complex target features from shallow edge and texture features.

[0122] In some embodiments, the detection head includes a first sub-detection head and a second sub-detection head.

[0123] The inputting the first backbone feature and the second backbone feature into the detection head to obtain a detection result includes:

[0124] Fusing and detecting the multiple first backbone features through the first sub-detection head to obtain multiple initial detection results. As Figure 2 shown, the first sub-detection head includes multiple convolutional blocks Conv and multiple RepCSPELAN4.

[0125] Fusing and detecting the second backbone feature and the initial detection results through the second sub-detection head to obtain the final detection result.

[0126] In some embodiments, as Figure 6 shown, the second sub-detection head includes multiple convolutional blocks, a pooling convolutional fusion module (PCRC), and a feature recombination module (FRM).

[0127] The fusing and detecting the second backbone feature and the initial detection results through the second sub-detection head to obtain the final detection result includes:

[0128] Convolving the second backbone feature and the initial detection results respectively through the multiple convolutional blocks to obtain target features. Here, the convolutional block includes an upsampling block, Maxpool, and multiple a*a Conv.

[0129] Performing max pooling and average pooling operations on the target features through the pooling convolutional fusion module to obtain key information. The specific structure of the pooling convolutional fusion module is as Figure 7 shown, and it includes Conv, Avgpool, DWConv, Maxpool, etc.

[0130] Performing upsampling and downsampling operations on the key information through the feature recombination module to obtain recombined features, and obtaining a prediction result based on the recombined features. Here, the feature recombination module includes a downsampling module and an upsampling module.

[0131] In addition, the second sub-detection head utilizes the distributed focal loss DFLoss algorithm during the training process, which is achieved by converting the input channels through a convolutional layer. Its purpose is to optimize the loss processing mechanism of the network, thereby improving the efficiency and accuracy of processing distributed data.

[0132]

[0133] where p t is the prediction rate of the model for the correct class, α t and γ are parameters for adjusting the prediction difficulty, is the distributed focal loss (DFLoss). This loss function is used during the training process to optimize the loss processing mechanism of the model and improve the effectiveness and accuracy of distributed data.

[0134] The feature recombination module aggregates and recombines features at different levels through operations such as conv, upsample, downsample, and softmax, thereby combining feature maps of different scales. Then, the recombined feature maps are used for further image processing.

[0135] F out = Conv(Upsample(Concat(F down ,F up )))

[0136] Among them, F down and F up respectively represent the feature images after downsampling and upsampling, and F out represents the final output feature map result obtained through operations such as convolution, upsampling, and downsampling.

[0137] The boundary dynamic perception detection method for the weak light environment of railways provided by the embodiments of the present invention performs aggregation analysis and reprocessing on the second backbone feature and the initial detection result through the second sub-detection head. The second sub-detection head can optimize the detection performance of dynamic and small-size targets, and improve the detection accuracy and robustness through the feature recombination module (FRM) and the pooling convolution fusion module (PCRC). It performs intrusion recognition on the monitored area for illegal intruders on the railway perimeter in the weak light environment, and focuses on detection in key throat areas, tunnels and other areas of the railway.

[0138] In the experiment, BDPNet was deployed on a railway scene dataset containing various low light conditions (such as at night, inside tunnels, etc.). According to the detection accuracy comparison results of the railway perimeter intrusion detection method in the Exdark weak light public dataset shown in Table 1, and the detection accuracy detection comparison of the railway perimeter intrusion detection method in the railway scene image dataset shown in Table 2, BDPNet provided by the present invention is significantly superior to the traditional YOLOv9 and other low light detection algorithms in terms of indicators such as accuracy and recall. In addition, BDPNet shows high stability and low false alarm rate when dealing with dynamic scenes (such as moving trains and operating personnel).

[0139] Table 1. Comparison of detection effects on the Exdark dataset

[0140]

[0141] Table 2. Comparison of detection effects on the railway weak light dataset

[0142]

[0143] As can be seen from Table 1 and Table 2 above, the average precision of BDPNet provided by the present invention in human detection in the railway low-light dataset reaches 0.866, which is significantly higher than that of MAET (0.693), IA-YOLO (0.72), DENet (0.734) and PE-YOLO (0.763). In addition, BDPNet also maintains a good balance in the floating-point operations per second (FLOPS) metric, with a value of 103.8. This indicates that BDPNet not only performs excellently in detection accuracy but also has competitiveness in computing efficiency, meeting the application conditions in the railway field. These results show that BDPNet is suitable for application in complex railway scenarios under low-light environments, can realize monitoring visualization and high-precision intrusion target detection models under low-light environments, and ensure the safe operation of railways. In addition, Figure 8 rows A - F in Figure 8 are the original images of multiple images to be detected under different low-light conditions and schematic diagrams of their related detection results, Figure 8 and columns (a) - (f) in Figure 8 are schematic diagrams before and after detection using different object detection algorithms. Figure 9 Specifically, columns (a) - (f) in Figure 9 are the original images to be detected, schematic diagrams of the detection results of BDPNet (i.e., the boundary dynamic perception detection network provided by the present invention), schematic diagrams of the detection results of YOLOv9, schematic diagrams of the detection results of IA-YOLO, schematic diagrams of the detection results of PE-YOLO, and schematic diagrams of the detection results of RD-YOLO. As

[0144] In summary, the boundary dynamic perception detection method for railway low-light environments provided by the present invention has the following advantages: (1) Enhanced low-light performance. Through the illumination correction network module (SCINet), the brightness and clarity of images can be effectively improved, enabling the algorithm to accurately perform target detection even in dim environments. (2) Improved high recognition ability in complex scenarios. Adopting a dual backbone network design enhances the understanding ability of complex railway scenarios. For example, it can still effectively identify targets in backgrounds with various interference objects such as trees and buildings. (3) High-precision intrusion target detection. Through the distributed focal loss (DFL) and feature recombination module, the algorithm optimizes feature extraction and fusion during detection, improving the detection accuracy, especially in identifying small or partially occluded targets. (4) The present invention can maintain stable performance under different lighting conditions and has good adaptability to weather changes and environmental dynamics.

[0145] Embodiment 2

[0146] Based on Embodiment 1, this Embodiment 2 provides a boundary dynamic perception detection device for railway low-light environments. This boundary dynamic perception detection device for railway low-light environments corresponds to the above-mentioned boundary dynamic perception detection method for railway low-light environments and specifically includes:

[0147] An image acquisition module, configured to acquire an image to be detected of the railway perimeter in a low-light environment and input it into the boundary dynamic perception detection network. The boundary dynamic perception detection network includes an illumination correction network module and a detection module;

[0148] An image correction module, configured to perform global brightening, shading correction, and edge detail enhancement on the image to be detected through the illumination correction network module to obtain a corrected image;

[0149] A detection module, configured to perform target detection on the corrected image through the detection module to obtain a detection result.

[0150] For specific details, refer to the description in the part of the boundary dynamic perception detection method for railway low-light environments, which will not be elaborated here.

[0151] Embodiment 3

[0152] Embodiment 3 of the present invention provides an electronic device, including a memory and a processor. The processor and the memory communicate with each other. The memory stores program instructions executable by the processor, and the processor calls the program instructions to execute the boundary dynamic perception detection method for railway low-light environments. This method includes the following process steps:

[0153] Acquire an image to be detected of the railway perimeter in a low-light environment and input it into the boundary dynamic perception detection network. The boundary dynamic perception detection network includes an illumination correction network module and a detection module;

[0154] The illumination correction network module performs global brightening, shading correction, and edge detail enhancement on the image to be detected, so as to obtain a corrected image;

[0155] The detection module performs target detection on the corrected image to obtain a detection result.

[0156] Embodiment 4

[0157] Embodiment 4 of the present invention provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements a method for boundary dynamic perception detection in a weak light environment of a railway. The method includes the following process steps:

[0158] Obtain an image to be detected of the railway perimeter in a weak light environment and input it into the boundary dynamic perception detection network. The boundary dynamic perception detection network includes an illumination correction network module and a detection module;

[0159] The illumination correction network module performs global brightening, shading correction, and edge detail enhancement on the image to be detected, so as to obtain a corrected image;

[0160] The detection module performs target detection on the corrected image to obtain a detection result.

[0161] Embodiment 5

[0162] Embodiment 5 of the present invention provides a computer program product, which includes a computer program. When the computer program is executed by a processor, it implements a method for boundary dynamic perception detection in a weak light environment of a railway. The method includes the following process steps:

[0163] Obtain an image to be detected of the railway perimeter in a weak light environment and input it into the boundary dynamic perception detection network. The boundary dynamic perception detection network includes an illumination correction network module and a detection module;

[0164] The illumination correction network module performs global brightening, shading correction, and edge detail enhancement on the image to be detected, so as to obtain a corrected image;

[0165] The detection module performs target detection on the corrected image to obtain a detection result.

[0166] Those of ordinary skill in the art can understand that: The drawings are only schematic diagrams of an embodiment, and the modules or processes in the drawings are not necessarily essential for implementing the present invention.

[0167] Each embodiment in this specification is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other, and the key point of each embodiment is to illustrate the differences from other embodiments. In particular, for method or system embodiments, since they are basically similar to method embodiments, they are described relatively simply, and reference can be made to the corresponding parts of the method embodiments for the relevant content. The method and system embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. A person of ordinary skill in the art can understand and implement it without creative efforts.

[0168] The above is only a preferred specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present invention should be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.

Claims

1. A boundary dynamic perception detection method in a railway weak light environment, characterized in that: include: Acquire an image to be detected of the railway perimeter in a weak light environment and input it into a boundary dynamic perception detection network, wherein the boundary dynamic perception detection network includes a lighting correction network module and a detection module; Performing global brightening, shadow correction and edge detail enhancement on the image to be detected through the illumination correction network module to obtain a corrected image; The detection module performs target detection on the corrected image to obtain a detection result.

2. The method according to claim 1, characterized in that The lighting correction network module includes a CBR module, an SC module, an I module and a CB module connected in sequence; The method of performing global brightening, dark correction and edge detail enhancement on the image to be detected by the illumination correction network module to obtain a corrected image includes: The CBR module and the SC module are used to extract features of the image to be detected, so as to obtain a feature extraction result g(x k ): g(x k )=SC(CBR 1,1 (x k )) Among them, x k represents the kth frame of the image to be detected, g(x k ) is the feature extraction result of the kth frame of the image to be detected. The CBR module consists of a convolutional layer, a batch normalization layer, and a ReLU activation function. The SC module is used for image feature extraction operations. t represents the index of a stage or time step, which is used to describe the multi-stage iteration in the feature processing process. y represents the image or feature input, and x t represents the input of element-by-element multiplication of the intermediate features obtained by step-by-step processing in stage t and y, z t Represents element-wise multiplication The result after processing, is the introduced parameterized operator, s t Indicates that The calculated transformed input, v t Indicated by y+s t Another conversion stage output composed of; The feature extraction result g(x k ) to perform illumination correction to obtain the illumination-corrected feature I(x k ): Among them, g t represents the characteristics at stage t, represents the parameterized mapping function, u t is the residual term at stage t, and the illumination-related information is obtained through the residual u t Accumulate to the current feature g t In order to gradually correct the illumination, g t+1 It represents the characteristics of the t+1 stage; The illumination correction feature I(x k ) is processed to obtain the corrected image F(x k ): F(x k )=CB 1,1 (I(g(x k )))。 3. The method according to claim 1, characterized in that The detection module includes a first trunk feature extraction module, a second trunk feature extraction module and a detection head; The performing target detection on the corrected image by the detection module to obtain a detection result includes: Performing multi-scale feature extraction on the corrected image by the first trunk feature extraction module to obtain a plurality of first trunk features; Performing multi-scale feature extraction on the corrected image and the first trunk features by the second trunk feature extraction module to obtain a plurality of second trunk features; The first trunk feature and the second trunk feature are input into the detection head to obtain a detection result.

4. The method according to claim 3, characterized in that The step of performing multi-scale feature extraction on the corrected image by using the first backbone feature extraction module to obtain a first backbone feature includes: The multi-scale feature extraction is performed on the corrected image using the first backbone feature extraction formula to obtain the first backbone feature: P3,P4,P5=L main (x i ) Where RepCSP is a multi-stage enhanced feature extraction module, which includes segmented convolution, residual connection, cross-stage connection, feature fusion and hierarchical operation. P3, P4, and P5 are the first backbone features of different scales respectively; L main (·) represents the main loss function, which is used to optimize the learning process of the first backbone feature extraction module; The step of performing multi-scale feature extraction on the corrected image and the first trunk feature by the second trunk feature extraction module to obtain the second trunk feature includes: Multi-scale feature extraction is performed on the corrected image and the first backbone feature using the second backbone feature extraction formula to obtain the second backbone feature: C3,C4,C5=L aux (P3,P4,P5,x i ) Wherein, Laux(·) is an auxiliary loss function, which is obtained by i After processing, the second backbone features C3, C4, C5 of different scales are obtained, and X includes the corrected image x i Together with the first backbone features P3, P4, and P5, CBL is a basic convolutional feature extraction module, which includes the combined operations of convolution, batch normalization, and activation functions, and DS represents the downsampling operation.

5. The method according to claim 3, characterized in that: The detection head comprises a first sub-detection head and a second sub-detection head; The step of inputting the first trunk feature and the second trunk feature into the detection head to obtain a detection result comprises: fusing and detecting the plurality of the first trunk features by the first sub-detection head to obtain a plurality of initial detection results; The second trunk feature and the initial detection result are fused and detected by the second sub-detection head to obtain a final detection result.

6. The method according to claim 5, characterized in that The second sub-detection head includes a plurality of convolution blocks, a pooling convolution fusion module and a feature recombination module; The fusing and detecting the second trunk feature and the initial detection result by the second sub-detection head to obtain a final detection result includes: Convolving the second trunk feature with the initial detection result respectively through the multiple convolution blocks to obtain target features; Performing maximum pooling and average pooling operations on the target features through the pooling convolution fusion module to obtain key information; The key information is upsampled and downsampled by the feature recombination module to obtain a recombination feature, and a prediction result is obtained based on the recombination feature.

7. A boundary dynamic perception detection device for a railway weak light environment, characterized in that: include: An image acquisition module is used to acquire an image to be detected of the railway perimeter in a weak light environment and input it into a boundary dynamic perception detection network, wherein the boundary dynamic perception detection network includes a lighting correction network module and a detection module; An image correction module, used for performing global brightening, shadow correction and edge detail enhancement on the image to be detected through the illumination correction network module to obtain a corrected image; The detection module is used to perform target detection on the corrected image through the detection module to obtain a detection result.

8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the method according to any one of claims 1 to 6 is implemented.

9. A computer-readable storage medium, characterized in that: It stores a computer program, which, when executed by a processor, implements the method according to any one of claims 1 to 6.

10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the method according to any one of claims 1 to 6 is implemented.

Citation Information

Cited By

  • Insect situation monitoring camera and monitoring method thereof

    CN121170461A