Foreign matter intrusion detection method for rail transit low-illumination operation scene

By using the image decomposition and enhancement network of multi-scale perception module and lighting attention mechanism module in the operating area of ​​rail transit trains, combined with the characteristic pyramid network and the second-stage target detection network of hollow space pyramid pooling, the problem of foreign object intrusion detection in low-illumination environments is solved, and accurate detection and classification of foreign objects of different sizes is achieved, which significantly reduces safety hazards.

CN120220028APending Publication Date: 2025-06-27NANJING UNIV OF SCI & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510309226.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-17
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The prior art is difficult to effectively detect foreign object invasion in the operating area of ​​rail transit trains in low-illumination environments, especially in night driving or extremely harsh weather conditions, and the detection performance of small foreign objects is poor, which poses a major safety hazard.

Method used

The multi-scale perception module is used to combine it with the image decomposition network, and by designing the lighting attention mechanism module and the image enhancement network, the image decomposition and enhancement image are improved to improve the model's ability to detect foreign objects in low-illumination environments. Combining the characteristic pyramid network and the hollow space pyramid pooling, a two-stage target detection network is built to achieve accurate detection and classification of foreign objects of different sizes.

Benefits of technology

In low-illumination environment, the accuracy and efficiency of foreign object intrusion detection are significantly improved, and invasive foreign objects of different sizes can be detected, and effectively classified, reducing safety hazards in rail transit operation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120220028A_ABST
    Figure CN120220028A_ABST
Patent Text Reader

Abstract

The invention discloses a rail transit low-illumination operation scene foreign matter intrusion detection method. The method comprises the following steps: firstly, deploying high-definition IPC cameras at two sides of a track to collect video streams, and carrying out frame reduction and frame extraction through an NVIDIA GPU codec to obtain high-definition images; secondly, constructing an image decomposition network containing a multi-scale perception module, extracting multi-dimensional features through parallel convolution and a pooling layer, and decomposing the image into an illumination component and a reflection component in combination with a channel-space attention mechanism; then designing a double-channel parallel illumination attention mechanism module, adaptively adjusting the feature weight through compression-excitation operation, and enhancing image details; and finally, providing a two-stage target detection network fusing a feature pyramid network and cavity space pyramid pooling, and realizing accurate positioning and classification of foreign matter invasion. According to the method, multi-scale feature learning and an illumination optimization mechanism are combined, the robustness and the real-time performance of foreign matter detection in a low-illumination environment are remarkably improved, an intrusion target can be automatically marked, and a classification result can be output.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of rail transit train operation safety detection, and particularly relates to a method for detecting foreign object intrusion in a low-illumination operation scenario of rail transit. Background Technique

[0002] With the continuous expansion of the rail transit operation scale, train operation safety accidents caused by foreign object intrusion occur from time to time. Once an accident occurs, it usually brings very serious consequences. Safety accidents caused by foreign objects invading the track have become major safety hazards for urban rail transit train operation.

[0003] In order to solve the problem of safety accidents caused by foreign object intrusion into the train operation area, at present, many scholars have studied rail transit foreign object intrusion detection. However, most of the studies are based on the normal daylight environment, and insufficient consideration is given to the environmental impacts such as night driving and extremely bad weather. Moreover, most of the detected foreign objects are relatively large foreign objects such as pedestrians, stones, and trees, and the performance of the algorithm for small foreign objects and driving scenarios with insufficient light drops significantly.

[0004] Therefore, if the scenario of train operation in a low-illumination environment is not considered during detection, it is impossible to ensure normal detection in low-illumination scenarios such as night driving, and small invading foreign objects cannot be detected in time, and there are still quite a few safety hazards. Summary of the Invention

[0005] The purpose of the present invention is to provide a method for detecting foreign object intrusion in a low-illumination operation scenario of rail transit to solve the above technical problems.

[0006] To achieve the purpose of the present invention, the present invention provides a method for detecting foreign object intrusion in a low-illumination operation scenario of rail transit, including the following steps:

[0007] Step 1: Install high-definition cameras on both sides of the road in the rail transit operation area to capture video streams, and obtain high-definition images by decimating frames of the video streams through an encoding and decoding device;

[0008] Step 2: Construct a multi-scale perception module and connect it to an image decomposition network to obtain a new image decomposition network. The high-definition image is decomposed into a light component and a reflection component by the new image decomposition network;

[0009] The construction of the multi-scale perception module is specifically as follows: The initial 1×1 convolutional layer and the initial 2×2 pooling layer are in parallel. The initial 1×1 convolutional layer extracts the features of the image, and the initial 2×2 pooling layer captures the wide-area information of the image. The two work in parallel to jointly process the multi-dimensional features of the image and complete the learning of multi-scale features. Three 3×3 convolutional layers are tightly connected after the initial 1×1 convolutional layer to extract image features at three different abstraction levels: large, medium, and small. After the initial 2×2 pooling layer, a 1×1 convolutional layer and a channel-spatial attention module are tightly connected. The 1×1 convolutional layer enhances the original features of the image through convolutional operations, and the channel-spatial attention module adaptively adjusts the feature response of the entire multi-scale module to complete the construction of the multi-scale perception module.

[0010] Step 3: By designing a light attention mechanism module, construct an image enhancement network, and input the light component and the reflection component into the image enhancement network to obtain an enhanced image.

[0011] Step 3.1: Design the compression channel of the light attention mechanism: Compress the three-dimensional space of the image reflection component and the image light component into a one-dimensional space through the compression channel of the light attention mechanism, so as to fully capture the feature information of the image reflection component and the image light component.

[0012] Step 3.2: Design the excitation channel of the light attention mechanism: Determine the feature weights according to the importance of the input image reflection component and the image light component. The feature weights guide the light attention mechanism module to use computing resources and image features to improve the performance of the model's visual tasks.

[0013] Step 3.3: Design a dual-channel parallel structure of the compression channel and the excitation channel of the light attention mechanism to construct a light attention mechanism module. Use the light attention mechanism module as the main network to construct an image enhancement network, and input the light component and the reflection component into the image enhancement network to obtain an enhanced image.

[0014] Step 4: Combine the feature pyramid network and the atrous spatial pyramid pooling to propose a two-stage object detection network. Connect the two-stage object detection network after the image enhancement network to obtain a foreign object intrusion detection model. Input the enhanced image into the foreign object intrusion detection model to complete the detection and classification of foreign object intrusion.

[0015] An electronic device includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the above-mentioned method for detecting foreign object intrusion in a low-illumination operation scenario of rail transit.

[0016] A non-transitory computer-readable storage medium stores computer instructions for causing a computer to execute the above-mentioned foreign object intrusion detection method for low-illumination operation scenarios of rail transit.

[0017] Compared with the prior art, the remarkable progress of the present invention lies in: (1) The present invention constructs a multi-scale perception module, and connects the multi-scale perception module to the image decomposition network to obtain a new image decomposition network. The new image decomposition network can expand the feature perception range of the input high-definition image, so as to better capture the texture details, brightness, contrast, edge information features and structures in the input high-definition image, and improve the model's perception ability of different scales of large, medium and small information; (2) The present invention designs a light attention mechanism module to compress channels and a light attention mechanism module to stimulate channels, and realizes the complete light attention mechanism module in parallel through two channels. Using the light attention mechanism module as the main network to construct an image enhancement network, improving the brightness enhancement effect and the retention effect of image texture details of the image enhancement network, and inputting the light component and the reflection component into the image enhancement network to obtain a clear and high-quality enhanced image; (3) The present invention combines a feature pyramid network and a dilated spatial pyramid pooling to construct a two-stage object detection network, and tightly connects the two-stage object detection network behind the image enhancement network to obtain a foreign object intrusion detection model, improving the detection accuracy of the foreign object intrusion detection model for different sizes of large, medium and small intrusion foreign objects, and can detect foreign objects invading the train operation area in the normal operation scenario of rail transit under extreme illumination scenarios and classify them.

[0018] To more clearly illustrate the functional characteristics and structural parameters of the present invention, the following further explains in conjunction with the drawings and specific embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] The drawings described herein are used to provide a further understanding of the present invention, and constitute a part of this application. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings:

[0020] Figure 1 is the method flow chart of the present invention;

[0021] Figure 2 is the structural diagram of the multi-scale perception module constructed in Python of the present invention;

[0022] Figure 3 is the structural diagram of the light attention mechanism module constructed in Python of the present invention;

[0023] Figure 4 is the two-stage object detection network constructed in Python of the present invention;

[0024] Figure 5 They are images with different illuminations in three different scenarios selected when verifying the invention effect of the present invention;

[0025] Figure 6 They are the enhancement effects of the present invention on the selected images with different illuminations in three different scenarios;

[0026] Figure 7 They are the detected intrusion foreign objects when the present invention inputs images of intrusion foreign objects of large, medium and small different sizes into the foreign object intrusion detection network. Specific embodiments

[0027] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments; based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0028] A method for detecting intrusion foreign objects in a low-illumination operation scenario of rail transit according to the present invention, in combination with Figure 1 , includes the following steps:

[0029] Step 1: Install high-definition IPC cameras on both sides of the road in the rail transit operation area, and the high-definition IPC cameras capture video streams in real time; use the NVIDIA GPU video codec to process the video streams: perform frame reduction processing on the video streams, and then perform frame extraction processing on the video streams after the frame reduction processing is completed to obtain high-definition images.

[0030] Step 2: Construct a multi-scale perception module and connect it to the image decomposition network to obtain a new image decomposition network. The high-definition image is decomposed into a light component and a reflection component through the new image decomposition network;

[0031] In combination with Figure 2 , the construction of the multi-scale perception module is specifically as follows:

[0032] Parallelize the initial 1×1 convolutional layer and the initial 2×2 pooling layer. The initial 1×1 convolutional layer extracts the features of the image, and the initial 2×2 pooling layer captures the wide-area information of the image. The two are parallel to jointly process the multi-dimensional features of the image and complete the learning of multi-scale features; three 3×3 convolutional layers are closely connected after the initial 1×1 convolutional layer to extract image features at three different abstraction levels of large, medium and small; a 1×1 convolutional layer and a channel-spatial attention module are closely connected after the initial 2×2 pooling layer. The 1×1 convolutional layer enhances the original features of the image through convolutional operations, and the channel-spatial attention module adaptively adjusts the feature response of the entire multi-scale module to complete the construction of the multi-scale perception module;

[0033] The decomposition formula of the new image decomposition network is as follows:

[0034] ;

[0035] Wherein, represents a free variable; represents the high-definition image; represents the reflection component of the image; represents the illumination component of the image; represents the Gaussian function; is the scale parameter of the high-definition image.

[0036] Step 3: By designing an illumination attention mechanism module, construct an image enhancement network, input the illumination component and the reflection component into the image enhancement network, and obtain an enhanced image;

[0037] Step 3.1: Design the compression channel of the illumination attention mechanism: Compress the three-dimensional space of the image reflection component and the image illumination component into a one-dimensional space through the compression channel of the illumination attention mechanism, so as to fully capture the feature information of the image reflection component and the image illumination component;

[0038] The compression process of the compression channel of the illumination attention mechanism is the formula:

[0039] ;

[0040] Wherein, represents a free variable; represents a compression function, represents the feature description operator obtained by compressing through the compression channel of the illumination attention mechanism; represents a three-dimensional matrix in the th two-dimensional matrix; the subscript represents the number of image features; , respectively represent the height and width of the feature map.

[0041] Step 3.2: Design the excitation channel of the illumination attention mechanism: Determine the feature weights according to the importance of the input image reflection component and the image illumination component, and the feature weights guide the illumination attention mechanism module to use computing resources and image features to improve the performance of the model's visual tasks;

[0042] The excitation formula of the excitation channel of the illumination attention mechanism is:

[0043] ;

[0044] Among them, represents a free variable; represents an activation function, represents the finally obtained weight parameter; represents the sigmoid activation function; 、 are respectively the artificially set parameters in the dimension-increasing layer and the dimension-decreasing layer.

[0045] Step 3.3, combine Figure 3 , design a dual-channel parallel construction of the compression channel and the excitation channel of the light attention mechanism to construct a light attention mechanism module, use the light attention mechanism module as the main network to construct an image enhancement network, input the light component and the reflection component into the image enhancement network, and obtain a clear and high-quality enhanced image.

[0046] Step 4, combine Figure 4 , combine the feature pyramid network and the atrous spatial pyramid pooling, including a 3×3 convolutional layer, a ReLu layer, an MSSM layer, a Sigmoid layer and a luminance attention module, propose a two-stage object detection network, connect the two-stage object detection network after the image enhancement network to obtain a foreign object intrusion detection model, after inputting the enhanced image into the foreign object intrusion detection model, the foreign object intrusion detection model automatically identifies the intrusion foreign object, marks the intrusion foreign object with a yellow detection box and automatically completes the classification.

[0047] Embodiment

[0048] Select three datasets of LOL, MIT-Adobe FiveK, and MOT20, randomly extract 600 images and 2 video clips to form the dataset for this experiment, and randomly divide them into a training set, a validation set and a test set according to the ratio of 7:1:2. The number of samples input to the model at the same time when training the model iteratively is set to 10, the size of the small images or regions cropped from the original images is set to 64×64, and the Adam optimizer is used for network optimization. The configuration of the computer used in the experiment is: AMD R7 7800X3D CPU and GeForce RTX 4090 D GPU, and the computer program described by the model is implemented using the Pytorch package.

[0049] Taking a certain rail transit driving section of a certain subway in a certain city as an example:

[0050] Install high-definition IPC cameras on both sides of the roads in the rail transit operation area. The high-definition IPC cameras capture video streams in real time. Use the NVIDIA GPU video codec to process the video streams: perform frame rate reduction on the video, and then perform frame extraction on the video after frame rate reduction to obtain high-definition images. For easy comparison, three images with different illuminations in different scenarios are set as Figure 5 , Figure 5 The three images in the figure are scenes of different train operation intervals under different light intensities.

[0051] Design a multi-scale perception module to enhance the perception range of the model. Connect the multi-scale module to the mainstream image decomposition network D-NET to construct a new image decomposition network. Input the high-definition image into the image decomposition network, and decompose the high-definition image with high fidelity to obtain the reflection component and the illumination component of the image. Design an illumination attention mechanism module, including designing the compression channel and the excitation channel of the illumination attention mechanism module. The two channels are parallel to improve the brightness enhancement effect of the model and the retention effect of image texture details. Use the illumination attention mechanism module as the main network to construct an image enhancement network. Input the illumination component and the reflection component into the image enhancement network to obtain a clear and high-quality enhanced image. The image enhancement effect is as Figure 6 , it can be seen that this method can effectively enhance the brightness and contrast of low images while perfectly preserving the detailed information of image texture contours, restoring image colors and improving the naturalness of images.

[0052] Combine the feature pyramid network and the atrous spatial pyramid pooling of the neural network to construct a two-stage object detection network. Connect the two-stage object detection network tightly after the image enhancement network to obtain a foreign object intrusion detection model. Input the enhanced image into the foreign object intrusion detection model, and the foreign object intrusion detection model automatically identifies the intruding foreign objects, marks the intruding foreign objects with yellow detection frames and automatically completes classification. Input images of intruding foreign objects of different sizes (large, medium, and small) into the foreign object intrusion detection network to detect intruding foreign objects as Figure 7 , it can be seen that this method can accurately identify intruding foreign objects of three different sizes (large, medium, and small), and the yellow detection frames complete the framing and classification.

[0053] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the embodiments provided in the present application can include non-volatile and / or volatile memories. Non-volatile memories can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memories can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.

[0054] It should be noted that, in this document, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprise", "include" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device.

[0055] Although the embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A method for detecting foreign body intrusion in low-light operation scenarios of rail transit, characterized in that: The following steps are involved: Step 1: Install high-definition cameras on both sides of the road in the rail transit operation area to capture video streams, and extract frames after reducing the frames of the video streams through a codec to obtain high-definition images; Step 2: construct a multi-scale perception module and connect it to an image decomposition network to obtain a new image decomposition network, wherein the high-definition image is decomposed into an illumination component and a reflection component through the new image decomposition network; Step 3: construct an image enhancement network by designing an illumination attention mechanism module, input the illumination component and the reflection component into the image enhancement network, and obtain an enhanced image; Step 4: Combining the feature pyramid network and the void space pyramid pooling, a two-stage target detection network is proposed. The two-stage target detection network is connected to the image enhancement network to obtain a foreign body intrusion detection model. The enhanced image is input into the foreign body intrusion detection model to complete foreign body intrusion detection and classification.

2. According to claim 1, a method for detecting foreign body intrusion in low-light operation scenarios of rail transit is characterized in that: The step 1 is specifically as follows: installing high-definition IPC cameras on both sides of the road in the rail transit operation area, and the high-definition IPC cameras capture video streams in real time; using the NVIDIA GPU video codec to process the video stream: performing frame reduction processing on the video stream, and then performing frame extraction processing on the video stream that has completed the frame reduction processing to obtain a high-definition image.

3. According to claim 1, a method for detecting foreign body intrusion in low-light operation scenarios of rail transit is characterized in that: The construction of the multi-scale perception module is specifically as follows: The initial 1×1 convolution layer and the initial 2×2 pooling layer are performed in parallel. The initial 1×1 convolution layer extracts the features of the image, and the initial 2×2 pooling layer captures the wide-area information of the image. The two layers jointly process the multi-dimensional features of the image in parallel to complete the learning of multi-scale features. The initial 1×1 convolutional layer is followed by three 3×3 convolutional layers that are tightly connected to extract image features at three different abstraction levels: large, medium, and small. The initial 2×2 pooling layer is tightly connected to a 1×1 convolution layer and a channel-space attention module. The 1×1 convolution layer enhances the original features of the image through convolution operations, and the channel-space attention module adaptively adjusts the feature response of the entire multi-scale module to complete the construction of the multi-scale perception module.

4. According to claim 1, a method for detecting foreign body intrusion in low-light operation scenarios of rail transit is characterized in that: The new image decomposition network decomposition formula of step 2 is as follows: ; in, represents a free variable; representing said high definition image; represents a reflection component of the image; represents an illumination component of the image; represents the Gaussian function; is the scale parameter of the high-definition image.

5. According to claim 1, a method for detecting foreign body intrusion in low-light operation scenarios of rail transit, characterized in that: The step 3 specifically includes the following steps: Step 3.1, designing a compression channel of the illumination attention mechanism: compressing the three-dimensional space of the image reflection component and the image illumination component into a one-dimensional space through the compression channel of the illumination attention mechanism, so as to fully capture the feature information of the image reflection component and the image illumination component; Step 3.2, designing an excitation channel of the illumination attention mechanism: determining feature weights according to the importance of the reflection component and the illumination component of the input image; Step 3.3, design the compression channel and excitation channel of the illumination attention mechanism to construct an illumination attention mechanism module in parallel, build an image enhancement network with the illumination attention mechanism module as the main network, input the illumination component and the reflection component into the image enhancement network to obtain an enhanced image.

6. A method for detecting foreign body intrusion in low-light operation scenarios of rail transit according to claim 5, characterized in that: The compression process of the illumination attention mechanism compression channel is as follows: ; in, represents a free variable; represents the compression function; represents the feature description operator obtained by compression of the compression channel of the illumination attention mechanism; Represents a three-dimensional matrix Middle A two-dimensional matrix; subscript Represents the number of image features; , Represent the height and width of the feature map respectively.

7. According to claim 5, a method for detecting foreign body intrusion in low-light operation scenarios of rail transit is characterized in that: The excitation formula of the excitation channel of the illumination attention mechanism is: ; in, represents a free variable; represents the activation function, Represents the final weight parameter; Represents the sigmoid activation function; , The parameters are set manually in the dimension-upgrading layer and the dimension-downgrading layer respectively.

8. The method for detecting foreign body intrusion in low-light operation scenarios of rail transit according to claim 1 is characterized in that: After the enhanced image is input into the foreign body intrusion detection model in step 4, the foreign body intrusion detection model automatically identifies the intruding foreign body, marks the intruding foreign body with a yellow detection frame and automatically completes the classification.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the method according to any one of claims 1 to 8 is implemented.

10. A non-transitory computer-readable storage medium, characterized in that: The non-transitory computer-readable storage medium stores computer instructions for causing the computer to execute the method of any one of claims 1 to 8.