Multi-target detection method

By repairing railway track features through multi-directional filtering and generative residual compensation modules, and combining 3D track domain spatial attention, the robustness and accuracy problems of track detection in existing technologies are solved, and efficient multi-target detection is achieved.

CN122066967AActive Publication Date: 2026-05-19QINGDAO UNIV OF TECH +3
View PDF 6 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
QINGDAO UNIV OF TECH
Filing Date
2026-04-17
Publication Date
2026-05-19

AI Technical Summary

Technical Problem

Existing railway track inspection technologies have poor robustness in complex environments, are difficult to adapt to the anisotropic characteristics of tracks, and lack global topological continuity and three-dimensional physical space constraints, resulting in insufficient inspection accuracy.

Method used

A multi-directional filtering module is used to decouple the track geometric modal features, a generative residual compensation module is used to repair feature breaks, and a 3D track domain spatial attention mechanism is introduced to construct a track feasible domain mask, thereby realizing track geometry reconstruction and foreign object detection.

Benefits of technology

It significantly improves detection accuracy under complex working conditions, reduces false alarm rate, and ensures the topological connectivity and physical spatial accuracy of detection results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122066967A_ABST
    Figure CN122066967A_ABST
Patent Text Reader

Abstract

The invention relates to a multi-target detection method, which comprises the following steps of: based on multidirectional anisotropic filtering, effectively decoupling track main body texture and environmental background noise, and remarkably improving the feature expression capability and the signal-to-noise ratio in a low-contrast and complex texture environment; the fracture area of the track caused under the severe working conditions of partial shielding, light reflection and the like is positioned and repaired, and the connectivity and integrity of the output track on the topological level are guaranteed; by constructing a feasible region mask surrounding an orbit, spatial semantic ambiguity caused by a monocular vision perspective effect is eliminated, it is ensured that foreign matter alarm is strictly limited in a physical dangerous region, and the false alarm rate is greatly reduced. Through simultaneous extraction of the railway track topological structure and troubleshooting of the foreign matter risk in the track area, the detection precision under the complex working condition is significantly improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of railway safety monitoring technology, and in particular to a multi-target detection method. Background Technology

[0002] As the physical foundation for train operation, the integrity of railway tracks and the clearance environment of the track area are core elements for ensuring train safety. With the continuous expansion of the high-speed railway network, traditional manual inspection can no longer meet the needs of high-frequency, large-scale inspection. In recent years, non-contact inspection technology based on machine vision has become a research hotspot due to its low cost and high efficiency. However, existing inspection methods still face severe challenges in complex field environments.

[0003] Current mainstream methods are mostly based on convolutional neural networks, but their general designs are difficult to adapt to the specific characteristics of railway scenarios. First, the isotropic square convolutional kernels are fundamentally mismatched with the anisotropic geometric features of the track, which extends continuously over long distances in the longitudinal direction and is modulated laterally by sleepers. This makes the main track signal easily homogenized by background noise such as gravel and ballast. Second, existing models rely on local receptive fields for pixel-level classification, lacking global reasoning capabilities for the long-range topological continuity of the track. When encountering local occlusion or reflective interference, they are prone to outputting broken track segments, failing to meet the stringent continuity requirements of train operation control. Furthermore, multi-task joint detection suffers from feature conflicts and inaccurate localization, further limiting detection accuracy. Summary of the Invention

[0004] To address the aforementioned technical problems, this disclosure provides a multi-target detection method.

[0005] In a first aspect, embodiments of this disclosure provide a multi-target detection method, including: The railway scene image is acquired, and features are extracted from the railway scene image through a pre-constructed neural network to obtain a basic feature map. The railway scene image includes the track. By pre-constructing a multi-directional filtering module, the target feature information used to characterize the orbital geometric mode in the basic feature map is decoupled in the frequency domain, and the target feature information is enhanced to obtain an enhanced feature map; The feature fracture region of the track is located in the enhanced feature map, and the geometric extension information of the track is inferred by a pre-built generative residual compensation module, so as to repair the feature fracture region in the feature space and obtain the repaired feature map. Based on the camera parameters when the target camera captures images of the railway scene, the repaired feature map is mapped onto a three-dimensional physical space to construct a track feasible region mask that surrounds the track. Multi-target detection is performed based on the track feasible region mask to generate detection results. The multi-target detection includes track detection and foreign object detection. The detection results include the geometric topology information of the track and the detection information of foreign objects that pose an operational risk to the track.

[0006] Secondly, embodiments of this disclosure provide a multi-target detection device, including: The feature extraction unit is used to acquire railway scene images and extract features from the railway scene images through a pre-built neural network to obtain a basic feature map, wherein the railway scene images include tracks; The feature enhancement unit is used to decouple the target feature information representing the orbital geometric mode in the basic feature map in the frequency domain by pre-constructing a multi-directional filtering module, and enhance the target feature information to obtain an enhanced feature map; The feature repair unit is used to locate the feature fracture region of the track in the enhanced feature map, and infer the geometric extension information of the track through a pre-built generative residual compensation module, so as to repair the feature fracture region in the feature space and obtain the repaired feature map. The feasible region construction unit is used to map the repaired feature map to a three-dimensional physical space based on the camera parameters when the target camera captures the railway scene image, and to construct a track feasible region mask that surrounds the track. The multi-target detection unit is used to perform multi-target detection based on the track feasible region mask and generate detection results. The multi-target detection includes track detection and foreign object detection. The detection results include the geometric topology information of the track and the detection information of foreign objects that pose an operational risk to the track.

[0007] Thirdly, embodiments of this disclosure provide an electronic device, including: Memory; Processor; and Computer programs; The computer program is stored in memory and configured to be executed by a processor to implement the first aspect of the method described above.

[0008] Fourthly, embodiments of this disclosure provide a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the method described in the first aspect above.

[0009] This disclosure provides a multi-target detection method, comprising: acquiring a railway scene image and extracting features from the railway scene image using a pre-constructed neural network to obtain a basic feature map, wherein the railway scene image includes the track; decoupling target feature information representing the track geometric mode in the basic feature map in the frequency domain using a pre-constructed multi-directional filtering module, and enhancing the target feature information to obtain an enhanced feature map; locating the feature fracture region of the track in the enhanced feature map, and inferring the geometric extension information of the track using a pre-constructed generative residual compensation module to repair the feature fracture region in the feature space to obtain a repaired feature map; mapping the repaired feature map to a three-dimensional physical space according to the camera parameters when the target camera captures the railway scene image, and constructing a track feasible region mask surrounding the track; performing multi-target detection based on the track feasible region mask to generate detection results, wherein the multi-target detection includes track detection and foreign object detection, and the detection results include the geometric topology information of the track and the detection information of foreign objects that pose an operational risk to the track. This application, based on multi-directional anisotropic filtering, effectively decouples the track texture from environmental background noise, significantly improving feature representation and signal-to-noise ratio in low-contrast and complex texture environments. It locates and repairs broken areas caused by partial occlusion, reflection, and other adverse conditions, ensuring the connectivity and integrity of the output trajectory at the topological level. By constructing a feasible domain mask surrounding the track, it eliminates spatial semantic ambiguity caused by monocular vision perspective effects, ensuring that foreign object alarms are strictly limited to physical danger zones, greatly reducing false alarm rates. By simultaneously extracting the railway track topology and investigating foreign object risks in the track area, it significantly improves detection accuracy under complex conditions. Attached Figure Description

[0010] The accompanying drawings, which are incorporated in and form a part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure.

[0011] To more clearly illustrate the technical solutions in the embodiments of this disclosure or the prior art, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0012] Figure 1 A flowchart illustrating a multi-target detection method provided in an embodiment of this disclosure; Figure 2 A schematic diagram of a 2D track mask region provided in an embodiment of this disclosure; Figure 3 A schematic diagram of the establishment of a 3D danger zone for a track provided in an embodiment of this disclosure; Figure 4 This is a schematic diagram of the detection result of a foreign object on a track provided in an embodiment of the present disclosure; Figure 5 A schematic diagram of detection accuracy provided in an embodiment of this disclosure; Figure 6 A flowchart illustrating another multi-target detection method provided in this embodiment of the disclosure; Figure 7 This is a schematic diagram of the structure of a multi-target detection device provided in an embodiment of this disclosure; Figure 8 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this disclosure. Detailed Implementation

[0013] To better understand the above-mentioned objectives, features, and advantages of this disclosure, the solutions disclosed herein will be further described below. It should be noted that, unless otherwise specified, the embodiments and features described herein can be combined with each other.

[0014] Numerous specific details are set forth in the following description in order to provide a full understanding of this disclosure, but this disclosure may also be implemented in other ways different from those described herein; obviously, the embodiments in the specification are only some, and not all, of the embodiments of this disclosure.

[0015] Specifically, railway tracks, as the physical foundation for train operation, require a high degree of geometric integrity and a clear environment in the track area to ensure safe train operation. With the expansion of high-speed railway networks, traditional manual inspections can no longer meet the demands of high-frequency, large-scale inspections. In recent years, non-contact inspection technologies based on machine vision have become a research hotspot due to their low cost and high efficiency. However, existing inspection technologies still face severe challenges in complex field environments. Early traditional vision methods mainly relied on edge detection operators or Hough transforms to extract straight-line features. These methods are extremely sensitive to environmental noise, easily fail against rust, reflective, or gravel backgrounds, and cannot adapt to curved tracks. With the development of deep learning, semantic segmentation and object detection algorithms based on convolutional neural networks have been widely applied. Although their performance is superior to traditional algorithms, existing technologies still have the following significant drawbacks when dealing with the special "long-distance continuous geometric manifold" and "unstructured open environment" of railways: First, there is a mismatch between the general convolutional kernel and the anisotropic features of the track. Existing backbone networks generally use isotropic square convolutional kernels. This design assumes that features have equal weights in all directions. However, railway tracks exhibit strong geometric anisotropy: rails show long-distance continuity in the longitudinal direction, but exhibit periodicity in the transverse direction modulated by sleepers. Isotropic convolutional kernels cannot distinguish between these two orthogonal geometric patterns during sampling, causing the main rail signal to be easily homogenized by the gravel texture in the background, making it difficult for the network to capture key longitudinal extension features. Secondly, existing methods lack robust repair capabilities for local topological breaks. In actual operation, tracks often face local occlusion or reflective interference. Existing detection models are mostly based on pixel-level classification using local receptive fields, lacking the ability to reason about global geometric topology. Once local features are lost due to interference, the network often outputs broken track segments. For train control systems that require strict continuity, such topological breaks are unacceptable. At the same time, there are serious semantic ambiguities between two-dimensional visual features and three-dimensional physical space. Existing foreign object detection is mainly performed in the two-dimensional image plane. Due to perspective projection effects, distant roadside objects may overlap with track pixels in the image, leading to a very high false detection rate. Based solely on two-dimensional visual features, the model struggles to distinguish between "risk obstacles" that are actually located on the track and "safe background objects" that are only visually close, leading to an extremely high false detection rate. Existing methods lack the ability to map two-dimensional features back to three-dimensional physical space, failing to utilize physical priors such as constant track gauge and smooth roadbed to constrain the detection range. Moreover, existing methods often suffer from multi-task feature conflicts and inaccurate localization. Existing joint detection frameworks often share detection heads, ignoring the contradiction between "the need for translation invariance in foreign object classification" and "the need for translation sensitivity in position regression," resulting in limited detection accuracy. Furthermore, discrete prediction leads to unacceptable geometric non-smoothness.In track regression tasks, commonly used anchor-frame or pixel classification methods output discrete results, often producing jagged, oscillating trajectories or unnatural curves. These fail to meet the stringent curve smoothness requirements of railway engineering design and are therefore unsuitable for direct use in high-precision navigation. In summary, a novel detection method is urgently needed that integrates directional geometry sensing, topological continuity restoration, and three-dimensional physical space constraints to overcome these technical bottlenecks.

[0016] To address the problems of poor robustness and high false detection rate of foreign objects in existing track detection technologies under complex backgrounds, occlusion, and strong light conditions, this disclosure provides a multi-target detection method. First, a four-directional convolution enhancement module based on directional geometric constraints is proposed to extract anisotropic features of the track. Second, a local structure recovery module repairs topological breaks caused by occlusion. Then, a 3D track domain spatial attention mechanism based on camera parameters is introduced to generate a physically feasible domain mask to suppress background interference. Finally, an asymmetric interaction mechanism injects track geometric priors into the foreign object detection branch, and a topological constraint polyline decoder is used to achieve continuous and smooth track prediction. This method simultaneously performs geometric perception and physical constraint-driven multi-target detection for railway track topology extraction and foreign object risk screening in the track area, generating detection results. This disclosure achieves end-to-end track geometry reconstruction and risk target screening, significantly improving detection accuracy under complex conditions. Detailed descriptions are provided through one or more of the following embodiments.

[0017] The multi-target detection method provided in this disclosure is applicable to multi-target detection scenarios in railway environments. This method can be executed by a multi-target detection device, which can be implemented in software and / or hardware and integrated into an electronic device. The electronic device can include, but is not limited to, mobile terminals such as smartphones, laptops, digital radio receivers, personal digital assistants (PDAs), tablet computers (Tablet PCs), PMPs (portable multimedia players), in-vehicle terminals (e.g., in-vehicle navigation terminals), wearable devices, and fixed terminals such as digital televisions, desktop computers, and smart home devices.

[0018] Figure 1 This is a flowchart illustrating a multi-target detection method provided in an embodiment of the present disclosure, specifically including as follows: Figure 1 The following steps are shown: S101. Obtain railway scene images and extract features from the railway scene images using a pre-built neural network to obtain basic feature maps.

[0019] The railway scene images include the tracks.

[0020] Understandably, in the early stages of railway visual perception, facing extremely complex open environments, targets in images exhibit a huge scale span, from tiny roadside signal lights at a distance to massive track textures at a close distance. This drastic scale variation makes it difficult for single-level feature extraction to simultaneously address global semantic understanding and local detail localization. Therefore, this disclosure constructs a composite feature encoding architecture with CSPDarknet as its core backbone, deeply integrating feature pyramid networks and path aggregation networks, denoted as the feature extraction module (e.g., a feature extraction module constructed using a deep convolutional neural network). Its input is a high-resolution image, and its output is multi-scale features. Specifically, a high-resolution railway scene image is acquired, where the railway scene information includes the track. Subsequently, multi-scale feature extraction is performed on the railway scene image using a deep convolutional neural network to obtain the basic feature map of the railway scene image.

[0021] S102. By pre-constructing a multi-directional filtering module, the target feature information used to characterize the orbital geometric mode in the basic feature map is decoupled in the frequency domain, and the target feature information is enhanced to obtain an enhanced feature map.

[0022] Understandably, based on the above S101, railway tracks often exhibit extremely long and thin continuous geometric manifolds with significant anisotropy. Existing standard convolutional neural networks generally use 3×3 isotropic convolutional kernels. However, this type of standard neural network often causes the longitudinal extension signal of the rail to be submerged by the disordered texture of the background gravel. Therefore, this disclosure designs a spatial prior feature enhancement module, also known as a multi-anisotropic filtering module. Specifically, the basic feature map is input into the spatial prior feature enhancement module, which includes multiple sets of anisotropic filters, each set of filters including parallel anisotropic convolutional layers. Subsequently, the anisotropic filters are used to decouple the target feature information in the basic feature map used to characterize the geometric modes of track-related elements such as rails, sleepers, and curves in the frequency domain, that is, to determine the relevant features of the track. The relevant features are then recombined through a gating mechanism that includes adaptive weight calculation to enhance the track's main direction response, resulting in an enhanced feature map.

[0023] The anisotropic filtering module includes multiple sets of filters, each set of filters including parallel anisotropic convolutional layers.

[0024] Understandably, the spatial prior feature enhancement module based on anisotropic filter banks first enhances the input features... (The basic feature map) is sent to a set of parallel anisotropic convolutional layers. Each filter bank A group of convolutional kernels containing four specific directions Among them, the longitudinal core Physical dimensions designed for A long, narrow rectangular shape, in which Specifically designed to respond to the low-frequency continuity of rail and roadbed edges along the vertical axis of the image, this convolution kernel acts as a bandpass filter for high-frequency components in the vertical direction. Furthermore, by aggregating long-distance pixels vertically, it effectively suppresses high-frequency noise in the horizontal direction. (Horizontal kernel) Physical dimensions designed for A flat rectangular shape is specifically designed to extract the lateral periodic texture of sleepers. Although sleepers are not directly detected targets, their periodic signals provide strong contextual clues confirming the presence of the track area. (Diagonal kernel) and: Corresponding to and direction In the curve scenario, a tangent vector sensitive matching mechanism is introduced. That is, the network automatically activates the corresponding diagonal kernel based on the main direction of the local gradient, thereby realizing the matching of the curve tangent. High-response matching in direction. The kernel design of the anisotropic convolutional layers included in each filter bank is not limited and can be adjusted according to user needs.

[0025] Optionally, by pre-constructing a multi-directional filtering module, the target feature information used to characterize the orbital geometric modes in the basic feature map is decoupled in the frequency domain, and the target feature information is enhanced to obtain an enhanced feature map, including: Anisotropic convolutional layers are used to convolve the base feature map to obtain response maps in each orientation. The convolution operation involves weighted summation of the base feature map's local regions with the convolutional kernel at each spatial location. The convolutional layers then map pixels in each orientation response map to activation values ​​in each orientation, and normalize these activation values ​​to obtain orientation weights. These weights represent the probability that the current pixel position belongs to a geometric orientation structure. Finally, the orientation response maps are weighted and summed to enhance the attributes of each pixel in the orientation response maps, resulting in enhanced feature maps. The pixel attributes include straight tracks, curved tracks, and the background environment.

[0026] Understandable, space curve This represents the convolutional response of kernel K at (x,y) to the input base feature map F. This response is obtained through the convolution operation. This projects the mixed visual features onto independent geometric subspaces, yielding response maps for each orientation. The convolution operation involves a weighted summation of the convolution kernel K with the local region of the input base feature map F at each spatial location (x, y). Furthermore, while the directional filter bank provides rich geometric cues, the contributions of different directional components to different topological regions can exhibit significant spatial heterogeneity. To address this, a lightweight self-gating mechanism is introduced into the spatial prior feature enhancement module. Specifically, it utilizes two layers... Convolution maps directional response features to directional activation values ​​(denoted as ). ), and use the Softmax function to calculate the normalized confidence weights based on the directional activation values ​​(denoted as ). ), as shown in formula (1).

[0027] Formula (1) In the formula, Represents an exponential function. For the first The original activation values ​​of each directional channel, denominator term The sum of the exponential values ​​of the activation values ​​in all directions. Calculated based on formula (1). satisfy Its physical meaning is that the current pixel position belongs to the first... The probability confidence of a geometric structure is calculated. Subsequently, gating is performed using directional weights to select features that better characterize the orbital geometric modes in the longitudinal or lateral directions.

[0028] Understandably, the response maps of each direction are weighted and summed using the weights of each direction, and the original features are injected through the residual structure to enhance the attributes of each pixel in the response maps of each direction, resulting in an enhanced feature map, as shown in formula (2). The attributes of the pixel include straight track, curved track, and background environment. The subsequent network can automatically determine whether the current pixel belongs to a straight track (enhanced), curved track (enhanced), or cluttered background (uniform distribution or suppression) based on the enhanced local texture features.

[0029] Formula (2) In the formula, Represents the enhanced feature map, weights It represents the probability that the current pixel position belongs to the corresponding geometric orientation structure.

[0030] Understandably, this utilizes weights in each direction. Convolutional responses in all directions After performing a weighted summation, it can also be done by... Convolutional dimensionality reduction, followed by injection of basic features in the form of residuals. To obtain the final output As shown in formula (3).

[0031] Formula (3) Understandably, residual connections ensure lossless gradient propagation, while weighted terms act as directional guiding fields. During backpropagation, for geometrically consistent continuous structures, gradients are amplified along the tangent direction; for random noise, gradients decay exponentially, thus achieving rapid convergence during training.

[0032] S103. Locate the characteristic fracture region of the track in the enhanced feature map, and infer the geometric extension information of the track through a pre-built generative residual compensation module, so as to repair the characteristic fracture region in the feature space and obtain the repaired feature map.

[0033] Understandably, based on the above S102, in actual working conditions, rails often exhibit visual feature breaks due to reflections or partial occlusions. Therefore, this disclosure designs a local structure restoration module, which aims to repair topological damage through a manifold completion mechanism. Specifically, the decoupled features (enhanced feature map) are input into the local structure restoration module. The broken areas of the rail are located using a pixel-level saliency map, and the local geometric extension pattern of the rail is inferred through a generative residual compensation branch to achieve implicit repair of the rail manifold in the feature space, thereby obtaining a repaired feature map.

[0034] The characteristic fracture region includes at least one fracture point.

[0035] Optionally, the feature fracture regions of the track are located in the enhanced feature map, and the geometric extension information of the track is inferred through a pre-built generative residual compensation module to repair the feature fracture regions in the feature space and obtain a repaired feature map, including: The enhanced feature map is mapped to a single-channel space through a convolutional layer, and a saliency map is generated using a first activation function. Each value in the saliency map quantitatively represents the structural contribution of the corresponding pixel position to the geometric reconstruction of the track. The enhanced feature map is then denoised using the saliency map to obtain a denoised feature map. The denoised feature map is then upsampled to restore the geometric edge features of the track in the denoised feature map. Based on the second activation function, the upsampled denoised feature map, and preset convolution weights, the geometric extension information of the center position is inferred from the clear texture in the neighborhood of the break point, and a compensation feature map is generated. Based on the compensation feature map and the enhanced feature map, a repaired feature map is obtained.

[0036] Understandably, through a Convolutional layer The enhanced feature map output by the spatial prior feature enhancement module Mapped to a single-channel space and using a first activation function (such as Sigmoid). Generate saliency map As shown in formula (4).

[0037] Formula (4) In the formula, the significance graph Each value in the graph quantitatively represents the structural contribution of the corresponding location to the orbital geometry reconstruction; the network learns automatically through end-to-end training, enabling the real orbital region to be accurately represented. The value approaches 1, while the reflective spots or flying insect noise areas The value approaches 0. It means The weights of the convolutional layers are used to transform multi-scale features into single-channel spatial features.

[0038] Understandably, when A value close to 1 indicates a high-confidence orbital structure at that location. When... When the value approaches 0, it indicates that the location is background noise or a reflective bright spot. The enhancement feature map is multiplied element-wise using this saliency map to obtain the denoised feature map. By constructing a differentiable soft threshold gate, the discrete high-frequency noise in the background is suppressed while preserving the continuity of the feature manifold, thus achieving soft threshold denoising, as shown in formula (5).

[0039] Formula (5) In the formula, This represents the denoised feature map.

[0040] Understandably, the generative residual compensation branch, in order to repair weak texture regions that were accidentally damaged by denoising or regions that were originally occluded, constructs a branch in the local structure restoration module that includes an upsampling operator. The generative branch of the residual refinement block is used to upsample the weighted / filtered features (i.e., the denoised feature map), such as... It aims to restore geometric edges at a finer granular level.

[0041] Understandably, after upsampling, through stacking The convolutional layer (also known as the residual thinning block) extracts the local geometric extension pattern in the upsampled denoised feature map. Its logic is based on the receptive field characteristics of convolution: even if the central pixel is occluded, the convolutional kernel can still cover the clear orbital pixels around it. Subsequently, by learning the linear or curvilinear arrangement of the surrounding pixels, the feature distribution that the central pixel should have is inferred, a texture compensation signal is generated, and the compensation feature map is obtained, as shown in formula (6).

[0042] Formula (6) In the formula, This refers to the compensation signal. , To preset the convolution weights, The second activation function (e.g., ReLU) is used to infer the geometric connectivity of the center location from the clear texture of the neighborhood of the breakpoint.

[0043] Understandably, based on the compensated feature map and the upsampled enhanced feature map, the repaired feature map is obtained, as shown in formula (7). That is, the compensated signal is superimposed back onto the original feature stream, and the repaired feature is output. Through the above superposition operation, the information gaps caused by physical occlusion or illumination interference are implicitly filled in the feature manifold, ensuring that the output feature is topologically connected.

[0044] Formula (7) In the formula, This refers to repairing the feature map.

[0045] S104. Based on the camera parameters when the target camera captures the railway scene image, the repaired feature map is mapped to a three-dimensional physical space to construct a track feasible region mask surrounding the track.

[0046] Understandably, based on the above S103, a 3D orbital spatial attention module is introduced. This module utilizes the camera imaging model to establish a differentiable inverse mapping from a 2D image plane to a 3D physical space, explicitly mapping the 2D feature map to the 3D physical space. It then constructs a 3D hazard zone entity surrounding the orbit and projects it to generate a physically feasible region mask, thus resolving depth ambiguity in monocular vision. Subsequently, this mask can be used to perform hard spatial gating on the features, outputting the 3D hazard zone.

[0047] Optionally, based on the camera parameters when the target camera captured the railway scene image, the repaired feature map is mapped to a three-dimensional physical space to construct a track feasible region mask surrounding the track, including: The system acquires camera parameters when capturing railway scene images from the target camera and calculates the correspondence between the camera coordinate system and the actual coordinate system based on these parameters. It predicts the initial track line based on the repair feature map and constructs a three-dimensional hazard zone entity of the track in three-dimensional physical space based on the initial track line and the correspondence. It projects the three-dimensional hazard zone entity onto a two-dimensional image plane and performs rasterization to generate a binary mask. It performs linear interpolation on the binary mask to obtain a soft mask and constructs a track feasible region mask surrounding the track based on the soft mask.

[0048] Understandably, the camera parameters of the target camera when capturing images of a railway scene are pre-calibrated or acquired. These camera parameters include intrinsic and extrinsic parameters, such as the focal length. He Guangxin External parameters include installation height, etc. and pitch angle etc. When setting the world coordinate system with the road surface as the reference plane... In this case, define the rotation matrix of the camera coordinate system. Description of the camera pitch angle The resulting coordinate transformation is shown in formula (8).

[0049] Formula (8) Understandably, after defining the rotation matrix in the camera coordinate system, an intermediate projection variable is introduced. It is used to describe the pixel's ordinate. With physical depth The coupling relationship between them is shown in formula (9).

[0050] Formula (9) In the formula, The equivalent focal length of the camera in the vertical direction. The ordinate of the pixel relative to the optical center is the centered ordinate.

[0051] Understandably, after determining the intermediate projection variables, the camera projection equation, the ground plane geometric constraint equation, and the geometric constraint of the camera installation height are combined to derive the direct closed-form solution of the physical depth, as shown in formula (10).

[0052] Formula (10) In the formula, the combined camera projection equation is: The geometric constraint equations for the ground plane are as follows: The camera installation height is Physical depth is This formula shows that, under the assumption of a flat roadbed, there is a one-to-one monotonic mapping relationship between the ordinate of a pixel in an image and its physical depth, and this mapping is completely differentiable with respect to camera parameters and pixel coordinates. For example, pixels further down the image (larger) are closer in physical distance, while pixels further up the image (smaller) are farther away.

[0053] Understandably, based on the calculated physical depth and lateral coordinates, the complete three-dimensional world coordinates corresponding to each pixel are recovered using the back projection formula, as shown in formula (11).

[0054] , Formula (11) In the formula, where Centered x-axis Horizontal focal length It is the transpose of the rotation matrix; thus, it realizes the accurate inverse reconstruction from two-dimensional visual observation to three-dimensional physical space, and determines the correspondence between the camera coordinate system and the actual coordinate system.

[0055] Understandably, after completing the 3D reconstruction, the next step is to construct the three-dimensional danger zone and modulate the linear features. Based on the 2D repair feature map, the initial track line is predicted, and the roadbed plane is fitted in 3D space, such as... Figure 2 As shown. Then, along the roadbed normal vector... Raise a preset physical height in three-dimensional physical space This constructs a semi-transparent three-dimensional prism, essentially creating a three-dimensional hazardous area entity for the track in three-dimensional physical space, such as... Figure 3 As shown, this entity precisely defines the clearance area required for train operation in physical space. The preset physical height is determined based on a correspondence and the fixed height of the train (e.g., 2m-3m), essentially mapping the train's height in physical space to the image space. Subsequently, the three-dimensional hazard zone entity is reprojected onto the two-dimensional image plane, and a rasterization operation is performed to generate a binary mask, denoted as [image image]. To match the feature map resolution and avoid gradient truncation, bilinear interpolation is performed on the binary mask to obtain a soft mask. And construct the orbital feasible domain mask that surrounds the orbit based on the soft mask.

[0056] Optionally, a soft mask is obtained by linear interpolation of the binary mask, and an orbital feasible region mask surrounding the orbit is constructed based on the soft mask, including: Based on the preset base gain coefficient, the preset enhancement coefficient of the mask region, and the preset non-zero safety factor, the features in the physical feasible region of the soft mask that have the risk of train collision are enhanced to construct the track feasible region mask surrounding the track. Among them, the preset base gain coefficient is used as the reference response of the position features, the preset enhancement coefficient is used to amplify the gradient weight of the features located in the physical feasible region, and the preset non-zero safety factor is used to retain the gradient of the non-track feasible region.

[0057] Understandably, after obtaining the soft mask, a linear spatial gating mechanism with baseline gain is introduced to control the fundamental features. Modulation is performed as shown in formulas (12) and (13).

[0058] Formula (12) Formula (13) In the formula, It is a soft mask after interpolation and smoothing. It is the base gain coefficient, used to maintain the baseline response of the characteristic; For the enhancement coefficient of the masked region, it is used to significantly amplify the gradient weights of features located within the physically feasible region; This is a non-zero baseline term used to preserve weak gradients in non-track regions, preventing neuron death during backpropagation and thus preventing the background gradient from completely disappearing. Through this gating mechanism, the network is forced to focus on spatial regions where there is a physical risk of train collision during the feature extraction stage, thereby eliminating semantic ambiguity between the roadside background and the foreground track caused by perspective effects.

[0059] Understandably, through this mechanism, the track domain gains stronger gradient-driven capabilities due to its high gain, while the responses of physically infeasible areas such as the roadside background and sky are effectively suppressed. Simultaneously, it constrains the network to focus its optimization on geometrically consistent regions jointly established by 3D geometric priors, track gauge constraints, and camera parameters, fundamentally eliminating visual semantic ambiguity.

[0060] S105. Perform multi-target detection based on the orbital feasible region mask and generate detection results.

[0061] Multi-target detection includes track detection and foreign object detection. The detection results include the geometric topology information of the track and the detection information of foreign objects that pose an operational risk to the track.

[0062] Understandably, based on the above S104, in order to resolve the feature conflict between the classification task's focus on texture semantics and the regression task's focus on edge contours in object detection, this disclosure adopts a decoupled detection head design. Specifically, the multi-scale feature map from the path aggregation network first passes through a convolutional layer to adjust the number of channels, and then, under the modulation of the physical feasible region mask, is divided into two parallel convolutional branches, such as a classification branch (i.e., the foreign object detection branch) and a regression branch. Specifically, the physical feasible region mask is injected as a hard spatial prior into the foreign object detection branch to locate and classify risk targets within the physical constraints, outputting a foreign object detection bounding box, i.e., foreign object detection information. Simultaneously, the features gated by the physical space are input into the topological constraint polyline decoder, and a differentiable expectation regression mechanism is used to transform discrete predictions into continuous integrals, combined with vertical monotonic interpolation, to output a track geometric polyline, i.e., geometric topological information.

[0063] Optionally, multi-target detection is performed based on the orbital feasible region mask to generate detection results, including: The basic feature map is modulated using a track feasible region mask to obtain an track feature map with enhanced features within the track feasible region; the probability of foreign objects existing at each pixel position in the track feature map is predicted by the first detection module, and the generated foreign object detection information is obtained; the geometric parameters of the bounding box used to select the track feasible region are predicted by the second detection module based on the track feature map, and the generated geometric topology information is obtained.

[0064] Understandably, the basic feature map is modulated using a track feasible region mask to obtain a track feature map with enhanced features within the track feasible region. The classification branch is denoted as the first detection module, and the regression branch as the second detection module. The classification branch contains two layers. Convolution, the final output channel number is The feature map is used to predict the probability of various foreign objects being present at each grid point using the sigmoid function, and to generate foreign object detection information. The regression branch contains two layers. Convolution, the final output is a feature map with 4 channels, used to predict the geometric parameters of the bounding box and generate geometric topology information.

[0065] Optionally, predict the probability of various foreign objects existing at each pixel location in the orbital feature map, and obtain the generated foreign object detection information, including: Predict the probability of a foreign object at each pixel location in the track feature map and output the initial coordinates of the foreign object detection box; decode the initial coordinates into target coordinates in the image coordinate system based on the downsampling step size and preset anchor box size of the track feature map; calculate the center point coordinates of the foreign object detection box based on the target coordinates and determine whether the center point of the foreign object detection box is located within the track feasible region; if the center point is located outside the track feasible region, the target object selected by the foreign object detection box is determined as a background object; or, if the center point is located within the track feasible region, the target object is determined as an alarm object; obtain the foreign object detection information of the alarm object, wherein the foreign object detection information includes the object category, the confidence level of the foreign object detection box, and the target coordinates.

[0066] Understandably, the rectangle coordinate decoding and acquisition network does not directly predict the absolute coordinates of the rectangle, but rather predicts the offset based on the grid. For the position on the feature map... The classification branch predicts the probability of a foreign object being present at each pixel location in the orbital feature map and outputs four parameters. This refers to the initial coordinates of the foreign object detection bounding box. Subsequently, based on the downsampling step size of the orbital feature map and the preset anchor box size, the initial coordinates are decoded into the target coordinates in the image coordinate system, thus determining the foreign object detection bounding box in the image coordinate system. As shown in formula (14).

[0067] Formula (14) In the formula, For the Sigmoid function, The downsampling step size of the current feature map. , The width and height of the preset anchor frame are then obtained. The precise location of the foreign object on the image plane was determined, and the final output effect is as follows. Figure 4 As shown, the target object appearing within the feasible orbital region is precisely selected, and a notification is issued.

[0068] Understandably, after determining the foreign object bounding box, geometric filtering based on the physical feasible region yields initial candidate bounding boxes. Subsequently, an asymmetric interaction mechanism is introduced to perform secondary filtering of the bounding boxes. Although feature-level physical mask injection has suppressed most background responses, explicit geometric filtering is performed to ensure zero false alarms. This involves calculating the center point of each candidate bounding box based on the target coordinates and determining whether the center point is located within the non-zero region of the physical feasible region mask in spatial geometric modeling, i.e., whether the center point is within the track's physical feasible region. If the center point is outside the physical feasible region, even if its confidence score is high, it is classified as visual background and directly eliminated. Only foreign object detection targets located within the physical safety boundary (physical feasible region) are retained as the final risk alarm targets.

[0069] The foreign object detection bounding box output by the detection head includes the target category, confidence score, and bounding box coordinates. Only when the center point of the bounding box is within the two-dimensional projection range of the entity in the three-dimensional hazard zone is it marked as a valid risk target and an alarm is triggered, thereby significantly reducing the false detection rate caused by roadside trees, utility poles, etc.

[0070] Understandably, during the training phase, the loss function for foreign object detection is composed of a weighted average of classification loss, regression loss, and confidence loss. The classification loss... Binary cross-entropy loss is used to supervise the prediction accuracy of the target category. Regression loss. The IoU loss is used to directly measure the overlap between the predicted bounding box and the ground truth bounding box. The formula is as follows: Furthermore, this loss is invariant to scale changes. Confidence loss. The BCE loss function, used for binary classification problems, is employed to monitor the presence of a target within the grid. The final total loss function is... By jointly optimizing, geometric reconstruction and target detection can achieve synergistic convergence.

[0071] Optionally, the geometric parameters of the bounding box used to select the feasible region of the orbit are predicted based on the orbit feature map, and the generated geometric topology information is obtained, including: A target vector of a set length is generated based on the ordinate of the image from the orbit feature map. A logarithmic probability value at a horizontal coordinate index is predicted for each row of pixels in the orbit feature map, and this logarithmic probability value is normalized to a probability density function based on a temperature parameter. The range of the horizontal coordinate index is determined based on the target vector. The expected value of the horizontal coordinates is calculated using the probability density function based on the target vector and the horizontal coordinate index, yielding the predicted horizontal coordinates of consecutive sampling points. The start and end points of the orbital polyline are predicted using a pre-trained parameterized monotonic model, and the predicted vertical coordinates of all sampling points between the start and end points are generated using linear interpolation. Geometric topology information includes the predicted horizontal and vertical coordinates of different sampling points.

[0072] Understandably, this disclosure employs a row-based expectation regression mechanism, utilizing a topologically constrained piecewise linear decoder to decode high-dimensional features into the final orbital curve. In the topologically constrained piecewise linear decoder, a differentiable expectation regression mechanism replaces discrete classification. Specifically, for the vertical coordinate of the orbital feature map, the network predicts a length of... The vector is denoted as the target vector. Subsequently, for each row of the orbit feature map, the decoder predicts a horizontal position index. The logarithmic odds value is calculated, and a Softmax function with a temperature parameter is introduced to normalize the logarithmic odds value into a probability density function. As shown in formula (15), the index range of the horizontal coordinate index is determined based on the target vector, i.e. .

[0073] Formula (15) In the formula, For temperature coefficient; when At that time, the distribution tended to be similar to that of Dirac. Functions that improve positioning accuracy; when At this time, the distribution tends to be uniform, preserving neighborhood information. Understandably, in the early stages of training, temperature can be... Setting the temperature to a larger value smooths the distribution and facilitates global search; in the later stages of training, the temperature can be... Setting it to a smaller value makes the distribution tend towards the Dirac distribution, thus improving positioning accuracy.

[0074] Understandably, after determining the probability density function, the expected value of the lateral coordinates is calculated using the probability distribution to obtain the predicted lateral coordinates of the continuous sampling points. As shown in formula (16).

[0075] Formula (16) In the formula, Normalize the discrete index to The expected value calculation process is fully differentiable with respect to the feature map input, allowing the gradient to propagate directly back to the feature extraction layer, thereby achieving sub-pixel level localization accuracy. This step transforms discrete grid indices into continuous normalized coordinates. And it is differentiable throughout the entire process.

[0076] Understandably, to prevent loops or out-of-order predictions in the vertical axis of the predicted trajectory, the decoder does not directly predict the ordinate of each point. Instead, it uses a parametric monotonic model to predict the starting and ending points of the trajectory polyline. With the termination endpoint And constrain it to [the specified value] using the Sigmoid function. Within the interval, the longitudinal predicted coordinates of all intermediate sampling points are then generated using a linear interpolation formula. As shown in formula (17).

[0077] Formula (17) In the formula, The preset number of sampling points, The formula inherently hard-codes the longitudinal monotonicity constraint into the network structure, fundamentally ensuring the continuity of the generated orbital geometric polylines in the physical extension direction, eliminating the need for additional post-processing sorting operations. This parametric design hard-codes the longitudinal monotonicity into the network structure, requiring no additional loss function constraints.

[0078] Understandably, to ensure a smooth output curve and conformity to dynamics, the multi-order differential geometric loss function employs a combined loss function during training. This includes zero-order position loss, first-order tangential consistency loss, and second-order curvature loss, among which, zero-order position loss... The L1 norm is used to measure the absolute positional deviation between the predicted point set and the true point set in Euclidean space, as shown in formula (18).

[0079] Formula (18) In the formula, For the index of the track sampling points, and These represent the predicted coordinates and true coordinates of the i-th sampling point, respectively.

[0080] First-order tangential consistency loss To constrain the local orientation of the orbit, the cosine similarity between the predicted tangent vector and the true tangent vector is calculated, as shown in formula (19).

[0081] Formula (19) In the formula, For predicting the tangent vector With the true tangent vector The angle between them; this loss term maximizes the dot product of the tangent vectors, forcing the direction of the predicted trajectory to be consistent with the actual trajectory; Second-order curvature loss To punish unnatural trajectory abrupt changes, the norm of the rate of change of adjacent tangent vectors is minimized, as shown in Equation (20).

[0082] Formula (20) In the formula, and Let be the unit tangent vector between adjacent points, and be the second-order difference. The loss approximates the local curvature of the trajectory; this loss condition causes the predicted trajectory to satisfy... and even Continuity is maintained, effectively suppressing sawtooth oscillations and ensuring conformity to the smooth railway geometry design specifications. Penalizing drastic changes in the tangent vector, this term effectively minimizes the second derivative (curvature), ensuring the generated curve satisfies... even Continuity is important; avoid jagged, jagged vibrations.

[0083] Understandably, global Chamfer loss and Focal Loss are used to address sample imbalance and global shape matching issues. Furthermore, while the foreign object detection branch employs a detection head architecture similar to YOLO, it introduces an asymmetric interaction mechanism. In fact, the mask generated by the aforementioned 3D module is downsampled and directly concatenated or element-wise multiplied with the feature map of the detection branch. During the inference phase, the bounding boxes output by the detection head undergo an additional geometric filter, calculating whether the center point of the bounding box lies within the 3D danger zone projection range. If not, even with high confidence, it will be classified as background and directly filtered out. This design allows GRMNet to significantly reduce the false positive rate while maintaining high recall.

[0084] In one embodiment, the algorithm model is deployed on a workstation equipped with an NVIDIA GeForce RTX 4080 SUPER GPU (12GB VRAM). The software environment is based on the PyTorch 2.2.2 deep learning framework, coupled with the CUDA 12.1 acceleration library and cuDNN 8.9 to achieve tensor parallel computation. The operating system is Windows 10. Regarding data input, the received raw images are uniformly adjusted to a resolution of 1280×1280 pixels to adapt to the receptive field design of the feature extraction network and balance the computational load. Secondly, to enhance the model's generalization ability, a "real-synthetic" dual-domain mixing strategy is adopted for the training data. Synthetic data is used to provide diverse geometric topological samples, while real data is used to provide high-fidelity texture features; the mixing ratio is approximately 1:6. Furthermore, this embodiment employs a two-stage transfer learning strategy. In the first stage, during the first 50 rounds, the backbone network parameters are frozen, and only the four-directional convolutional enhancement module based on directional geometric constraints, the local structure recovery module, the topological constraint piecewise linear decoder, and the detector head are trained. The batch size is set to 8, and the learning rate is set to... The goal is to allow the randomly initialized head network to quickly adapt to pre-trained features, avoiding disruption of the backbone weights. Phase two involves the last 250 rounds, during which all layers are unfrozen for full parameter fine-tuning, the batch size is adjusted to 4, and the learning rate is decayed using a cosine annealing strategy. To simulate complex outdoor environments, various augmentation techniques were introduced during training. For example, Mosaic augmentation stitched together four images to enrich background diversity; Mixup augmentation linearly superimposed two images to improve feature robustness; and random perspective transformation simulated camera shake or distortion from different viewpoints. The total loss function, configured with weights, consists of geometric loss. and detection loss Weighted composition. For example... Figure 5 As shown, settings and It can achieve the best balance and ensure the dominance of geometric reconstruction.

[0085] The multi-object detection method provided in this disclosure constructs a geometric perception railway modeling network that deeply couples visual perception with physical constraints. Addressing the "illness" and "unstructured interference" problems in railway scene perception under monocular vision, it establishes a closed-loop inference paradigm of "feature decoupling—manifold repair—physical inverse projection—topological constraint regression." This paradigm aims to formalize the rigid geometric priors in railway engineering into differentiable operators in deep neural networks, thereby achieving accurate mapping from discrete pixel space to a topologically constrained 3D Euclidean space within an end-to-end framework.

[0086] Based on the above embodiments, Figure 6This is a flowchart illustrating another multi-object detection method provided in this embodiment. The multi-object detection method is applied to a geometrically perceptive railway modeling network, which is an end-to-end dual-task learning framework designed to simultaneously perform track geometry topology reconstruction and foreign object risk target detection. The network mainly consists of five interconnected functional modules. The feature extraction backbone network, employing a CSPDarknet combined with a path aggregation network structure, is responsible for extracting multi-scale basic visual features from the original image. The spatial prior feature enhancement module, located after the backbone network, is responsible for introducing directional inductive bias to enhance anisotropic features. The local structure restoration module is responsible for repairing topological breaks caused by occlusion or reflection. The 3D track-domain spatial attention module is responsible for generating a physical feasibility mask to eliminate spatial semantic ambiguity. The dual-branch decoding head includes a topological constraint polyline decoder and an asymmetric interactive foreign object detection head.

[0087] Understandably, regarding the feature extraction backbone network, this disclosure constructs a composite feature encoding architecture with CSPDarknet as the core backbone, deeply integrating a feature pyramid network and a path aggregation network. The input is a high-resolution image, and the output is multi-scale features. In this architecture, the backbone network first performs multi-level convolutional downsampling on the input high-resolution image, generating a series of basic feature maps with different receptive fields. Subsequently, the feature pyramid network constructs a top-down feature propagation path, progressively transferring the strong semantic information contained in the high-level feature maps to the lower levels through upsampling operations, endowing the lower-level features with semantic discriminative capabilities. Complementarily, the path aggregation network constructs a bottom-up feature aggregation path, mapping the strong localization information retained in the lower-level feature maps back to the higher levels through lateral connections. This bidirectional flowing feature pyramid structure effectively achieves the organic fusion of semantic and positional information, generating a multi-scale feature pyramid that is rich in detail and scale-robust. This mechanism provides a high-quality feature base for subsequent fine-tuning, ensuring that the model can obtain matching feature resolution whether detecting distant small foreign objects or reconstructing close-range fine track structures.

[0088] Understandably, although the aforementioned multi-scale pyramid solves the scale sensitivity problem, its output feature map is still limited at the operator level by the isotropic assumption of the standard convolution kernel. Conventional square convolution kernels have rotational invariance in spatial sampling, a characteristic that is difficult to adapt to targets with extremely anisotropic geometric features, such as railway tracks. Specifically, they appear as an approximately one-dimensional low-frequency continuous manifold in the longitudinal direction and a periodic high-frequency texture modulated by sleepers in the transverse direction. If isotropic convolution is used directly, the longitudinal extension signal of the rail is easily submerged by the random background texture of the roadbed gravel, resulting in a significant reduction in the feature signal-to-noise ratio. To address this core technical challenge, this application designs a Spatial Prior Feature Enhancement (SPFE) module, the core of which lies in explicitly introducing directional inductive bias in the convolution operation. This module constructs a set of anisotropic filter banks based on specific physical geometric scales to achieve bandpass filtering in specific directions in the frequency domain. Specifically, a longitudinal kernel configured as a long, narrow rectangle is used to respond to the low-frequency extension characteristics of the rail and roadbed edges along the image's longitudinal axis, suppressing high-frequency noise in the lateral direction by aggregating long-distance pixels along the longitudinal direction. A transverse kernel configured as a flat, rectangular shape resonates with the lateral periodic arrangement of sleepers, thereby extracting environmental context textures confirming the presence of the track area. A diagonal kernel configured at a specific angle is used to capture the oblique energy distribution in curved track scenarios. Furthermore, for nonlinear curved track scenarios, this application introduces a tangent vector sensitive matching mechanism, treating the curve as a spatial curve on a Riemannian manifold. This allows the diagonal convolution kernel to perform high-response matching of the instantaneous tangent direction of the curve along the sampled Riemannian metric path, achieving a curvature-dependent dynamic structural expression. Finally, to achieve adaptive fusion of multi-directional features, this application constructs a probabilistic weighted mechanism of "decomposition-gating-recombination". An adaptive gated sub-network is used to predict the activation probability of each directional subspace at the current pixel position, and the calibrated directional guiding features are injected into the main stream through a residual structure. This design allows the model's optimization path to naturally converge to the main direction of the geometric structure, fundamentally avoiding the gradient vanishing problem caused by texture dominance.

[0089] Understandably, environmental interference is unavoidable in practical applications of railway visual perception. Factors such as weed cover, bridge shadows, and high-reflection on rail surfaces often cause continuous tracks to appear fragmented and broken, exhibiting an unstructured state. To address this issue, this disclosure innovatively models it as a manifold completion task in a high-dimensional feature space and proposes a local structure recovery module. This module first abandons the hard threshold filtering strategy that relies on manual experience and instead constructs an adaptive saliency generation subnetwork based on a learning-based projection layer. This subnetwork generates a pixel-level saliency map by performing channel compression and nonlinear mapping on the input features. This map can quantitatively represent the confidence level of each pixel belonging to the real track structure. A soft threshold gating mechanism constructed using this saliency map can adaptively identify and suppress instantaneous high-frequency noise caused by reflections or flying insects in a differentiable manner, while preserving potential weak texture structure regions. Furthermore, to recover lost topological information, this module introduces a generative residual compensation branch. This branch leverages the receptive field diffusion characteristic of deep convolutional neural networks to establish a contextual reasoning mechanism from local to global. Even if the features of a key point are lost due to occlusion, the surrounding features still retain the geometric extension trend of the track. This branch can extract geometric patterns from the clear track segment features in the neighborhood of the break point, and infer the connection signals that should exist at the information gaps through the nonlinear transformation of the residual refinement block. This design essentially performs an implicit "geometric patching" on the feature manifold. It does not rely on explicit interpolation formulas, but generates directional texture compensation signals by learning the inherent continuity of the data distribution. Finally, this compensation signal is superimposed back onto the original feature stream, ensuring that the feature map input to the subsequent decoder maintains the integrity and continuity of the differential homeomorphism even under weak texture and strong interference conditions, fundamentally improving the model's robustness to harsh environments.

[0090] Understandably, the inherent "dimensional collapse" characteristic of monocular vision systems can cause overlap when different objects in three-dimensional physical space are projected onto a two-dimensional image plane. In railway scenarios, this phenomenon manifests as the distant roadside background often being visually confused with the foreground track area, leading to a very high false alarm rate in foreign object detection. To overcome this limitation of two-dimensional perception, this application introduces a 3D track-domain spatial attention module. Its core idea is to explicitly inject the hard geometric constraints of the physical world into the feature extraction process of the neural network. First, based on a standard pinhole camera model and the ground plane assumption unique to railway scenarios, this module constructs a closed inverse mapping solution path from two-dimensional pixel coordinates to three-dimensional world coordinates. Using this path, along with camera intrinsic and extrinsic parameters, the model can recover the physical depth of each pixel in the image in the camera coordinate system, thus achieving a transition from "image space" to "physical space." Furthermore, this depth recovery process is fully differentiable and supports end-to-end gradient backpropagation. Based on the recovered dense 3D point cloud, this application further utilizes strong prior knowledge from railway engineering, namely the standard track gauge (1.435m) and train clearance height, to reconstruct a three-dimensional hazardous zone entity surrounding the track in 3D space. This entity precisely defines the physical clearance area required for train operation. Subsequently, the system reprojects this 3D entity back onto the 2D image plane and performs polygon rasterization to generate a precise 2D physical feasibility mask. By constructing a linear spatial gating mechanism, this mask is used to modulate the feature map extracted by the backbone network pixel by pixel. This operation is equivalent to putting a "physical filter" on the neural network, giving the network "physical perception" capabilities: it not only significantly enhances the response intensity of features located inside the track domain, but more importantly, as a hard spatial prior, it fundamentally suppresses background interference located outside the physical feasible domain. This means that no matter how similar the texture of trees or background on the roadside is to foreign objects, as long as their physical coordinates are outside the hazardous zone entity, their feature response will be forced to zero. This deep coupling of physics and vision completely eliminates the negative impact of visual semantic ambiguity on foreign object detection, greatly improving the system's detection accuracy and security.

[0091] Understandably, this application also constructs a dual-branch architecture to achieve multi-task collaboration. For the track geometry branch, to overcome the quantization error and non-smoothness issues caused by discrete classification, a topological constraint piecewise linear decoder is proposed. This decoder uses differentiable expectation regression to replace the traditional maximum point calculation operation, achieving sub-pixel-level coordinate prediction, and hard-encodes the physical extension attributes of the track using longitudinal monotonicity interpolation. To ensure that the generated trajectory conforms to railway engineering dynamics specifications, this application constructs a multi-order differential geometric loss system. In addition to the zero-order position constraint, a first-order tangential consistency loss is introduced to constrain the local orientation, and a second-order curvature loss is introduced to penalize tangent vector mutations, forcing the output curve to meet or even be continuous. Furthermore, addressing the multi-task feature conflict problem mentioned in the background technology, the foreign object detection branch adopts a "decoupled detection head" architecture, performing category classification and bounding box regression separately through parallel convolutional branches, effectively balancing the translation invariance required for classification tasks and the translation sensitivity required for regression tasks. On this basis, this branch introduces an asymmetric interaction mechanism, injecting the aforementioned generated physical feasibility mask unidirectionally into the detection head as a spatial attention barrier. In the final output stage, geometric filtering logic is further executed: it calculates whether the center point of the predicted rectangle is within the two-dimensional projection range of the three-dimensional hazard zone entity. Only targets whose center point is within this physical safety boundary are judged as valid risks, thus accurately identifying foreign objects within the geometrically prior "guardrail" and realizing a logical closed loop from geometric reconstruction to risk perception.

[0092] The multi-target detection method provided in this disclosure, through the organic integration of the above modules, has significant beneficial effects. Specifically, the four-directional convolution enhancement module based on directional geometric constraints effectively decouples the main track texture from environmental background noise by introducing directional inductive bias, significantly improving feature representation ability and signal-to-noise ratio in low-contrast and complex texture environments. The implicit repair mechanism of the local structure restoration module endows the model with inference completion ability under adverse conditions such as partial occlusion and reflection, ensuring the connectivity and integrity of the output trajectory at the topological level. The introduction of 3D track domain spatial attention eliminates spatial semantic ambiguity caused by monocular vision perspective effect, ensuring that foreign object alarms are strictly limited to physical danger zones, greatly reducing the false alarm rate. The multi-order differential geometric constraints of the topological constraint polyline decoder make the generated track curve naturally meet the requirements of dynamic smoothness, and can be directly applied to high-precision trajectory planning and geometric parameter calculation without complex post-processing.

[0093] Figure 7 This is a schematic diagram of a multi-target detection device provided in an embodiment of the present disclosure. The multi-target detection device provided in this embodiment can execute the processing flow provided in the multi-target detection method embodiment, such as… Figure 7 As shown, the multi-target detection device 700 includes: The feature extraction unit 701 is used to acquire railway scene images and extract features from the railway scene images through a pre-built neural network to obtain a basic feature map, wherein the railway scene images include tracks; The feature enhancement unit 702 is used to decouple the target feature information representing the orbital geometric mode in the basic feature map in the frequency domain by pre-constructing a multi-directional filtering module, and enhance the target feature information to obtain an enhanced feature map; The feature repair unit 703 is used to locate the feature fracture region of the track in the enhanced feature map, and infer the geometric extension information of the track through a pre-built generative residual compensation module, so as to repair the feature fracture region in the feature space and obtain the repaired feature map. The feasible region construction unit 704 is used to map the repair feature map to a three-dimensional physical space based on the camera parameters when the target camera captures the railway scene image, and to construct a track feasible region mask that surrounds the track. The multi-target detection unit 705 is used to perform multi-target detection based on the track feasible region mask and generate detection results. The multi-target detection includes track detection and foreign object detection. The detection results include the geometric topology information of the track and the detection information of foreign objects that pose an operational risk to the track.

[0094] The anisotropic filtering module includes multiple sets of filters, each set of filters including parallel anisotropic convolutional layers.

[0095] Optionally, the feature enhancement unit 702 is used for: Anisotropic convolutional layers are used to perform convolution operations on the basic feature map to obtain response maps in each direction. The convolution operation refers to the weighted summation of the anisotropic convolution kernel with the local region of the basic feature map at each spatial location. Convolutional layers are used to map pixels in each directional response map to activation values ​​in each directional direction, and the activation values ​​in each directional direction are normalized to obtain the weights of each directional direction. The weights of each directional direction are used to characterize the probability that the current pixel position belongs to the geometric directional structure. The response maps of each direction are weighted and summed to enhance the attributes of each pixel in the response maps of each direction, resulting in an enhanced feature map. The attributes of the pixels include straight tracks, curved tracks, and background environment.

[0096] The characteristic fracture region includes at least one fracture point.

[0097] Optionally, the feature repair unit 703 is used for: The enhanced feature map is mapped to a single-channel space through a convolutional layer, and a saliency map is generated using the first activation function. Each value in the saliency map quantitatively represents the structural contribution of the corresponding pixel position to the orbital geometry reconstruction. The enhanced feature map is denoised using a saliency map to obtain a denoised feature map. Upsampling is performed on the denoised feature map to restore the geometric edge features of the track in the denoised feature map; Based on the second activation function, the upsampled denoised feature map, and the preset convolution weights, the geometric extension information of the center position is inferred from the clear texture of the neighborhood of the break point, and a compensation feature map is generated. Based on the compensation feature map and the enhancement feature map, the repair feature map is obtained.

[0098] Optionally, the feasible domain building unit 704 is used for: Obtain the camera parameters when the target camera captures images of the railway scene, and calculate the correspondence between the camera coordinate system and the actual coordinate system based on the camera parameters; Predict the initial trajectory line based on the repair feature map, and construct the three-dimensional danger zone entity of the trajectory in three-dimensional physical space based on the initial trajectory line and the corresponding relationship; The three-dimensional hazardous area is projected onto a two-dimensional image plane, and a rasterization operation is performed to generate a binary mask; A soft mask is obtained by linear interpolation of the binary mask, and an orbital feasible region mask that surrounds the orbit is constructed based on the soft mask.

[0099] Optionally, the feasible domain building unit 704 is used for: Based on the preset base gain coefficient, the preset enhancement coefficient of the mask region and the preset non-zero guarantee term, the features in the physical feasible region of the soft mask that have the risk of train collision are enhanced to construct the track feasible region mask surrounding the track. Among them, the preset base gain coefficient is used as the reference response of the position feature, the preset enhancement coefficient is used to amplify the gradient weight of the feature located in the physical feasible region, and the preset non-zero safety term is used to retain the gradient outside the orbital feasible region.

[0100] Optionally, the multi-target detection unit 705 is used for: By modulating the basic feature map using the orbit feasible region mask, an orbit feature map with enhanced features within the orbit feasible region is obtained; The first detection module predicts the probability of foreign objects at each pixel location in the orbit feature map and obtains the generated foreign object detection information. The second detection module predicts the geometric parameters of the bounding box used to select the feasible region of the track based on the track feature map, and obtains the generated geometric topology information.

[0101] Optionally, the multi-target detection unit 705 is used for: Predict the probability of a foreign object being present at each pixel location in the orbit feature map, and output the initial coordinates of the foreign object detection box; Based on the downsampling step size of the orbit feature map and the preset anchor frame size, the initial coordinates are decoded into target coordinates in the image coordinate system; Calculate the coordinates of the center point of the foreign object detection frame based on the target coordinates, and determine whether the center point of the foreign object detection frame is located within the feasible region of the track. If the center point is outside the feasible region of the track, the target object selected by the foreign object detection box is determined as a background object; or, if the center point is within the feasible region of the track, the target object is determined as an alarm object. Obtain foreign object detection information for the alarm object, including object category, confidence level of the foreign object detection box, and target coordinates.

[0102] Optionally, the multi-target detection unit 705 is used for: A target vector of a set length is generated based on the ordinate of the image of the orbit feature map; Based on each row of pixels in the orbit feature map, a log-odds value at a horizontal coordinate index is predicted, and the log-odds value is normalized to a probability density function based on the temperature parameter. The index range of the horizontal coordinate index is determined according to the target vector. The expected value of the horizontal coordinates is calculated using the probability density function based on the target vector and the horizontal coordinate index, thus obtaining the predicted horizontal coordinates of continuous sampling points. Predict the start and end points of the orbital polyline using a pre-trained parametric monotonic model, and generate the longitudinal predicted coordinates of all sampling points between the start and end points using linear interpolation. The geometric topology information includes the predicted lateral and longitudinal coordinates of different sampling points.

[0103] Figure 7 The multi-target detection device shown in the embodiment can be used to execute the technical solutions of the above method embodiments. Its implementation principle and technical effect are similar, and will not be repeated here.

[0104] Figure 8 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this disclosure. See below for details. Figure 8 The diagram illustrates a structural schematic suitable for implementing the electronic device 800 in the embodiments of this disclosure. The electronic device 800 in the embodiments of this disclosure may include, but is not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), in-vehicle terminals (e.g., in-vehicle navigation terminals), wearable electronic devices, etc., as well as fixed terminals such as digital TVs, desktop computers, smart home devices, etc. Figure 8 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein.

[0105] like Figure 8As shown, the electronic device 800 may include a processing device 801 (e.g., a central processing unit, a graphics processor, etc.), which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 802 or a program loaded from a storage device 808 into a random access memory (RAM) 803 to implement the multi-target detection method as described in the embodiments of this disclosure. The RAM 803 also stores various programs and data required for the operation of the electronic device 800. The processing device 801, ROM 802, and RAM 803 are interconnected via a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.

[0106] Typically, the following devices can be connected to I / O interface 805: input devices 806 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 807 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 808 including, for example, magnetic tapes, hard disks, etc.; and communication devices 809. Communication device 809 allows electronic device 800 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 8 An electronic device 800 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively.

[0107] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts, thereby implementing the multi-target detection method as described above. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 809, or installed from a storage device 808, or installed from a ROM 802. When the computer program is executed by the processing device 801, it performs the functions defined in the methods of embodiments of this disclosure.

[0108] It should be noted that the computer-readable medium described in this disclosure can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.

[0109] In some implementations, clients and servers can communicate using any currently known or future-developed network protocol such as HTTP (Hypertext Transfer Protocol) and can interconnect with digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), the Internet (e.g., the Internet of Things), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any currently known or future-developed networks.

[0110] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.

[0111] Optionally, when one or more of the above-described procedures are executed by the electronic device, the electronic device may also perform other steps described in the above embodiments.

[0112] Computer program code for performing the operations of this disclosure can be written in one or more programming languages ​​or a combination thereof, including but not limited to object-oriented programming languages ​​such as Java, Smalltalk, and C++, as well as conventional procedural programming languages ​​such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0113] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0114] The units described in the embodiments of this disclosure can be implemented in software or hardware. The names of the units are not, in some cases, intended to limit the specific unit.

[0115] The functions described above in this document can be performed at least in part by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), system-on-a-chip (SoCs), complex programmable logic devices (CPLDs), and so on.

[0116] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0117] It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or gateway that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or gateway. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or gateway that includes said element.

[0118] The above description is merely a specific embodiment of this disclosure, enabling those skilled in the art to understand or implement it. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this disclosure. Therefore, this disclosure is not to be limited to the embodiments described herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A multi-target detection method, characterized in that, include: A railway scene image is acquired, and features are extracted from the railway scene image using a pre-constructed neural network to obtain a basic feature map, wherein the railway scene image includes the track; By pre-constructing a multi-directional filtering module, the target feature information used to characterize the orbital geometric mode in the basic feature map is decoupled in the frequency domain, and the target feature information is enhanced to obtain an enhanced feature map; The characteristic fracture region of the track in the enhanced feature map is located, and the geometric extension information of the track is inferred through a pre-built generative residual compensation module, so as to repair the characteristic fracture region in the feature space to obtain a repaired feature map; Based on the camera parameters when the target camera captures the railway scene image, the repaired feature map is mapped onto a three-dimensional physical space to construct a track feasible region mask surrounding the track; Multi-target detection is performed based on the track feasible region mask to generate detection results. The multi-target detection includes track detection and foreign object detection. The detection results include the geometric topology information of the track and foreign object detection information indicating that the track poses an operational risk.

2. The method according to claim 1, characterized in that, The anisotropic filtering module includes multiple sets of filters, each set of filters including parallel anisotropic convolutional layers. The process of decoupling the target feature information representing the orbital geometric mode from the basic feature map in the frequency domain through the pre-constructed anisotropic filtering module, and enhancing the target feature information to obtain an enhanced feature map, includes: The anisotropic convolutional layer performs convolution operations on the basic feature map to obtain response maps in each direction. The convolution operation refers to the weighted summation of the anisotropic convolutional kernel with the local region of the basic feature map at each spatial location. Convolutional layers are used to map pixels in the response maps of each direction to activation values ​​in each direction, and the activation values ​​in each direction are normalized to obtain weights in each direction. The weights in each direction are used to characterize the probability that the current pixel position belongs to the geometric direction structure. The response maps of each direction are weighted and summed using the weights of each direction to enhance the attributes of each pixel in the response maps of each direction, thereby obtaining an enhanced feature map. The attributes of the pixels include straight tracks, curved tracks, and background environment.

3. The method according to claim 1, characterized in that, The characteristic fracture region includes at least one fracture point. Locating the characteristic fracture region of the track in the enhanced feature map and inferring the geometric extension information of the track using a pre-built generative residual compensation module to repair the characteristic fracture region in the feature space to obtain a repaired feature map includes: The enhanced feature map is mapped to a single-channel space through a convolutional layer, and a saliency map is generated using a first activation function, wherein each value in the saliency map quantitatively represents the structural contribution of the corresponding pixel position to the orbital geometry reconstruction. The enhanced feature map is denoised using the saliency map to obtain a denoised feature map; The denoised feature map is upsampled to recover the geometric edge features of the track in the denoised feature map; Based on the second activation function, the upsampled denoised feature map, and the preset convolution weights, the geometric extension information of the center position is inferred from the clear texture of the neighborhood of the break point, and a compensation feature map is generated. Based on the compensation feature map and the enhancement feature map, the repair feature map is obtained.

4. The method according to claim 1, characterized in that, The step of mapping the repaired feature map onto a three-dimensional physical space based on the camera parameters captured by the target camera when the railway scene image was taken, and constructing a track feasible region mask surrounding the track, includes: The camera parameters of the target camera when capturing the railway scene image are obtained, and the correspondence between the camera coordinate system and the actual coordinate system is calculated based on the camera parameters; The initial trajectory line is predicted based on the repair feature map, and a three-dimensional danger zone entity of the trajectory is constructed in three-dimensional physical space based on the initial trajectory line and the correspondence. The three-dimensional hazardous area entity is projected onto a two-dimensional image plane, and a rasterization operation is performed to generate a binary mask; A soft mask is obtained by linear interpolation of the binary mask, and an orbital feasible region mask surrounding the orbit is constructed based on the soft mask.

5. The method according to claim 4, characterized in that, The step of performing linear interpolation on the binary mask to obtain a soft mask, and constructing an orbital feasible region mask surrounding the orbit based on the soft mask, includes: Based on the preset base gain coefficient, the preset enhancement coefficient of the mask region, and the preset non-zero guarantee term, the features in the physical feasible region of the soft mask where there is a risk of train collision are enhanced to construct a track feasible region mask that surrounds the track. The preset base gain coefficient is used as the reference response of the position feature, the preset enhancement coefficient is used to amplify the gradient weight of the feature located within the physical feasible region, and the preset non-zero safety term is used to retain the gradient outside the orbital feasible region.

6. The method according to claim 1, characterized in that, The multi-target detection based on the orbital feasible region mask, generating detection results, includes: The basic feature map is modulated using the track feasible region mask to obtain a track feature map with enhanced features within the track feasible region; The first detection module predicts the probability of a foreign object at each pixel position in the orbit feature map and obtains the generated foreign object detection information. The second detection module predicts the geometric parameters of the bounding box used to select the feasible region of the track based on the track feature map, and obtains the generated geometric topology information.

7. The method according to claim 6, characterized in that, Predict the probability of various foreign objects existing at each pixel location in the orbit feature map, and obtain the generated foreign object detection information, including: Predict the probability of a foreign object being present at each pixel location in the orbit feature map, and output the initial coordinates of the foreign object detection box; Based on the downsampling step size and preset anchor frame size of the orbit feature map, the initial coordinates are decoded into target coordinates in the image coordinate system; Calculate the center point coordinates of the foreign object detection frame based on the target coordinates, and determine whether the center point of the foreign object detection frame is located within the feasible orbital region; If the center point is located outside the feasible region of the track, the target object selected by the foreign object detection box is determined as a background object; or, if the center point is located within the feasible region of the track, the target object is determined as an alarm object. Obtain foreign object detection information of the alarm object, wherein the foreign object detection information includes object category, confidence level of the foreign object detection box, and target coordinates.

8. The method according to claim 6, characterized in that, Based on the orbit feature map, predict the geometric parameters of the bounding box used to select the feasible region of the orbit, and obtain the generated geometric topology information, including: A target vector of a set length is generated based on the ordinate of the image of the orbit feature map; Based on each row of pixels in the orbit feature map, a log-odds value at a horizontal coordinate index is predicted, and the log-odds value is normalized to a probability density function based on the temperature parameter, wherein the index range of the horizontal coordinate index is determined according to the target vector; The expected value of the horizontal coordinates is calculated using the probability density function based on the target vector and the horizontal coordinate index, thereby obtaining the predicted horizontal coordinates of the continuous sampling points; The starting and ending points of the orbital polyline are predicted by a pre-trained parametric monotonic model, and the longitudinal predicted coordinates of all sampling points between the starting and ending points are generated by linear interpolation. The geometric topology information includes the lateral predicted coordinates and the longitudinal predicted coordinates of different sampling points.

9. An electronic device, characterized in that, include: Memory; processor; as well as Computer programs; The computer program is stored in the memory and configured to be executed by the processor to implement the multi-target detection method as described in any one of claims 1 to 8.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the multi-target detection method as described in any one of claims 1 to 8.