A track foreign matter multi-modal classification and identification method and system fusing point cloud intensity and geometric contour features

CN122530713APending Publication Date: 2026-08-07SICHUAN RUNYINGDA ELECTRONIC TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SICHUAN RUNYINGDA ELECTRONIC TECH CO LTD
Filing Date
2026-07-09
Publication Date
2026-08-07

AI Technical Summary

Technical Problem

现有方案中,点云反射强度易受环境光照、探测距离、入射角度的复合干扰,强度特征稳定性与可比性差,难以稳定支撑异物的材质分类识别;几何轮廓提取多采用固定方向、单一尺度的梯度算子实现,无法完整捕捉不同朝向、不同尺寸异物的轮廓信息,几何特征的表征全面性不足;多模态特征融合多采用简单通道拼接或单向注意力加权方式,难以实现强度特征与几何特征的双向互补增强,多模态特征的协同价值未得到充分挖掘;同时多数检测方案仅输出基础的异物类别与位置信息,未与轨道安全限界标准形成联动,无法自动判定异物侵限等级并输出对应分级告警信号,难以直接接入现有轨道运营管控体系,无法满足工程化应用的实际需求

Benefits of technology

[0015]本发明的有益效果:通过强度解耦校正消除多源环境干扰,通过多方向多尺度提取强化几何特征表征,通过双向交叉注意力实现两类特征互补增强,最终结合轨道安全限界输出分级告警,能够系统性提升异物识别的准确率与环境适应性,同时具备直接对接运营管控体系的工程落地能力。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122530713A_ABST
    Figure CN122530713A_ABST
Patent Text Reader

Abstract

The application discloses a track foreign matter multi-modal classification and identification method and system fusing point cloud intensity and geometric contour features, and belongs to the technical field of track traffic safety monitoring. The method comprises the following steps: acquiring track area three-dimensional point cloud data, correcting and generating a corrected intensity map through an adaptive compensation model for reflection intensity normalization; projecting the three-dimensional point cloud into a two-dimensional depth map and extracting a multi-direction and multi-scale geometric contour feature map; realizing cross-modal fusion of feature layers of two types of features through a bidirectional cross-attention mechanism to generate a joint foreign matter representation map; outputting foreign matter three-dimensional bounding boxes and category information through a multi-task network, and outputting a graded alarm signal in combination with a track three-dimensional safety limit model. Through intensity adaptive compensation, multi-scale contour extraction, bidirectional cross-modal fusion and small sample domain adaptive recognition, the application significantly improves the precision, robustness and environmental adaptability of track foreign matter identification under complex light and variable distance working conditions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of rail transit safety monitoring technology, specifically to a method and system for multimodal classification and identification of foreign objects on rails that integrates point cloud intensity and geometric contour features. Background Technology

[0002] With the continuous increase in the operating mileage and train density of my country's rail transit network, foreign object intrusion along the tracks has become a significant risk source threatening train safety and causing operational delays. Efficient and reliable foreign object detection technology is the core support for ensuring rail operation safety. Currently, foreign object detection mainly relies on two solutions: manual inspection and video image recognition. Manual inspection is inefficient and lacks real-time performance, failing to meet the needs of all-weather operation. Video image recognition is significantly affected by environmental factors such as lighting, rain, snow, and fog, with detection accuracy dropping sharply at night and in inclement weather, making it difficult to guarantee detection reliability. Point cloud detection technology based on 3D LiDAR, with its advantages of high ranging accuracy, independence from visible light conditions, and direct acquisition of target 3D geometric information, is gradually becoming the mainstream technology for rail foreign object detection. However, existing point cloud-based rail foreign object identification solutions still have several shortcomings, directly restricting detection performance and engineering feasibility. In existing solutions, point cloud reflection intensity is easily affected by the combined interference of ambient light, detection distance, and incident angle, resulting in poor stability and comparability of intensity features, making it difficult to reliably support the material classification and identification of foreign objects. Geometric contour extraction often uses gradient operators with fixed direction and single scale, which cannot fully capture the contour information of foreign objects with different orientations and sizes, resulting in insufficient comprehensiveness of geometric feature representation. Multimodal feature fusion often uses simple channel splicing or unidirectional attention weighting, making it difficult to achieve bidirectional complementary enhancement of intensity and geometric features, and the synergistic value of multimodal features has not been fully explored. At the same time, most detection solutions only output basic foreign object category and location information, without linking with track safety clearance standards, and cannot automatically determine the level of foreign object intrusion and output corresponding graded alarm signals. This makes it difficult to directly integrate into the existing track operation and control system and cannot meet the actual needs of engineering applications. Summary of the Invention

[0003] To address the shortcomings of existing technologies, the purpose of this invention is to provide a multimodal classification and identification method and system for track foreign objects that integrates point cloud intensity and geometric contour features, effectively improving the accuracy and robustness of track foreign object identification under complex working conditions.

[0004] To achieve the above objectives, the embodiments of this invention provide the following technical solutions:

[0005] This application provides a method for multimodal classification and identification of orbital foreign objects that integrates point cloud intensity and geometric contour features, including: S1, acquiring three-dimensional point cloud data of the orbital region using a three-dimensional LiDAR, wherein the three-dimensional point cloud data includes the spatial coordinates and reflection intensity information of each point; S2, establishing an adaptive compensation model for ambient light interference, distance attenuation, and incident angle, normalizing and correcting the reflection intensity information, and generating a corrected intensity map decoupled from ambient light, detection distance, and incident angle; S3, projecting the three-dimensional point cloud data into a two-dimensional depth map, and utilizing a rotatable gradient operator group. S4. Extract multi-directional, multi-scale geometric contour feature maps; S5. Through an attention-guided bidirectional cross-attention mechanism, fuse the corrected intensity map and the geometric contour feature map across modalities at the feature layer to generate a joint foreign object representation map in which intensity features and contour features enhance each other; S6. Based on the joint foreign object representation map, use a multi-task network to simultaneously perform foreign object localization and foreign object classification, output the three-dimensional bounding box of the foreign object, the foreign object category and confidence level, and determine whether the foreign object constitutes an intrusion based on the geometric parameters of the three-dimensional bounding box and the preset three-dimensional safety boundary model of the track, and output a graded alarm signal.

[0006] Furthermore, the formula for the adaptive compensation model in S2 is expressed as follows: ;in, This represents the corrected reflection intensity value; This represents the reflection intensity value obtained from the original measurement by the lidar. Represents the ambient light compensation coefficient; Represents the real-time ambient light background intensity value; Represents the distance attenuation compensation coefficient; Represents the incident angle compensation coefficient; Represents the laser incident angle; Represents the cosine factor of the incident angle; This represents the actual measured distance at which the laser-guided radar reaches the target. The reference distance is represented by the parameters of the adaptive compensation model in S2, which are dynamically updated through an online targetless automatic calibration process. Specifically, this includes: utilizing the statistical consistency of the reflection intensity of stationary objects in continuous multi-frame point cloud data, and combining the incident angle cosine factor, iteratively optimizing the distance attenuation coefficient, ambient light compensation coefficient, and incident angle compensation coefficient, thereby eliminating the influence of ambient light intensity changes, distance attenuation, and incident angle tilt on the comparability of reflection intensity.

[0007] Furthermore, S2 also includes: analyzing the multi-echo pulse signal of the three-dimensional lidar, distinguishing transparent or semi-transparent foreign objects from the background based on the intensity difference and waveform width difference between the first and last echoes, and compensating for transmission loss of the reflection intensity of the foreign objects to enhance the identifiability of the corresponding foreign objects in the correction intensity map.

[0008] Further, S3 specifically includes: constructing a Gaussian difference pyramid of the two-dimensional depth map; performing convolution along multiple preset angles using a set of directionally adjustable filters on each scale layer; taking the maximum response value of each pixel in each direction as the contour feature of that scale layer; upsampling and fusing the contour features extracted from each scale layer; wherein the angle interval of the directionally adjustable filters is no greater than 30°, and the filter scale corresponding to each scale layer matches the scale factor of the Gaussian difference pyramid.

[0009] Furthermore, S3 also includes: constructing a structure tensor matrix and calculating eigenvalues ​​for the upsampled and fused contour feature map; using the normalized ratio of the difference between eigenvalues ​​as the angle confidence level; and performing angle-preserving guided filtering to refine the contour boundary so that the sharp geometric features of the foreign object are preserved, while suppressing pseudo-contours caused by depth noise.

[0010] Further, S4 specifically includes: dividing the corrected intensity map and the geometric contour feature map into non-overlapping image blocks and embedding positional codes; using the image block features of the corrected intensity map as a first query matrix, and the image block features of the geometric contour feature map as a first key matrix and a first value matrix, obtaining contour-enhanced intensity features through multi-head cross-attention calculation; simultaneously using the image block features of the geometric contour feature map as a second query matrix, and the image block features of the corrected intensity map as a second key matrix and a second value matrix, obtaining intensity-enhanced contour features through another set of multi-head cross-attention calculations; and concatenating and fusing the intensity features and the contour features.

[0011] Furthermore, the multi-task network in S5 is pre-trained using a few-sample domain adaptation method, specifically including: using a general point cloud public dataset as the source domain and actual track scene samples as the target domain, and reducing the recognition bias caused by scene differences through maximum mean difference domain feature alignment training; the multi-task network includes a shared feature encoder, a foreign object localization sub-network and a foreign object classification sub-network, and the shared feature encoder performs deep encoding on the joint foreign object representation map.

[0012] Furthermore, during the pre-training process of the multi-task network, a dynamic weighted loss function based on task homoscedasticity uncertainty is constructed. The dynamic weighted loss function is a weighted sum of foreign object localization loss and foreign object classification loss. The weight coefficients are adaptively adjusted according to the uncertainty parameters of the corresponding tasks, so that the shared feature encoder learns a shared feature representation that is balanced and optimized for localization and classification tasks.

[0013] Furthermore, S5 also includes: matching and removing the inherent structure point cloud of the track based on the prior template of the track standard section; fitting the three-dimensional contour surface of the remaining target point cloud after removal; extracting the contour curvature, shape size and spatial distribution features as geometric auxiliary verification features; and performing temporal consistency verification on the object localization and classification recognition results of multiple consecutive frames, and outputting the final recognition result based on the temporal voting mechanism to eliminate the influence of single-frame recognition.

[0014] Accordingly, this application also provides a multimodal classification and recognition system for orbital foreign objects that integrates point cloud intensity and geometric contour features, including: a lidar data acquisition module for acquiring three-dimensional point cloud data of the orbital region using a three-dimensional lidar, wherein the three-dimensional point cloud data includes the spatial coordinates and reflection intensity information of each point; an intensity correction module for establishing an adaptive compensation model for ambient light interference, distance attenuation, and incident angle, normalizing and correcting the reflection intensity information, and generating a corrected intensity map decoupled from ambient light, detection distance, and incident angle; and a multi-scale geometric contour extraction module for projecting the three-dimensional point cloud data into a two-dimensional depth map and utilizing a rotatable gradient operator set. The system extracts multi-directional, multi-scale geometric contour feature maps; an intensity-contour cross-modal fusion module is used to fuse the corrected intensity map and the geometric contour feature map at the feature layer through an attention-guided bidirectional cross-attention mechanism to generate a joint foreign object representation map in which intensity features and contour features enhance each other; a multi-task foreign object identification module is used to simultaneously perform foreign object localization and foreign object classification based on the joint foreign object representation map using a multi-task network, and output the three-dimensional bounding box, foreign object category, and confidence level of the foreign object; a foreign object intrusion determination module is used to determine whether the foreign object constitutes an intrusion based on the geometric parameters of the three-dimensional bounding box and a preset track three-dimensional safety boundary model, and output a graded alarm signal.

[0015] The beneficial effects of this invention are as follows: by eliminating multi-source environmental interference through intensity decoupling correction, by enhancing geometric feature representation through multi-directional and multi-scale extraction, by achieving complementary enhancement of two types of features through bidirectional cross-attention, and finally by combining track safety clearance output with graded alarms, it can systematically improve the accuracy and environmental adaptability of foreign object identification, and at the same time has the engineering implementation capability to directly connect with the operation and control system. Attached Figure Description

[0016] Figure 1 A flowchart illustrating a multimodal classification and recognition method for orbital foreign objects that integrates point cloud intensity and geometric contour features, provided in an embodiment of this application;

[0017] Figure 2 This is a schematic diagram of the structure of a multimodal classification and recognition system for track debris that integrates point cloud intensity and geometric contour features, provided in an embodiment of this application. Detailed Implementation

[0018] The specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are for illustration and explanation only and are not intended to limit the scope of the present invention.

[0019] In this invention, the terms "system" and "network" are used interchangeably. "Multiple" refers to two or more; therefore, in this invention, "multiple" can also be understood as "at least two." "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, or B alone. Additionally, the character " / ", unless otherwise specified, generally indicates that the preceding and following related objects have an "or" relationship. Furthermore, it should be understood that in the description of this invention, terms such as "first" and "second" are used only for descriptive purposes and should not be construed as indicating or implying relative importance or order.

[0020] like Figure 1 As shown in the embodiment of this application, a multimodal classification and identification method for orbital foreign objects that integrates point cloud intensity and geometric contour features is provided, including: S1, acquiring three-dimensional point cloud data of the orbital region through a three-dimensional LiDAR, wherein the three-dimensional point cloud data includes the spatial coordinates and reflection intensity information of each point; S2, establishing an adaptive compensation model for ambient light interference, distance attenuation, and incident angle, normalizing and correcting the reflection intensity information, and generating a corrected intensity map decoupled from ambient light, detection distance, and incident angle; S3, projecting the three-dimensional point cloud data into a two-dimensional depth map, and using a rotatable gradient operator. S4. Extract multi-directional, multi-scale geometric contour feature maps; S5. Through an attention-guided bidirectional cross-attention mechanism, fuse the corrected intensity map and the geometric contour feature map across modalities at the feature layer to generate a joint foreign object representation map in which intensity features and contour features enhance each other; S6. Based on the joint foreign object representation map, use a multi-task network to simultaneously perform foreign object localization and foreign object classification, output the three-dimensional bounding box of the foreign object, the foreign object category and confidence level, and determine whether the foreign object constitutes an intrusion based on the geometric parameters of the three-dimensional bounding box and the preset track three-dimensional safety boundary model, and output a graded alarm signal.

[0021] In one possible embodiment, a three-dimensional lidar deployed on track-line pillars or inspection vehicles is first used to continuously scan the track area, acquiring three-dimensional point cloud data covering the track bed and the boundary areas on both sides. This three-dimensional point cloud data refers to a dataset composed of a large number of discrete spatial points, each carrying corresponding three-dimensional spatial coordinates and reflection intensity information. The spatial coordinates characterize the spatial location of the target, and the reflection intensity information reflects the target surface's ability to reflect laser signals, thus revealing the target's material properties. After the initial point cloud acquisition is completed, an intensity correction stage is performed. A pre-built adaptive compensation model is used to normalize and correct the initial reflection intensity, eliminating ambient light and detection interference. The influence of three types of interference factors—distance, incident angle, and so on—on the intensity values ​​is investigated, generating a corrected intensity map reflecting the target's own reflectivity. This corrected intensity map is a two-dimensional grayscale mapping formed by mapping the corrected intensity values ​​of the 3D point cloud according to the scanning azimuth. It can carry intensity features in an image-like form, facilitating subsequent feature extraction and fusion processing. In the synchronous or subsequent process of intensity correction, the 3D point cloud data is converted into a two-dimensional depth map using a bird's-eye view projection. The value of each pixel in the depth map corresponds to the distance and depth between the target and the radar in that azimuth. Then, a set of rotatable gradient operators is used to process the depth map, extracting edge contour information at multiple directions and scales, and finally fusing them to form a geometric contour. The feature map, in which the rotatable gradient operator set refers to a set of gradient calculation operators that can flexibly adjust the detection direction, can overcome the limitation of fixed-direction operators that can only detect edges with a specific orientation, and capture the contour information of foreign objects with different orientations. After constructing the two types of feature maps, the two modalities are fused at the feature level through an attention-guided bidirectional cross-attention mechanism. This fusion process is not a simple pixel stitching of the original image, but first maps the two types of images to a high-dimensional feature space, and then achieves mutual enhancement of the two types of features through bidirectional attention guidance, finally generating a joint foreign object representation map that simultaneously enhances the ability to distinguish materials and the ability to represent shapes. Finally, the joint foreign object representation map is input into the pre-trained multi-tasking system. In the service network, the network synchronously completes two tasks: locating and classifying foreign objects. It outputs the three-dimensional bounding box corresponding to the foreign object, its category, and the corresponding confidence level. The three-dimensional bounding box can completely represent the spatial location, size, and orientation angle of the foreign object. Based on this, the spatial range of the three-dimensional bounding box is compared with the preset track three-dimensional safety clearance model to determine whether the foreign object has intruded into the spatial boundary of safe track passage. According to the degree of intrusion and the danger level of the foreign object, the corresponding level of alarm signal is output. The track three-dimensional safety clearance model is the minimum spatial boundary model required to ensure the safe passage of trains, constructed according to the rail transit design specifications. It is the core basis for determining whether a foreign object poses a safety threat.

[0022] By eliminating multi-source environmental interference through intensity decoupling correction, enhancing geometric feature representation through multi-directional and multi-scale extraction, achieving complementary enhancement of two types of features through bidirectional cross-attention, and finally combining with track safety clearance to output graded alarms, the system can systematically improve the accuracy and environmental adaptability of foreign object identification, while also having the engineering implementation capability to directly connect with the operation and control system.

[0023] The parameters of existing strength compensation models are mostly obtained through offline laboratory calibration. The parameters are fixed and cannot be adjusted. However, factors such as the diurnal fluctuation of ambient light, the offset of equipment installation angle, and the aging of components can lead to mismatch of calibration parameters and a gradual decrease in correction accuracy over time. Traditional online calibration schemes require the deployment of special calibration targets, which are difficult and costly to implement on operating lines and cannot be updated routinely.

[0024] In this embodiment of the application, the parameters of the adaptive compensation model in S2 are dynamically updated through an online targetless automatic calibration process. Specifically, this includes: utilizing the statistical consistency of the reflection intensity of stationary objects in continuous multi-frame point cloud data, and combining the incident angle cosine factor, iteratively optimizing the distance attenuation coefficient, ambient light compensation coefficient, and incident angle compensation coefficient, thereby eliminating the influence of ambient light intensity changes, distance attenuation, and incident angle tilt on the comparability of reflection intensity.

[0025] In one possible embodiment, the adaptive compensation model normalizes the original reflection intensity through step-by-step processing of ambient light subtraction, incident angle correction, and distance attenuation compensation, ultimately generating a corrected intensity value decoupled from ambient light, detection distance, and incident angle. The calculation formula corresponding to this adaptive compensation model is as follows:

[0026] ;in, This represents the corrected reflection intensity value, which is the intensity value reflecting the reflectivity attribute of the target itself after eliminating the three types of environmental interference; The value representing the original reflection intensity obtained from the lidar measurement is the measured echo intensity that includes environmental interference. Represents the ambient light compensation coefficient, used to quantify the superposition effect of ambient background light on the intensity of laser echo; This represents the real-time ambient light background intensity value, which is obtained in real-time by the matching ambient light sensing unit on site; This represents the distance attenuation compensation coefficient, used to quantify the attenuation of laser energy as the propagation distance increases; This represents the incident angle compensation coefficient, used to correct the fitting deviation of the effect of the incident angle on the intensity. This represents the laser incident angle, which is the angle between the propagation direction of the laser beam and the normal direction of the target surface; Represents the cosine factor of the incident angle, used to compensate for the attenuation of echo intensity caused by non-perpendicular incident; This represents the actual measured distance at which the laser-guided radar reaches the target. This represents the reference distance, which is the standard distance benchmark used for distance normalization.

[0027] The parameters of this adaptive compensation model are dynamically updated through online targetless automatic calibration. During implementation, spatial registration is first performed on point cloud data from multiple consecutive frames. By comparing the spatial position changes of point clouds in different frames, stationary object point clouds with fixed positions are selected. These stationary objects are mostly inherent facilities along the track, such as rails, guardrails, and overhead contact line supports, and can be used as reference samples without the need for additional calibration targets. Based on the characteristic that the true reflectivity of the same stationary object remains constant, its corrected intensity value should maintain stable statistical consistency across multiple frames. Using this as the core constraint, a calibration loss function is constructed by combining the cosine factor corresponding to the incident angle. This loss function quantifies the fluctuation of the corrected intensity; the smaller the fluctuation, the higher the accuracy of the correction parameters. Based on this, an iterative optimization method using gradient descent is employed to iteratively update three types of parameters: distance attenuation coefficient, ambient light compensation coefficient, and incident angle compensation coefficient, until the loss function converges to the preset convergence condition, thus completing one parameter update. The formula corresponding to this iterative optimization process is as follows:

[0028] , ;in, This represents the current value of the distance attenuation compensation coefficient at the t-th iteration. This represents the current value of the ambient light compensation coefficient at the t-th iteration. This represents the current value of the incident angle compensation coefficient at the t-th iteration; This represents the distance attenuation compensation coefficient value after the (t+1)th iteration update. This represents the ambient light compensation coefficient value after the (t+1)th iteration update. This represents the incident angle compensation coefficient value after the (t+1)th iteration update; This represents the learning rate, which controls the step size for each parameter update. This represents the online calibration loss function, which uses the mean squared error to measure the deviation between the measured and estimated values. This represents the total number of sampling points for the reflection intensity of stationary objects participating in the calibration. This represents the measured value of the original reflection intensity at the i-th sampling point; This represents the original reflection intensity estimate obtained by back-calculating from the current compensation model. By synchronously and dynamically optimizing the distance attenuation compensation coefficient, ambient light compensation coefficient, and incident angle compensation coefficient, the influence of ambient light intensity changes, distance attenuation, and incident angle tilt on the comparability of reflection intensity can be eliminated, ensuring the consistency and stability of intensity characteristics under different working conditions.

[0029] By adopting a targetless automatic calibration method, there is no need to deploy additional calibration equipment and targets. The existing static facilities along the track can be used as calibration references to iteratively update the parameters of the compensation model online. This allows for real-time adaptation to dynamic changes in ambient light and equipment status, continuously ensuring the accuracy of intensity correction. At the same time, it significantly reduces the operation and maintenance costs and implementation difficulty of calibration, meeting the routine application needs of operating lines.

[0030] Most intensity correction methods are designed for opaque objects, but transparent or semi-transparent foreign objects such as plastic sheeting, dust netting, rain and fog often appear in track scenarios. When lasers penetrate these objects, multiple echoes are generated, and transmission loss will cause the reflection intensity of foreign objects to be underestimated, making these foreign objects extremely difficult to identify in intensity maps and very easy to miss. These foreign objects are precisely the common safety hazards in track operation.

[0031] In this embodiment of the application, step S2 further includes: analyzing the multi-echo pulse signal of the three-dimensional lidar, distinguishing transparent or semi-transparent foreign objects from the background based on the intensity difference and waveform width difference between the first and last echoes, and compensating for transmission loss of the reflection intensity of the foreign objects to enhance the identifiability of the corresponding foreign objects in the calibration intensity map.

[0032] In one possible embodiment, the multi-echo pulse signal output by the 3D lidar is first analyzed, and the first and last echo data corresponding to each laser emission pulse are extracted. Simultaneously, the intensity values ​​and waveform width information corresponding to the two echoes are recorded. Here, the multi-echo pulse signal refers to the reflection echoes generated successively on the surfaces of different objects when the laser beam encounters multiple layers of objects in its propagation path. The lidar can sequentially record the relevant information for each echo, where the first echo corresponds to the reflection from the outermost object, and the last echo corresponds to the reflection from the innermost background object. The intensity difference between the first and last echoes is calculated, along with the waveform width information. The difference in shape width is used to determine whether the laser beam at a given location has penetrated the outer transparent or semi-transparent object and reached the inner background. This indicates the presence of a transparent or semi-transparent foreign object at that location. Otherwise, it is considered an ordinary opaque object. For point cloud data identified as transparent or semi-transparent foreign objects, the transmission loss coefficient is calculated based on the intensity ratio of the two echoes. This coefficient is then used to compensate for and correct the reflection intensity of the foreign object. The compensated intensity value is then incorporated into the corrected intensity map, thereby improving the intensity contrast between the transparent foreign object and the background and enhancing the identifiability of such foreign objects in the corrected intensity map.

[0033] By utilizing the multi-echo characteristics of lidar to identify and distinguish between transparent and semi-transparent foreign objects, and by performing targeted transmission loss compensation, the recognition of transparent and semi-transparent foreign objects in intensity characteristics can be significantly improved, effectively reducing the missed detection rate of these common track foreign objects and filling the gap in the detection capability of existing technologies for semi-transparent foreign objects.

[0034] In this embodiment of the application, S3 specifically includes: constructing the Gaussian difference pyramid of the two-dimensional depth map; performing convolution along multiple preset angles using a set of directional adjustable filters on each scale layer; taking the maximum response value of each pixel in each direction as the contour feature of that scale layer; upsampling and fusing the contour features extracted from each scale layer; wherein the angle interval of the directional adjustable filters is no greater than 30°, and the filter scale corresponding to each scale layer matches the scale factor of the Gaussian difference pyramid.

[0035] In one possible embodiment, the extraction of multi-directional, multi-scale geometric contour feature maps begins with the construction of a Gaussian difference pyramid. First, the two-dimensional depth map undergoes multiple Gaussian blurring and downsampling processes of varying degrees to construct a multi-level Gaussian pyramid structure. Then, the Gaussian blurring results of adjacent levels are subtracted to obtain the Gaussian difference pyramid. Each level of the difference image corresponds to a detection scale; higher levels correspond to a large-scale detection range, suitable for capturing the overall contour of large foreign objects, while lower levels correspond to a small-scale detection range, suitable for capturing the detailed contours of small foreign objects. The corresponding calculation formula for this process is:

[0036] ,in Representing pixels At a scale factor of The Gaussian difference response value under these conditions; The representative scaling factor is The two-dimensional Gaussian function is used to apply Gaussian blur to an image; It represents the scaling constant between adjacent scale factors and is used to control the scale interval between different levels; Represents the two-dimensional depth map at the pixel level The depth value at that location; This represents the convolution operator.

[0037] After completing the multi-scale spatial construction, a set of directionally adjustable filters is used to perform convolution operations along multiple preset angles at each scale layer. The maximum response value of each pixel in all directions is taken as the contour feature value of that pixel at the current scale layer. The corresponding calculation formula for this process is:

[0038] ,in Representing pixels At a scale factor of The maximum contour response value under the given conditions; This represents the direction angle of the directionally adjustable filter; The preset set of directions representing the directionally adjustable filter includes multiple uniformly distributed angular directions; The direction angle is represented as , scale factor is The directional adjustable filter is used. The angle interval of the directional adjustable filter is set within a small range to ensure that it can cover contours of all orientations. The filter size corresponding to each scale layer is matched with the scale factor of that layer to avoid feature distortion caused by scale mismatch. Finally, the contour feature maps extracted from all scale layers are uniformly upsampled to the size of the original depth map and merged through weighted fusion to form the final geometric contour feature map. This allows the fused feature map to contain both large-scale overall contour information and small-scale detailed contour information.

[0039] By constructing a multi-scale detection space using the difference of Gaussian pyramid, it can simultaneously adapt to the contour detection needs of foreign objects of different sizes. Combined with an directional adjustable filter, it can achieve edge capture in all directions, ensuring the integrity of the contours of foreign objects facing different directions, and significantly improving the comprehensiveness and accuracy of geometric contour feature representation.

[0040] In this embodiment of the application, S3 further includes: constructing a structure tensor matrix and calculating eigenvalues ​​for the upsampled and fused contour feature map, using the normalized ratio of the difference between eigenvalues ​​as the angle confidence level, and performing angle-preserving guided filtering to refine the contour boundary so that the sharp geometric features of the foreign object are preserved, while suppressing pseudo-contours caused by depth noise.

[0041] In one possible embodiment, the process begins with constructing a structure tensor matrix. For each pixel in the upsampled and fused contour feature map, the horizontal and vertical gradients within its neighborhood are calculated to construct a second-order structure tensor matrix. This structure tensor matrix describes the local texture direction characteristics of the image. By performing eigenvalue decomposition on it, the edge direction and saliency of the local region can be determined. Subsequently, eigenvalue decomposition is performed on the structure tensor matrix corresponding to each pixel to obtain two eigenvalues. The normalized ratio of the difference between the two eigenvalues ​​is used as the angle confidence of the pixel. The higher the angle confidence, the stronger the contour directionality at that location, and the closer it is to the real object edge. The lower the angle confidence, the more likely the location is a pseudo-contour caused by noise. Finally, guided filtering is applied to refine the contour feature map using the angle confidence as a weight. For real contour regions with high confidence, their original angle and sharpness are preserved without smoothing. For pseudo-contour regions with low confidence, smoothing suppression is applied. The final result is a refined contour feature map that retains sharp geometric features while suppressing pseudo-contour interference.

[0042] By calculating the angle confidence of the contour based on the structure tensor and using angle-preserving guided filtering for contour refinement, it is possible to effectively suppress pseudo-contours generated by depth noise while fully preserving the sharp geometric corner features of foreign objects. This resolves the contradiction between denoising and feature preservation, and improves the purity and discriminability of contour features.

[0043] In this embodiment, step S4 specifically includes: dividing the corrected intensity map and the geometric contour feature map into non-overlapping image blocks and embedding positional codes; using the image block features of the corrected intensity map as a first query matrix, and the image block features of the geometric contour feature map as a first key matrix and a first value matrix, obtaining contour-enhanced intensity features through multi-head cross-attention calculation; simultaneously using the image block features of the geometric contour feature map as a second query matrix, and the image block features of the corrected intensity map as a second key matrix and a second value matrix, obtaining intensity-enhanced contour features through another set of multi-head cross-attention calculations; and concatenating and fusing the intensity features and the contour features.

[0044] In one possible embodiment, the process begins with image patch segmentation and positional encoding. The calibration intensity map and the geometric contour feature map are each divided into several non-overlapping square image patches. Linear feature embedding is performed on each image patch, mapping it to a corresponding feature vector. Simultaneously, a corresponding two-dimensional positional encoding is added to each image patch to preserve the spatial location information corresponding to the features, preventing spatial location errors during the fusion process. After feature construction is completed, a first-way cross-attention calculation is performed. The image patch features of the calibration intensity map are used as the first query matrix, and the image patch features of the geometric contour feature map are used as the first key matrix and the first value matrix, respectively. This is achieved through multi-head cross-attention. Attention calculation yields enhanced contour intensity features. This process utilizes the positional information of the geometric contour to guide the intensity features to focus on the region where the foreign object contour is located, suppressing intensity interference from the background region. Simultaneously, a second-way cross-attention calculation is performed, using the image patch features of the geometric contour feature map as the second query matrix and the image patch features of the corrected intensity map as the second key matrix and second value matrix. Another set of multi-head cross-attention calculations yields enhanced contour features. This process utilizes the material differentiation information of the intensity to guide the contour features to strengthen the contour weight of highly identifiable foreign objects and weaken redundant contours in the background. The formula corresponding to the above bidirectional cross-attention calculation process is:

[0045] ;in, This represents the cross-attention output obtained by using intensity features as the query and contour features as the key, which is the intensity feature for contour enhancement; This represents the cross-attention output obtained by using contour features as the query and intensity features as the key, which is the contour feature with enhanced intensity. This represents the query matrix generated from the geometric contour feature map; This represents the bond matrix generated from the corrected intensity map; This represents the value matrix generated from the corrected intensity map; This represents the query matrix generated from the corrected intensity map; Represents the key matrix generated from the geometric contour feature map; This represents the value matrix generated from the geometric contour feature map; The feature dimension representing the key vector is used to scale the dot product result to stabilize the gradient during training and avoid gradient saturation of the Softmax function due to excessively large dot product values. Finally, the features from the two outputs are concatenated along the channel dimension, and a learnable fusion weight matrix is ​​used for dimensional transformation to obtain the final joint foreign object representation map.

[0046] By employing a bidirectional multi-head cross-attention mechanism, the intensity and contour features are enhanced in both directions, fully exploring the synergistic value of the two modalities. At the same time, image block segmentation and positional encoding are introduced to ensure the spatial alignment accuracy of the features. The generated joint representation map has both the material discrimination capability of intensity and the morphological representation capability of contour, significantly improving the discriminative power of the fused features and the accuracy of subsequent recognition.

[0047] In this embodiment of the application, the multi-task network in S5 is pre-trained using a few-sample domain adaptation method, specifically including: using a general point cloud public dataset as the source domain and actual track scene samples as the target domain, and reducing the recognition bias caused by scene differences through maximum mean difference domain feature alignment training; the multi-task network includes a shared feature encoder, a foreign object localization sub-network and a foreign object classification sub-network, and the shared feature encoder performs deep encoding on the joint foreign object representation map.

[0048] In one possible embodiment, the few-shot domain adaptation pre-training of the multi-task network begins with dataset construction, specifically building a source domain dataset and a target domain dataset. The source domain dataset uses a general public point cloud dataset containing a large number of labeled samples of different types of obstacles, providing the model with sufficient basic feature knowledge. The target domain dataset uses foreign object point cloud samples from actual track measurements, providing the model with track-specific features. Subsequently, the overall structure of the multi-task network is built, consisting of a shared feature encoder, a foreign object localization subnetwork, and a foreign object classification subnetwork. The shared feature encoder uses a deep convolutional neural network structure and is responsible for classifying the input foreign object. The representation map is used for deep feature extraction, outputting general shared high-dimensional features. The foreign object localization sub-network predicts the center position, size, and orientation angle parameters of the foreign object's 3D bounding box based on the shared features. The foreign object classification sub-network predicts the probability distribution of the foreign object belonging to each category based on the shared features. During training, source domain samples and target domain samples are input into the network simultaneously. The maximum mean difference algorithm is used to calculate the distribution difference between the source domain features and the target domain features. This distribution difference is added to the total loss function as the domain alignment loss, guiding the network to learn cross-domain general domain-invariant features, reducing the recognition bias caused by scene differences, and realizing knowledge transfer from general scenes to track scenes under small sample conditions.

[0049] By adopting a small-sample domain adaptation training strategy, the differences between scenarios are reduced through feature alignment between the source and target domains, enabling the effective transfer of general knowledge to the track scenario. Combined with the multi-task network structure of the shared encoder, a high-precision recognition model can be obtained with a small number of track test samples, effectively solving the problem of poor model generalization caused by the scarcity of track scenario samples, and significantly reducing the cost and cycle of data collection.

[0050] In this embodiment of the application, during the pre-training process of the multi-task network, a dynamic weighted loss function based on the homoscedastic uncertainty of the task is constructed. The dynamic weighted loss function is a weighted sum of the foreign object localization loss and the foreign object classification loss. The weight coefficients are adaptively adjusted according to the uncertainty parameters of the corresponding task, so that the shared feature encoder learns a shared feature representation that is balanced and optimized for localization and classification tasks.

[0051] In one possible embodiment, the construction and training of the dynamic weighted loss function first introduces the concept of homoscedastic uncertainty. A learnable uncertainty parameter is set for both the foreign object localization and foreign object classification tasks. This parameter reflects the learning difficulty and noise level of the corresponding task; higher uncertainty indicates greater learning difficulty and higher noise level. This parameter is automatically updated during training. Based on the principle of homoscedastic uncertainty, the weight coefficients of the corresponding loss terms are automatically calculated according to the uncertainty parameters of the two tasks. The higher the uncertainty of a task, the lower its corresponding loss weight, preventing the more difficult task from dominating the feature learning process and ensuring balanced learning for both tasks. The total loss function of the multi-task network is composed of a weighted average of the classification loss, localization loss, and domain alignment loss, and the corresponding calculation formula is:

[0052] ;in Represents the total training loss of the multi-task network; Cross-entropy loss, representing the task of classifying foreign objects, is used to measure the deviation between the classification prediction result and the true label. The bounding box regression loss represents the task of foreign object localization and is used to measure the deviation between the predicted bounding box and the true bounding box. The feature alignment loss represents the feature alignment loss between the source and target domains, and is used to measure the difference in feature distribution between the two domains. These represent the weight coefficients of classification loss, localization loss, and alignment loss, respectively. The weight coefficients of classification loss and localization loss are dynamically adjusted according to the uncertainty parameters of the corresponding tasks. During training, the entire network is optimized end-to-end, enabling the shared feature encoder to gradually learn shared feature representations that are balanced and optimized for both localization and classification tasks, while also adapting to the scene characteristics of the orbital target domain.

[0053] By constructing a dynamic weighted loss function based on the homoscedasticity uncertainty of each task, the corresponding loss weights can be automatically adjusted according to the uncertainty of each task during training. This achieves balanced optimization of localization and classification tasks without manual parameter tuning, enabling the shared feature encoder to learn the optimal general feature representation for both tasks and effectively improving the overall recognition performance of the model.

[0054] In this embodiment of the application, S5 further includes: matching and removing the point cloud of the inherent structure of the track based on the prior template of the track standard section; fitting the three-dimensional contour surface of the remaining target point cloud after removal; extracting the contour curvature, shape size and spatial distribution features as geometric auxiliary verification features; and performing temporal consistency verification on the object localization and classification recognition results of multiple consecutive frames, and outputting the final recognition result based on the temporal voting mechanism to eliminate the influence of single-frame recognition.

[0055] In one possible embodiment, the post-processing optimization of the identification results begins with the removal of inherent track structures. A standard track cross-section template is pre-constructed based on track design parameters. This template includes the standard three-dimensional shape of inherent track components such as rails, sleepers, and fasteners, representing the normal track shape in the absence of foreign objects. The real-time collected point cloud is registered and matched with the standard cross-section template. Point clouds with a matching degree higher than a threshold are identified as inherent track structures and removed, retaining only the remaining suspected foreign object point cloud clusters to reduce interference from background structures at the source. For the target point cloud clusters remaining after removing inherent structures, three-dimensional contour surface fitting is performed. Features such as contour curvature, shape size, and spatial distribution density are extracted as geometric auxiliary verification features and cross-validated with the recognition results output by the multi-task network to correct classification biases that clearly do not conform to geometric logic. Subsequently, a sliding time window is set to perform temporal consistency verification on the object localization and classification results of multiple consecutive frames within the window, eliminating abnormal detection results with excessively large positional jumps. Simultaneously, a weighted voting mechanism is used to weight the classification results of each frame within the window, with frames closer to the current time corresponding to higher weights. The category with the highest cumulative votes is taken as the final recognition result. The calculation formula corresponding to this temporal voting process is as follows:

[0056] ;in, Representative category The cumulative vote score within the time window; The length of the time window for sequential voting; Represents the index of the t-th frame within the time window; Represents the predicted foreign object category in frame t; This represents an indicator function; it takes the value 1 when the condition is true and 0 when the condition is false. Representing the The temporal weighting coefficient of a frame increases with its proximity to the current time. The mathematical operator for finding the maximum value of the independent variable means iterating through all candidate object categories to find the one that maximizes the cumulative voting score. The category that achieves the maximum value; The category of foreign object in the final output is the category with the highest cumulative voting score obtained through the above maximum value calculation. This time-series voting mechanism can effectively smooth the random fluctuations in single-frame recognition, eliminate random errors in single-frame recognition, and improve the stability of continuous monitoring results.

[0057] By using a priori templates of standard track sections, the inherent structure of the track is eliminated from the source, reducing background interference. At the same time, three-dimensional geometric auxiliary features are extracted for cross-validation. Combined with a time-series voting mechanism of multiple consecutive frames, random errors in single-frame recognition are effectively eliminated, significantly improving the stability and accuracy of recognition results and reducing the false alarm rate of the system.

[0058] Reference Figure 2 This application also provides a multimodal classification and recognition system for orbital foreign objects that integrates point cloud intensity and geometric contour features, including: a lidar data acquisition module for acquiring three-dimensional point cloud data of the orbital region using a three-dimensional lidar, wherein the three-dimensional point cloud data includes the spatial coordinates and reflection intensity information of each point; an intensity correction module for establishing an adaptive compensation model for ambient light interference, distance attenuation, and incident angle, normalizing and correcting the reflection intensity information, and generating a corrected intensity map decoupled from ambient light, detection distance, and incident angle; and a multi-scale geometric contour extraction module for projecting the three-dimensional point cloud data into a two-dimensional depth map and extracting the depth using a rotatable gradient operator set. The system includes: a multi-directional, multi-scale geometric contour feature map; an intensity-contour cross-modal fusion module, which uses an attention-guided bidirectional cross-attention mechanism to fuse the corrected intensity map and the geometric contour feature map at the feature layer to generate a joint foreign object representation map where intensity features and contour features enhance each other; a multi-task foreign object identification module, which uses a multi-task network to simultaneously perform foreign object localization and classification based on the joint foreign object representation map, and outputs the 3D bounding box, foreign object category, and confidence level of the foreign object; and a foreign object intrusion determination module, which determines whether the foreign object constitutes an intrusion based on the geometric parameters of the 3D bounding box and a preset track 3D safety boundary model, and outputs a graded alarm signal.

[0059] In one possible embodiment, the lidar data acquisition module consists of a 3D lidar device and a data acquisition unit, responsible for scanning the orbital area to acquire 3D point cloud data containing spatial coordinates and reflection intensity information, while outputting multi-echo pulse signals; the intensity correction module has a built-in adaptive compensation model and an online parameter identification unit, responsible for completing the normalization correction of reflection intensity, transmission compensation of transparent foreign objects, and online targetless calibration of model parameters, outputting a corrected intensity map decoupled from environmental interference; the multi-scale geometric contour extraction module has a built-in Gaussian difference pyramid unit, directional adjustable filter bank, and structure tensor filter unit, responsible for projecting the 3D point cloud into a 2D depth map and extracting multi-directional, multi-scale refined contours. The system includes a geometric contour feature map; an intensity-contour cross-modal fusion module with a built-in bidirectional multi-head cross-attention calculation unit, responsible for completing the bidirectional cross-attention fusion of two types of features and outputting a joint foreign object representation map that enhances the two types of features; a multi-task foreign object recognition module with a built-in multi-task network model trained with small sample domain adaptation, responsible for simultaneously completing the localization and classification of foreign objects, and outputting the 3D bounding box, category, and confidence level of the foreign object; and a foreign object intrusion determination module with a built-in preset track 3D safety boundary model, responsible for spatially comparing the 3D bounding box of the foreign object with the safety boundary, determining the intrusion level of the foreign object, and outputting the corresponding graded alarm signal. Each module executes in a preset process sequence to jointly complete the high-precision detection and early warning of foreign objects on the track.

[0060] The optional embodiments of the present invention have been described in detail above with reference to the accompanying drawings. However, the embodiments of the present invention are not limited to the specific details in the above embodiments. Within the scope of the technical concept of the embodiments of the present invention, various simple modifications can be made to the technical solutions of the embodiments of the present invention, and these simple modifications all fall within the protection scope of the embodiments of the present invention.

[0061] It should also be noted that the various specific technical features described in the above embodiments can be combined in any suitable manner without contradiction. To avoid unnecessary repetition, the embodiments of the present invention will not describe the various possible combinations separately.

[0062] Furthermore, various different implementations of the present invention can be combined arbitrarily, as long as they do not violate the spirit of the present invention, they should also be regarded as the content disclosed in the present invention.

Claims

1. A multimodal classification and recognition method for orbital foreign objects that integrates point cloud intensity and geometric contour features, characterized in that, include: S1. Obtain three-dimensional point cloud data of the orbital area using a three-dimensional lidar, wherein the three-dimensional point cloud data includes the spatial coordinates and reflection intensity information of each point; S2. Establish an adaptive compensation model for ambient light interference, distance attenuation, and incident angle, normalize and correct the reflection intensity information, and generate a corrected intensity map decoupled from ambient light, detection distance, and incident angle. S3. Project the three-dimensional point cloud data into a two-dimensional depth map, and use a rotatable gradient operator set to extract multi-directional, multi-scale geometric contour feature maps. S4. Through an attention-guided bidirectional cross-attention mechanism, the corrected intensity map and the geometric contour feature map are fused across modalities at the feature layer to generate a joint foreign object representation map in which intensity features and contour features enhance each other. S5. Based on the joint foreign object characterization map, use a multi-task network to simultaneously perform foreign object localization and foreign object classification, output the three-dimensional bounding box of the foreign object, the foreign object category and confidence level, and determine whether the foreign object constitutes an intrusion based on the geometric parameters of the three-dimensional bounding box and the preset three-dimensional safety boundary model of the track, and output a graded alarm signal.

2. The method for multimodal classification and recognition of orbital foreign objects by fusing point cloud intensity and geometric contour features according to claim 1, characterized in that, The formula for the adaptive compensation model in S2 is expressed as follows: ;in, This represents the corrected reflection intensity value; This represents the reflection intensity value obtained from the original measurement by the lidar. Represents the ambient light compensation coefficient; Represents the real-time ambient light background intensity value; Represents the distance attenuation compensation coefficient; Represents the incident angle compensation coefficient; Represents the laser incident angle; Represents the cosine factor of the incident angle; This represents the actual measured distance at which the laser-guided radar reaches the target. Represents a reference distance; The parameters of the adaptive compensation model in S2 are dynamically updated through an online, target-free automatic calibration process, specifically including: By utilizing the statistical consistency of the reflection intensity of stationary objects in continuous multi-frame point cloud data, and combining the incident angle cosine factor, the distance attenuation coefficient, ambient light compensation coefficient, and incident angle compensation coefficient are iteratively optimized to eliminate the influence of ambient light intensity changes, distance attenuation, and incident angle tilt on the comparability of reflection intensity.

3. The method for multimodal classification and recognition of orbital foreign objects by fusing point cloud intensity and geometric contour features according to claim 1, characterized in that, S2 also includes: The multi-echo pulse signal of the three-dimensional lidar is analyzed. Based on the intensity difference and waveform width difference between the first and last echoes, transparent or semi-transparent foreign objects are distinguished from the background. Transmission loss compensation is performed on the reflection intensity of the foreign objects to enhance the identifiability of the corresponding foreign objects in the corrected intensity map.

4. The method for multimodal classification and recognition of orbital foreign objects by fusing point cloud intensity and geometric contour features according to claim 1, characterized in that, S3 specifically includes: Construct the Gaussian difference pyramid of the two-dimensional depth map, and use a set of directional adjustable filters to perform convolution along multiple preset angles at each scale layer. Take the maximum response of each pixel in each direction as the contour feature of that scale layer. The contour features extracted from each scale layer are upsampled and fused; The angular interval of the directional adjustable filter is no greater than 30°, and the filter scale corresponding to each scale layer matches the scale factor of the Gaussian difference pyramid.

5. The multimodal classification and recognition method for orbital foreign objects based on the fusion of point cloud intensity and geometric contour features according to claim 4, characterized in that, S3 further includes: For the upsampled and fused contour feature map, a structure tensor matrix is ​​constructed and eigenvalues ​​are calculated. The normalized ratio of the difference of eigenvalues ​​is used as the angle confidence value to refine the contour boundary by angle-preserving guided filtering, so that the sharp geometric features of the foreign object are preserved, while suppressing the pseudo contours caused by depth noise.

6. The method for multimodal classification and recognition of orbital foreign objects by fusing point cloud intensity and geometric contour features according to claim 1, characterized in that, S4 specifically includes: The correction intensity map and the geometric contour feature map are divided into non-overlapping image blocks and embedded with position codes. The image block features of the correction intensity map are used as the first query matrix, and the image block features of the geometric contour feature map are used as the first key matrix and the first value matrix. The intensity features of contour enhancement are obtained by multi-head cross attention calculation. Simultaneously, using the image patch features of the geometric contour feature map as the second query matrix, and the image patch features of the corrected intensity map as the second key matrix and the second value matrix, the intensity-enhanced contour features are obtained through another set of multi-head cross attention calculations. The intensity feature and the contour feature are spliced ​​and fused together.

7. The method for multimodal classification and recognition of orbital foreign objects by fusing point cloud intensity and geometric contour features according to claim 1, characterized in that, The multi-task network in S5 is pre-trained using a few-shot domain adaptation method, specifically including: Using a general point cloud public dataset as the source domain and actual track scene samples as the target domain, we reduce the recognition bias caused by scene differences by training with feature alignment of the maximum mean difference domain. The multi-task network includes a shared feature encoder, a foreign object localization subnetwork, and a foreign object classification subnetwork. The shared feature encoder performs deep encoding on the joint foreign object representation map.

8. The method for multimodal classification and recognition of orbital foreign objects by fusing point cloud intensity and geometric contour features according to claim 7, characterized in that, During the pre-training process of the multi-task network, a dynamic weighted loss function based on task homoscedasticity uncertainty is constructed. The dynamic weighted loss function is a weighted sum of foreign object localization loss and foreign object classification loss. The weight coefficients are adaptively adjusted according to the uncertainty parameters of the corresponding tasks, so that the shared feature encoder learns a shared feature representation that is balanced and optimized for localization and classification tasks.

9. The method for multimodal classification and recognition of orbital foreign objects by fusing point cloud intensity and geometric contour features according to claim 1, characterized in that, The S5 also includes: Based on prior template matching of the standard track section, the point cloud of the inherent structure of the track is removed. The remaining target point cloud after removal is fitted with a three-dimensional contour surface, and the contour curvature, shape, size and spatial distribution features are extracted as geometric auxiliary verification features. Furthermore, the temporal consistency verification of the object localization and classification results of multiple consecutive frames is performed, and the final recognition result is output based on the temporal voting mechanism to eliminate the influence of single frame recognition.

10. A multimodal classification and recognition system for orbital foreign objects that integrates point cloud intensity and geometric contour features, characterized in that, The method for multimodal classification and recognition of orbital foreign objects by fusing point cloud intensity and geometric contour features as described in any one of claims 1 to 9 includes: The lidar data acquisition module is used to acquire three-dimensional point cloud data of the orbital area through a three-dimensional lidar. The three-dimensional point cloud data includes the spatial coordinates and reflection intensity information of each point. The intensity correction module is used to establish an adaptive compensation model for ambient light interference, distance attenuation, and incident angle, normalize and correct the reflection intensity information, and generate a corrected intensity map decoupled from ambient light, detection distance, and incident angle. The multi-scale geometric contour extraction module is used to project the three-dimensional point cloud data into a two-dimensional depth map and extract multi-directional, multi-scale geometric contour feature maps using a rotatable gradient operator set. The intensity-contour cross-modal fusion module is used to perform cross-modal fusion of the corrected intensity map and the geometric contour feature map at the feature layer through an attention-guided bidirectional cross-attention mechanism, generating a joint foreign object representation map in which intensity features and contour features enhance each other. The multi-task foreign object identification module is used to simultaneously perform foreign object localization and foreign object classification based on the joint foreign object characterization map using a multi-task network, and output the three-dimensional bounding box of the foreign object, the foreign object category and confidence level. The foreign object intrusion determination module is used to determine whether a foreign object constitutes an intrusion based on the geometric parameters of the three-dimensional bounding box and the preset three-dimensional safety clearance model of the track, and output a graded alarm signal.