A large language model-oriented multi-modal track foreign matter detection method and system

By improving the YOLOv11 network and dynamic gating fusion module, and combining the collaborative detection architecture of fine, medium and full receptive field branches, multimodal orbital foreign object detection is realized, which solves the problems of low detection accuracy and insufficient environmental adaptability in the existing technology, has intelligent judgment capability, and provides all-weather high-reliability safety detection.

CN122049607BActive Publication Date: 2026-06-26EAST CHINA JIAOTONG UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
EAST CHINA JIAOTONG UNIVERSITY
Filing Date
2026-04-17
Publication Date
2026-06-26

AI Technical Summary

Technical Problem

Existing foreign object detection technologies for railway tracks suffer from low detection accuracy, insufficient environmental adaptability, and an inability to quickly assess risk levels and provide targeted handling recommendations.

Method used

A multimodal orbital foreign object detection method for large language models is adopted. Through an improved YOLOv11 network and a dynamic gating fusion module, combined with a collaborative detection architecture of fine, medium and full receptive field branches, adaptive weighted fusion of multimodal data and lightweight channel dimensionality reduction are achieved, and the position coordinates and category confidence of the orbital foreign object are output.

Benefits of technology

It significantly improves the detection accuracy of small-sized and large-scale foreign objects, maintains stable detection performance in harsh environments, and has the ability to intelligently judge the properties and threat level of foreign objects, providing an all-weather, highly reliable security detection solution.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122049607B_ABST
    Figure CN122049607B_ABST
Patent Text Reader

Abstract

The application provides a kind of multi-modal track foreign matter detection method and system for large language model, it is related to track foreign matter detection technical field, method includes: obtaining the multi-modal data of track scene, multi-modal data includes visible light image and thermal imaging image, multiple modal characteristics are obtained by parallel inputing multi-modal data into multi-branch fusion network;The multiple modal characteristics obtained are respectively adaptively weighted and fused by dynamic gate fusion module to output multiple fusion characteristics, and the comprehensive characteristics are obtained by channel splicing to multiple fusion characteristics;The channel compression processing is carried out to comprehensive characteristics by light channel dimension reduction network, and the final detection feature is obtained by combining double residual connection mechanism, and the track foreign matter detection result is output by the neck network and detection head to realize track foreign matter detection, and the detection result includes foreign matter position coordinates and class confidence degree.The application also has the intelligent research and judgment ability to foreign matter attribute and threat level.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of orbital foreign object detection technology, and in particular to a multimodal orbital foreign object detection method and system for large language models. Background Technology

[0002] As the core component of a rail transit system, the safety of the track's operating environment directly determines the stability of train operations and the safety of passenger travel. During long-term operation, the track area is susceptible to various foreign object intrusions due to changes in the natural environment, human activities, and the surrounding environment. These intrusions include fallen tree branches, scattered construction debris, abandoned tools and equipment, wild animals, and sudden obstacles. These foreign objects pose a serious threat to the rail transit system: at best, they may cause emergency braking and operational delays; at worst, they may cause derailments, wheel and axle damage, and other major safety accidents, resulting not only in huge economic losses but also potentially serious casualties.

[0003] Currently, foreign object detection on railway tracks mainly relies on traditional detection technologies and single-modal visual perception technologies, which face numerous bottlenecks that urgently need to be addressed. Specifically, in traditional detection technologies, manual inspection methods are limited by the subjective judgment ability, physiological fatigue limits, and inspection efficiency of inspectors, making it difficult to achieve all-weather, full-coverage detection. Especially in harsh environments such as nighttime, heavy rain, and dense fog, detection accuracy can drop significantly, easily leading to missed detections and misjudgments. Furthermore, existing single-modal visual perception technologies generally lack the ability to deeply fuse multi-source sensing data, failing to fully exploit the complementary information between different modalities, resulting in insufficient environmental adaptability and robustness of the detection system. In addition, existing multimodal fusion methods often employ simple feature stitching or fixed-weight fusion strategies, failing to fully leverage the complementary advantages of visible light images and infrared images in different scenarios. Simultaneously, traditional detection systems lack the ability to intelligently assess the attributes and threat levels of foreign objects, making it difficult to quickly evaluate risk levels and provide targeted handling suggestions based on information such as the type, size, and location of foreign objects, thus failing to meet the safe operation requirements of "early detection, early assessment, and early handling" in rail transit systems. Summary of the Invention

[0004] Based on this, the purpose of this invention is to provide a multimodal orbital foreign object detection method and system for large language models, which solves the technical problem that existing detection methods lack the ability to intelligently judge the attributes and threat levels of foreign objects, making it difficult to quickly assess the risk level and provide targeted disposal suggestions based on information such as the type, size and location of foreign objects.

[0005] This invention provides a multimodal orbital foreign object detection method for large language models, comprising:

[0006] Multimodal data of the orbital scene is acquired, including visible light images and thermal imaging images. The multimodal data is input in parallel into a multi-branch fusion network to obtain multiple modal features. The multi-branch fusion network includes an improved YOLOv11 network, which includes a backbone network, a neck network, and a detection head.

[0007] The dynamic gating fusion module adaptively weights and fuses the multiple modal features obtained to output multiple fused features. The multiple fused features are then concatenated to obtain a comprehensive feature.

[0008] The comprehensive features are processed by channel compression through a lightweight channel dimensionality reduction network, and the final detection features are obtained by combining a dual residual connection mechanism. The final detection features are output as track foreign object detection results through the neck network and the detection head to realize track foreign object detection. The detection results include the location coordinates of the foreign object and the category confidence.

[0009] The aforementioned multimodal orbital foreign object detection method for large language models innovatively implements a collaborative detection architecture with three receptive field branches: fine, medium, and full. The first branch, a deep local attention mechanism, accurately captures the fine-grained features of small-sized foreign objects through deep convolution and local window attention, avoiding the detail dilution problem caused by global attention. The second branch, a cross-block attention mechanism, effectively establishes the correlation between medium-scale features. The third branch, a global position encoding attention mechanism, fully captures the global contextual information of the orbital environment. This differentiated design of the three attention mechanisms allows the network to simultaneously consider both the local detailed features and global contextual information of orbital foreign objects, significantly improving the detection accuracy for both small-sized and large-scale foreign objects. The system achieves adaptive weighted fusion of visible and infrared modes through a dynamic gating fusion module, effectively utilizing the complementary advantages of multimodal data under different environmental conditions. It maintains stable detection performance even in harsh environments such as rain, fog, and nighttime where single-mode failure occurs. A lightweight channel dimensionality reduction network ensures efficient channel compression while maintaining feature information integrity, overcoming the information loss problem caused by traditional 1×1 convolutional direct dimensionality reduction. The entire network structure design fully considers engineering deployment requirements, and the output features are fully compatible with standard detectors. It maintains high accuracy while meeting real-time processing requirements, and possesses intelligent judgment capabilities for foreign object attributes and threat levels, providing a 24 / 7, highly reliable, and intelligent safety detection solution for rail transit systems.

[0010] In addition, the multimodal orbital foreign object detection method for large language models according to the present invention may also have the following additional technical features:

[0011] Furthermore, the backbone network includes fine receptive field branches, medium receptive field branches, and full receptive field branches, with each receptive field branch outputting a modal feature, wherein:

[0012] The fine receptive field branch uses a 3×3 grid partitioning strategy to divide the feature map into 9 local regions, each with a size of H / 3×W / 3. It uses 3×3 convolutional kernels for local feature enhancement and employs a deep local attention mechanism to extract fine-grained features within a 3×3 sliding window.

[0013] The middle receptive field branch uses a 2×2 grid block strategy to divide the input feature map into 4 medium regions, and in conjunction with the cross-block attention mechanism, it models the feature relationships within the blocks and the spatial relationships between blocks.

[0014] The full receptive field branch uses a 1×1 grid block strategy to maintain the complete size of the feature map, and combines it with a global positional encoding attention mechanism to achieve contextual feature extraction across the entire map based on the introduction of learnable positional encoding;

[0015] The cross-block attention mechanism includes a first attention sub-branch and a second attention sub-branch. The first attention sub-branch corresponds to local attention within a block, and the second attention sub-branch corresponds to cross-block attention.

[0016] The calculation result of the first attention sub-branch is:

[0017]

[0018]

[0019]

[0020]

[0021] In the formula, intra out This is the result of the calculation of the first attention sub-branch; yes Normalized weights The initial weights computed for the first attention sub-branch; V intra Indicates the internal value; B indicates the corresponding batch; num_blocks indicates the number of blocks in the feature map; N indicates the number of pixels per block; C indicates the number of feature channels; where N is... × ; Represents the characteristic tensor;

[0022] The calculation result of the second attention sub-branch is:

[0023]

[0024]

[0025]

[0026]

[0027] In the formula, inter_out is the result of the calculation of the second attention sub-branch; yes Normalized weights The initial weights calculated for the second attention sub-branch; V inter B represents the cross-domain value; B represents the corresponding batch; num_blocks represents the number of blocks in the feature map; C represents the number of feature channels.

[0028] Furthermore, the steps of adaptively weighting and fusing the obtained multiple modal features through the dynamic gating fusion module to output multiple fused features include:

[0029] For each branch's visible light and infrared blocks, calculate their global average features along the spatial dimension to obtain feature vectors that characterize the overall information of each block;

[0030] The global features of the two modalities are concatenated and input into a gated network consisting of two fully connected layers and a GELU activation function. Three dynamic weight coefficients are generated through the Softmax output layer. At the same time, the cross-attention features between the two modal blocks are calculated to enhance the complementarity between the modalities. Finally, the fused features are calculated according to the fusion formula to output the fused features. The three dynamic weight coefficients include α, β and γ, where: α represents the fusion weight of the visible light modality, β represents the fusion weight of the infrared modality, and γ represents the fusion weight of the cross-attention features.

[0031] The fusion formula is:

[0032] In the formula, fused_block is the final weighted fused feature tensor; Represents visible light mode segmentation; Indicates infrared mode segmentation; X out This represents the output of the cross-attention mechanism.

[0033] Furthermore, the steps of concatenating multiple fused features to obtain a comprehensive feature include:

[0034] Multiple fused features are concatenated to obtain a tensor with 3C channels. Then, a gating network is used to generate dynamic weights: first, the concatenated tensor is globally averaged along the H×W dimension, flattened, subjected to two linear transformations, GELU nonlinear activation, and Softmax normalization to obtain branch weights that sum to 1. Then, its dimension is expanded to adapt to feature weighting. Next, the three branches are weighted and fused to obtain the core weighted feature. This is repeated three times along the channel dimension and added to the residual of the original concatenated tensor to obtain the fused feature fused_feat. Then, a lightweight dimensionality reduction network is used to gradually compress the number of channels from 3C to the target dimension C through 1×1 convolution, GELU activation, and BatchNorm normalization to obtain the dimensionality reduction feature out. Finally, a residual layer is used to directly reduce the original concatenated tensor from 3C to C, and the residual is added to out to obtain the final output comprehensive feature.

[0035] Furthermore, the step of outputting the track foreign object detection result via the neck network and the detection head to achieve track foreign object detection includes:

[0036] The fused features after channel recovery are input into the YOLO detection head. The precise location coordinates of the foreign object are obtained through the anchor point mechanism and bounding box regression. The category probability distribution of the foreign object is obtained through a multi-class classifier.

[0037] Based on multi-scale feature information and the relative position of the track, the system outputs complete detection results including the coordinates of the foreign object bounding box, category label, confidence score, and threat level, providing a basis for decision-making on graded handling by rail transit operators.

[0038] Furthermore, in the step of acquiring multimodal data of the orbital scene, the methods for acquiring multimodal data include:

[0039] The original images of the orbital scene are acquired, including visible light images and thermal imaging images. The original images are optimized using timestamp alignment and spatial registration techniques to ensure the spatiotemporal consistency of multimodal data.

[0040] The original image after achieving spatiotemporal consistency is preprocessed, and the preprocessed original image is scale-aligned using a bilinear interpolation algorithm to obtain bimodal feature data with consistent channel number and spatial size. The preprocessing includes illumination normalization, contrast enhancement, and noise filtering.

[0041] Another aspect of the present invention provides a multimodal orbital foreign object detection system for large language models, comprising:

[0042] The acquisition module is used to acquire multimodal data of the orbital scene, including visible light images and thermal imaging images. The multimodal data is input in parallel into a multi-branch fusion network to obtain multiple modal features. The multi-branch fusion network includes an improved YOLOv11 network, which includes a backbone network, a neck network, and a detection head.

[0043] The fusion module is used to adaptively weight and fuse multiple modal features obtained by the dynamic gating fusion module to output multiple fused features.

[0044] The splicing module is used to splice multiple fused features into a comprehensive feature. The comprehensive feature is then compressed through a lightweight channel dimensionality reduction network, and the final detection feature is obtained by combining a dual residual connection mechanism.

[0045] The detection module is used to output the track foreign object detection result through the neck network and the detection head to realize track foreign object detection. The detection result includes the location coordinates of the foreign object and the category confidence level.

[0046] In another aspect, the present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the multimodal orbital foreign object detection method for large language models as described above.

[0047] In another aspect, the present invention provides a data processing device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the multimodal orbital foreign object detection method for large language models as described above. Attached Figure Description

[0048] Figure 1 This is a flowchart of a multimodal orbital foreign object detection method for large language models in an embodiment of the present invention;

[0049] Figure 2 This is a schematic diagram illustrating the core grouping and segmentation concept of the multimodal fusion module in this embodiment of the invention;

[0050] Figure 3 This is a flowchart illustrating the multimodal fusion module in an embodiment of the present invention;

[0051] Figure 4 This is a flowchart illustrating the deep local attention process of the first branch in an embodiment of the present invention.

[0052] Figure 5 This is a flowchart illustrating the cross-block attention process of the second branch in an embodiment of the present invention.

[0053] Figure 6This is a flowchart illustrating the global positional attention process of the third branch in an embodiment of the present invention.

[0054] Figure 7 This is a schematic diagram of the dual-modal block fusion process in an embodiment of the present invention;

[0055] Figure 8 This is a schematic diagram of the architecture of a multimodal orbital foreign object detection system for large language models in an embodiment of the present invention;

[0056] The following detailed description, in conjunction with the accompanying drawings, will further illustrate the present invention. Detailed Implementation

[0057] To facilitate understanding of the present invention, a more complete description will be given below with reference to the accompanying drawings. Several embodiments of the invention are illustrated in the drawings. However, the invention can be implemented in many different forms and is not limited to the embodiments described herein. Rather, these embodiments are provided so that this disclosure will be thorough and complete.

[0058] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. The terminology used herein in the description of the invention is for the purpose of describing particular embodiments only and is not intended to be limiting of the invention. The term "and / or" as used herein includes any and all combinations of one or more of the associated listed items.

[0059] To facilitate understanding of the present invention, several embodiments are given below. However, the present invention can be implemented in many different forms and is not limited to the embodiments described herein. Rather, these embodiments are provided so that the disclosure of the present invention will be more thorough and complete.

[0060] Example 1

[0061] Please see Figure 1 The figure shows a multimodal orbital foreign object detection method for large language models in the first embodiment of the present invention, the method including steps S101 to S104:

[0062] S101. Acquire multimodal data of the orbital scene, including visible light images and thermal imaging images. Input the multimodal data in parallel into a multi-branch fusion network to obtain multiple modal features.

[0063] A dual-modal acquisition system, consisting of a high-resolution visible light industrial camera and a high-sensitivity infrared thermal imager, was used to acquire raw images of the track scene. The raw images included both visible light and thermal images. The raw images were optimized using timestamp alignment and spatial registration techniques to ensure the spatiotemporal consistency of the multimodal data. The spatiotemporally consistent raw images were preprocessed, and then scale-aligned using a bilinear interpolation algorithm to obtain dual-modal feature data with consistent channel numbers and spatial dimensions. The preprocessing included illumination normalization, contrast enhancement, and noise filtering.

[0064] Specifically, a synchronous triggering mechanism ensures that the original images of the track scene acquired by the visible light industrial camera and the infrared thermal imager are time-consistent. Then, the original images of the track scene undergo integrity verification, quality verification, and preprocessing to obtain the track surface image. Specifically, integrity verification detects whether the initial track surface image has missing, damaged, or abnormal conditions. Quality verification checks whether the initial track surface image meets set standards for clarity, contrast, and brightness to exclude low-quality data caused by factors such as acquisition equipment failure. Initial track surface images that fail integrity or quality verification are deemed invalid and deleted. Preprocessing operations include performing targeted correction and transformation on the two types of images respectively, using feature point matching and perspective transformation to eliminate viewpoint differences, unifying resolution through bilinear interpolation, and normalizing pixel values ​​to finally generate dual-modal feature data with consistent channel count and spatial size.

[0065] Secondly, dual-modal feature data with the same number of channels and spatial dimensions, including visible light modal images and infrared modal images, are input into the backbone network of the improved YOLOv11 network along with shared labels for feature extraction, resulting in five sets of feature maps at different scales. The label categories include fallen tree branches, scattered construction debris, abandoned tools and equipment, wild animals, and sudden obstacles. The latter three sets of feature maps at different scales are then fed into a fusion module to leverage the complementarity between different modalities to obtain a fused feature map with richer information. Specifically, this utilizes the idea of ​​grouping and partitioning; please refer to [link to relevant documentation]. Figure 2 As shown, three branches are used, each branch dividing the tensor according to its width and height dimensions, resulting in 9 equal-sized feature blocks, 4 equal-sized feature blocks, and 1 feature block, respectively. For the specific overall process of the group-block fusion mechanism, please refer to [link to documentation]. Figure 3 As shown.

[0066] The multi-branch fusion network employs a parallel architecture with non-shared parameters. In this embodiment, the multi-branch fusion network is an improved YOLOv11 network. The improved YOLOv11 network includes a backbone network, a neck network, and a detection head. Furthermore, the improved YOLOv11 network contains three structurally independent branch processing paths, each optimized for different receptive field scales. Specifically, the backbone network includes a fine receptive field branch, a medium receptive field branch, and a full receptive field branch, wherein:

[0067] The fine receptive field branch uses a 3×3 grid partitioning strategy to divide the feature map into 9 local regions, each with a size of H / 3×W / 3. It uses 3×3 convolutional kernels for local feature enhancement and employs a deep local attention mechanism to extract fine-grained features within a 3×3 sliding window, effectively enhancing the feature representation capability of small-sized foreign objects. The medium receptive field branch uses a 2×2 grid partitioning strategy to divide the input feature map into 4 medium-sized regions. It also uses a cross-partitioning attention mechanism to model the feature relationships within each partition and the spatial relationships between partitions. The full receptive field branch uses a 1×1 grid partitioning strategy to maintain the complete size of the feature map. It also uses a global positional encoding attention mechanism to achieve contextual feature extraction across the entire image by introducing learnable positional encoding.

[0068] Specifically:

[0069] (1) The calculation process of the fine receptive field branch is as follows:

[0070] In this embodiment, the fine receptive field branch is the first branch, namely branch 1. The visible light and infrared tensors of the first branch are respectively divided into 9 tensors of equal size. Among them, the visible light block of this branch is composed of B. 1r1 B 1r2 B 1r9 This indicates that the infrared blocks in this branch are represented by B; similarly, the infrared blocks in this branch are represented by B. 1t1 B 1t2 B 1t9 This indicates that for each pair of visible light blocks B... 1rj and infrared block B 1tj All will be initially extracted using 3×3 convolution, where j represents the block number. In branch 1, j = 1, 2, 3, ..., 9.

[0071] After extracting initial features through 3×3 convolutions, each pair of blocks after the convolution operation in this branch is further processed by a designed deep local attention mechanism. This attention mechanism aims to enhance local features through deep convolutions while further enhancing features using a window attention mechanism. For the execution flow of deep local attention, please refer to [link to relevant documentation]. Figure 4 As shown. Specifically:

[0072] For each input block, local features are first enhanced using depthwise convolution, and the block shape is adjusted to a three-dimensional tensor. The first dimension is batch size × number of blocks, the second dimension is block height × block width, and the third dimension is the number of channels to accommodate subsequent attention calculations. Next, a linear layer projection is used to project the adjusted tensor into a query (Q) tensor to capture the current pixel's needs, a key (K) tensor to describe the features of each pixel, and a value (V) tensor to describe the feature value of each pixel.

[0073] Then, a 3×3 window is used to extract Q, K, and V locally, i.e., divide them into smaller window sequences. Each window sequence includes a window query, a window key, and a window value. A window attention mechanism is then used to calculate this sequence, ensuring complete coverage of the core area of ​​the foreign object while minimizing background noise, thus balancing feature capture and anti-interference capabilities. This attention calculation can be expressed as:

[0074]

[0075]

[0076] Among them, Q win Indicates window query; K win Indicates the window key; V win Indicates window value; T indicates matrix transpose; This indicates that the square root of the number of channels is used as the scaling factor; Softmax is a normalization operation that maps a vector of real values ​​within a certain range to a probability distribution vector between 0 and 1, with a sum of 1; Attention indicates that Q is used... win and K win The initial window attention weights are obtained by multiplying the transpose of the expression and dividing by the scaling factor, and then normalized by Softmax. Output represents the normalized weights and the relationship between Attention and the window value V. win The output value after multiplication.

[0077] Finally, after attention calculation, all windows are merged and transformed into tensors of the same size as the original blocks through linear projection and dimension restoration.

[0078] Therefore, each block, after being processed by the aforementioned attention mechanism, yields an output of the same size as the original block. Among all the calculated output blocks, the visible light output block is composed of... , …, This indicates that the infrared output block is composed of , …, This indicates that... Further explanation is needed. or This indicates the blocks before each branch is processed. or This represents the block after processing each branch, where i is the branch number, which is branch 1 here, so i=1; j=1, 2, ..., 9.

[0079] (2) The calculation process of the receptive field branch is as follows:

[0080] In this embodiment, the receptive field branch is the second branch, namely branch 2. The visible light and infrared tensors of the second branch are respectively divided into four tensors of equal size. The visible light block of this branch is composed of... , , ..., This indicates that the infrared blocks in this branch are represented by... Similarly, the infrared blocks in this branch are... , , ..., This indicates that for each pair of visible light blocks B... 2rj and infrared block B 2tj All of them will undergo preliminary feature extraction through 5×5 convolution, where j represents the block number. In branch 2, j=1, 2, 3, 4.

[0081] After extracting initial features through 5×5 convolutions, each pair of blocks after the convolution operation in this branch undergoes a designed cross-block attention process. The core function of this attention is to perform deep modeling of medium-sized block features through a dual-attention collaborative mechanism of "intra-block local attention + inter-block cross-attention," simultaneously capturing pixel-level local details within blocks and global contextual relationships between blocks. The output is an enhanced feature map with dimensions identical to the input, providing complementary features in terms of detail and context for subsequent dual-modal dynamic fusion. For the execution flow of the cross-block attention, please refer to [link to relevant documentation]. Figure 5 As shown. This cross-block attention mechanism includes a first attention sub-branch and a second attention sub-branch. The first attention sub-branch corresponds to local attention within a block, and the second attention sub-branch corresponds to cross-block attention. Figure 5 As shown, the first attention sub-branch corresponds to Figure 5 The content on the left, the second attention sub-branch corresponds to Figure 5 The content on the right. Specifically:

[0082] The tensor format for the input of the first attention sub-branch is [B, num_blocks, C, b]. h ×b w ];

[0083] Where B represents the corresponding batch; num_blocks represents the number of blocks in the feature map, i.e., the number of blocks in the second branch; C represents the number of channels; b h Indicates the height of a single block; bw This represents the width of a single block, therefore b h ×b w This indicates the number of pixels in each block, i.e., the number of pixels per block.

[0084] First, transform the input tensor x from the above format using dimensionality transformation. ,in This represents a tensor after the sequence format has been converted, providing an adaptation dimension for subsequent pixel-level and intra-block local attention calculations, in the form of [B, num_blocks, N, C]. B represents the corresponding batch; num_blocks represents the number of blocks in the feature map; N represents the number of pixels per block; and C represents the number of feature channels; where N is... × .

[0085] Next, the Q-block, K-block, and V-block projections of the intra-block linear layer are generated, which are the internal queries, internal keys, and internal values ​​for the intra-block attention computation. This process can be represented by the following formula:

[0086]

[0087] Among them: Q intra Indicates an internal query; K intra Indicates an internal key; V intra Indicates an internal value; This represents a linear layer projection, and the tensor generated by this process is in the format [B, num_blocks, N, 3×C]; X seq This represents a tensor after converting the sequence format, which serves as the input for the internal query, internal key, and internal value; Split represents the split operation, which can generate queries, keys, and values.

[0088] This process uses a single linear layer to generate 3×C channels. Compared with the traditional method of generating query, key and value separately using 3 independent linear layers, the number of parameters is reduced by 1 / 3 and the computational efficiency is improved by 40%. Furthermore, the feature consistency of query, key and value is ensured through dimensional splitting, avoiding feature shift caused by multiple linear layers.

[0089] Secondly, the similarity between pixels is calculated by using the dot product of the query and the transpose of the key. The result of the initial weight calculation for intra-block attention is obtained by dividing by the scaling factor. Then, the normalized attention weights are obtained by Softmax normalization.

[0090] This process can be expressed by the following formula:

[0091]

[0092]

[0093]

[0094]

[0095] in, This indicates that the square root of the number of channels is used as the scaling factor; The initial weights calculated for the first attention sub-branch, i.e., the results of calculating the initial weights for local attention within the block; yes Normalized weights; Indicates the normalization operation; T represents the matrix transpose; Represents the characteristic tensor.

[0096] Furthermore, and Multiplying the results yields the computational result of the first attention sub-branch:

[0097]

[0098]

[0099]

[0100]

[0101] Intra out This is the result of the calculation of the first attention sub-branch.

[0102] The calculation of the first attention sub-branch is now complete. We will now proceed to the calculation of the second attention sub-branch.

[0103] First, the pixel sequence of each block of the tensor input to the first branch is averaged, that is, the features of N pixels are condensed into a global vector of one channel dimension, reducing the computational complexity of inter-block attention. This yields the second branch input sequence of dimension [B, num_blocks, C] to adapt to the second attention sub-branch, that is, the calculation of inter-block cross attention, where B represents the corresponding batch; num_blocks represents the number of blocks of the feature map; and C represents the number of feature channels.

[0104] Next, a second attention sub-branch is generated, namely the cross-domain query for inter-block cross-attention, with cross-domain keys and cross-domain values, corresponding to... Figure 5 This process can be expressed by the formula:

[0105]

[0106]

[0107]

[0108]

[0109]

[0110] In the formula, Q inter Indicates a cross-domain query; K inter Indicates a cross-domain key; V inter Indicates a cross-domain value; The initial weights calculated for the second attention sub-branch, i.e., the result of the initial weight calculation for inter-block attention; yes Normalized weights; This indicates a normalization operation;

[0111] The calculation process for the second attention sub-branch is the same as that for the first attention sub-branch. The output of the second attention sub-branch calculation is obtained below:

[0112]

[0113]

[0114]

[0115]

[0116] Here, inter_out is the result of the calculation of the second attention sub-branch.

[0117] Finally, the local attention outputs within a block and the cross-attention outputs between blocks are concatenated along the channel dimension to fuse local detail features and global correlation features. Then, linear projection is used to compress the channel dimension and enhance feature representation. Finally, an inverse dimensionality transformation is performed to restore the input block feature map format. It should be further noted that in the output blocks calculated in this branch, the visible light output block is composed of... , , ..., This indicates that the infrared output block is composed of , , ..., express.

[0118] (3) The calculation process of the full receptive field branch is as follows:

[0119] In this embodiment, the full receptive field branch is the third branch, i.e., branch 3. The visible light and infrared tensors of the third branch are each divided into one tensor block, i.e., maintaining their original size. The visible light block of this branch is... This indicates that; similarly, the infrared blocks in this branch are composed of... This indicates that they will then undergo preliminary feature extraction via 7×7 convolution.

[0120] After preliminary feature extraction using 7×7 convolutions, each pair of blocks resulting from the convolution operation of this branch undergoes a specially designed global positional attention mechanism. The core function of this attention mechanism is to perform full receptive field branch modeling on the complete feature map, simultaneously capturing global pixel-to-pixel correlation features and spatial location information, resulting in an enhanced feature map with strictly conserved input and output dimensions. This is suitable for capturing the global contours and cross-range correlations of large objects (over 80cm) in orbital foreign object detection. For the execution flow of this global positional attention mechanism, please refer to [link to relevant documentation]. Figure 6 As shown. The specific process is as follows:

[0121] First, the input tensor is converted from the block feature map format to the global sequence format, providing an adaptation dimension [B, N, C] for global pixel-level attention computation, where B represents the batch, C represents the number of feature channels, and N represents the total pixel value.

[0122] Next, based on these N pixels, the positional encoding of each pixel is extracted and adjusted to the current feature map size. Specifically, the positional encoding is expanded to the batch dimension to match the global sequence dimension. Thus, the final dimension of the positional encoding is also [B, N, C].

[0123] Secondly, the positional encoding is added pixel-by-pixel to the input tensor, thereby fusing feature and positional information. Essentially, this operation adds a unique positional identifier to each pixel globally using learnable parameters, addressing the limitation of global self-attention which only focuses on feature similarity and loses spatial relationships, ensuring that the spatial position of foreign objects does not shift during detection.

[0124] In this embodiment, the global sequence after injection position encoding is used This is represented by the process of projecting and splitting this sequence through a linear layer to obtain the query, key, and value required by the attention mechanism. This process can be expressed by the following formula:

[0125]

[0126]

[0127]

[0128]

[0129] in, It is a linear layer projection operation; Indicates a split operation; Q 全局 K 全局 V 全局 It is a global attention sequence The query, key, and value are obtained by performing a linear projection operation followed by a split operation.

[0130] Next, the full-pixel similarity is calculated by querying the dot product of the query and the transpose of the key. This is then multiplied by a scaling factor to mitigate the vanishing gradient in high dimensions, thus obtaining the initial global attention weights. Finally, a Softmax normalization operation is performed to obtain the normalized attention weights. This process can be expressed by the following formula:

[0131]

[0132]

[0133]

[0134]

[0135]

[0136] in, This represents the result of the initial weight calculation for global attention. yes Normalized weights. Then, the global attention weights are multiplied by the value vector, so that the features of each pixel are fused with the feature information of all related pixels globally, enhancing the global consistency of large foreign objects. This process can be expressed by the following formula:

[0137]

[0138]

[0139]

[0140]

[0141] in, To enhance the global attention output, a combination of linear layer projection, GELU activation function, and linear layer head projection is used to strengthen the features of the global attention output, avoiding feature degradation caused by global attention computation and improving the feature recognition of large foreign objects.

[0142] Finally, the dimension is restored to the feature map shape to obtain the result of this global attention calculation to realize the output of branch 3. Among them, the visible light output block of branch 3 output is composed of BO 3r1 This indicates that the infrared output block is composed of BO 3t1 This indicates that the calculation process for the three different branches has now ended.

[0143] S102. The obtained multiple modal features are adaptively weighted and fused using a dynamic gating fusion module to output multiple fused features.

[0144] As a specific example, the multiple modal features include visible light output blocks and infrared output blocks. Specifically, for each pair of visible light output blocks and infrared output blocks in each branch, they are fused to output a fused feature.

[0145] For the visible light blocks in each branch With infrared blocks This is called a pair, where i represents the branch number and j represents the block number. Taking the second branch as an example: the visible light output block of branch 2 is composed of... , , ..., This indicates that the infrared output block is composed of , , ..., Therefore and For the first pair, and For the second pair, and For the third pair, and This is the fourth pair. The purpose of step S102 is to merge each pair of blocks in each branch into one, for example, merging the first pair... and Merge into a single block The second pair and Merge into a single block .

[0146] The specific operation is as follows: Figure 7 As shown, this operation first divides the same pair of output blocks within the same branch. and output blocks Average pooling is performed along the block height and block width dimensions to obtain the average feature rgb_global of the visible light block and the average feature ir_global of the infrared block, thereby compressing the spatial dimension and preserving the global semantics.

[0147] Next, the average features of the two modalities are concatenated along the channel dimension to form global cross-modal features, which are then used as input to the gating network to generate three weights. This gating network is a neural network consisting of a linear layer, a GELU activation function, another linear layer, and a Softmax layer combined in sequence, thereby mapping the concatenated global features to three normalized weights.

[0148] The forward propagation process of the gating network can be described by the following formula:

[0149]

[0150] Among them, Linear 2C→C (﹒) represents the first linear layer, where the tensor compresses the channels to half their capacity; GELU represents the GELU activation function, which introduces non-linear changes to the entire network, adapting to complex weight mapping scenarios; Linear C→3 (﹒) represents the second linear layer, which outputs 3 weight values; Softmax is a normalization operation that ensures that the weights sum to 1, that is, the three slices of the weight tensor in the third dimension sum to 1.

[0151] Then, by dividing the original input into blocks, visible light can be segmented. Tensor Q obtained by linear layer projection 可见光 As a query, infrared segmentation Tensor K obtained by linear layer projection 红外 and V 红外 The cross-attention mechanism, which serves as both key and value, allows visible light modal features and thermal imaging modal features to complement each other, generating cross-modal enhanced features.

[0152] The generation method for queries, keys, and values ​​is consistent with the steps for generating queries, keys, and values ​​during the computation of the first or second attention sub-branch of the full receptive field branch. Then, the query Q of these two modalities... 可见光 and key K 红外 Initial weights are calculated through cross-attention and then normalized using Softmax. Finally, these weights are multiplied by the keys to obtain the output of the attention mechanism. Dimensionality reduction transforms the sequence features into block features, resulting in the output tensor. .

[0153] Finally, based on the obtained weight tensor slices, the three weight values ​​α, β, and γ, which sum to 1, are obtained respectively, and the output is segmented using the following fusion formula. and output blocks and attention output tensor The weighted summation, specifically, the fusion formula is:

[0154]

[0155]

[0156]

[0157]

[0158] Where fused_block is the final weighted fused feature tensor; α represents the fusion weight of the visible light mode, β represents the fusion weight of the infrared mode, and γ represents the fusion weight of the cross-attention features; This indicates the visible light mode output in blocks; Indicates infrared mode output blocks; X out This represents the output tensor of the cross-attention mechanism.

[0159] S103. Multiple fused features are spliced ​​together to obtain comprehensive features. The comprehensive features are then compressed using a lightweight channel dimensionality reduction network, and the final detection features are obtained by combining a dual residual connection mechanism.

[0160] In other words, the features of each branch are combined and merged, and then the weighted fusion is used to obtain the comprehensive features before dimensionality reduction and output.

[0161] Assume that after step S102, the fused feature tensor obtained by the first branch is The fused feature tensor obtained from the second branch is The fused feature tensor obtained from the third branch is They will all undergo the reverse operation of the original block operation, so that the feature tensors after fusion of each branch are spliced ​​together along the width and height dimensions and merged into the shape of the original input, resulting in three spliced ​​tensors fused1, fused2, and fused3, where fused1 is the tensor after splicing and merging the first branch; fused2 is the tensor after splicing and merging the second branch; and fused3 is the tensor after splicing and merging the third branch.

[0162] Furthermore, the entire process of three-branch fusion (fine receptive field branch, medium receptive field branch, and full receptive field branch) and weighted fusion of the merged tensors fused1, fused2, and fused3, followed by dimensionality reduction, aims to compress the number of channels from 3C to the target dimension C while maximizing the preservation of complementary features of each branch (fine-grained details, mesoscale correlations, and global contours), avoiding information loss due to dimensionality reduction. Its design is adaptable to multi-branch feature fusion scenarios (such as multi-scale foreign object capture in orbital foreign object detection), balancing feature representation capability and computational efficiency. It stabilizes the training process through dual residual connections, and dynamic weights enable adaptive complementarity of branch features.

[0163] First, the concatenated tensor of input fused1, fused2, and fused3 is processed through a gating network to generate weights. The gating network is designed to first perform global average pooling along the H×W dimension to obtain global features, which are then flattened, subjected to two linear transformations, nonlinear activation by the GELU function, and normalized by Softmax to output the normalized weights branch_weights.

[0164] Next, to adapt the weights to the spatial dimension of the feature tensors, the normalized weights are divided into three equal parts, and dimensionality expansion operations are performed on each part to convert them into tensors w1, w2, and w3 of shape [B, 1, 1, 1], to support subsequent element-wise weighted operations. These three tensors correspond to the weights of fused1, fused2, and fused3, respectively.

[0165] After weight expansion, the output tensors (fused1, fused2, fused3) of the three branches are weighted and fused based on w1, w2, and w3 to generate a core weighted feature tensor of shape [B, C, H, W], denoted as weighted_feat. This operation can adaptively integrate the feature advantage information of different branches, and its weighted fusion formula is as follows:

[0166] weighted_feat w1×fused1+w2×fused2+w3×fused3;

[0167] Where, weighted_feat∈ fused1, fused2, and fused3 are the three branch output tensors, each with the shape [B, C, H, W].

[0168] To achieve residual fusion of weighted features and the original multi-branch concatenated features, `weighted_feat` is stacked three times along the channel dimension to obtain a tensor of shape [B, 3C, H, W]: `weighted_feat_repeat` (matching the number of channels to 3C after concatenating the three branches). Then, `weighted_feat_repeat` is residually added to the feature tensors concatenated from `fused1`, `fused2`, and `fused3` to obtain the fused feature tensor `fused_feat`. This operation preserves the complete feature information of the original three branches while incorporating the dynamically weighted optimized features, effectively avoiding the loss of feature information.

[0169] Then, channel compression is performed on fused_feat using a lightweight dimensionality reduction network. The process of this lightweight dimensionality reduction network is as follows: the number of channels is reduced from 3C to 2C through 1×1 convolution, then activated by the GELU function, normalized by the BatchNorm operation, then compressed from 2C to C through 1×1 convolution, and then normalized again by BatchNorm to obtain the initial dimensionality reduction feature with shape [B, C, H, W].

[0170] To further enhance feature representation and alleviate the vanishing gradient problem, a dual residual connection is introduced: First, the 3C-channel features concatenated from fused1, fused2, and fused3 are directly reduced to the C-channel using a 1×1 convolution to obtain residual features. Then, the initial dimensionality-reduced features are added element-wise to these residual features, ultimately yielding an output tensor of shape [B, C, H, W]. This dual residual design, combined with nonlinear activation and batch normalization, stabilizes the training process, alleviates gradient vanishing, and enhances feature representation, balancing the complementarity of multi-scale foreign object detection with the computational efficiency of real-time inference.

[0171] S104. The final detection features are output through the neck network and the detection head to achieve track foreign object detection. The detection results include the location coordinates of the foreign object and the confidence level of the category.

[0172] In this embodiment, the fused features after channel recovery are input into the YOLO detection head. The precise location coordinates of the foreign object are obtained through the anchor point mechanism and bounding box regression. The category probability distribution of the foreign object is obtained through a multi-class classifier. Based on the relative position of the multi-scale feature information and the track, a complete detection result including the bounding box coordinates of the foreign object, category label, confidence score and threat level is output, providing a decision-making basis for the rail transit operation department for graded disposal.

[0173] As a concrete example, the track foreign object detection results output by the neck network and detection head in the YOLOv11 network are first subjected to semantic parsing processing adapted to the track scene. This process involves mapping pixel coordinates to track scene physical location information such as "track mileage marker XX + XX meters (up / down line) + left / right trackside area" based on target coordinates, category labels, confidence levels, and feature association vectors. Next, combined with a pre-defined track foreign object domain knowledge base, attribute labels such as "conductive / loose structure / flammable - risk level" are added to various target detection categories, such as "metal block," "loose bolt," and "plastic bag." Simultaneously, the target labeling status is determined based on a confidence threshold: targets with a confidence level of 0.9 or higher are labeled "confirmed valid," while those below 0.7 are labeled "pending verification," forming structured semantic input data that can be parsed by the Large Language Model (LLM). The collaborative process between LLM and the dual-modal detection network can be found in [link to relevant documentation]. Figure 8 .

[0174] LLM combines track operation data (operating / non-operating periods, weather, equipment parameters) and quantifies risk levels based on a three-dimensional risk matrix of foreign object category, confidence level, and operational status. It generates natural language warning text and visual annotation instructions containing the target's physical location, category attributes, and operational impact. It calls the emergency response knowledge base to match graded response instructions and constructs a "detection-assessment-warning-response" decision chain. At the same time, the risk decision results are integrated into multi-scenario structured reports (simplified warning and response entry points for operation and maintenance, and statistical analysis conclusions for management and R&D). The semantic amendment examples of the annotations are extracted and incorporated into the model's joint optimization dataset, and rule call logs are recorded to support the updating of the track foreign object knowledge base.

[0175] In summary, the multimodal orbital foreign object detection method for large language models described in the above embodiments of the present invention innovatively achieves a collaborative detection architecture with three receptive field branches: fine, medium, and full. The first branch's deep local attention mechanism accurately captures the fine-grained features of small-sized foreign objects through deep convolution and local window attention, avoiding the detail dilution problem caused by global attention. The second branch's cross-block attention mechanism effectively establishes the correlation between medium-scale features. The third branch's global position encoding attention mechanism fully captures the global contextual information of the orbital environment. The differentiated design of these three attention mechanisms enables the network to simultaneously consider both the local detailed features and global contextual information of orbital foreign objects, significantly improving the detection performance for small-sized foreign objects and large-scale objects. The system achieves high detection accuracy for foreign objects. It utilizes a dynamic gating fusion module to adaptively weightedly fuse visible and infrared modes, effectively leveraging the complementary advantages of multimodal data under different environmental conditions. This ensures stable detection performance, especially in harsh environments where single modes fail, such as rain, fog, or nighttime. A lightweight channel dimensionality reduction network ensures efficient channel compression while maintaining feature information integrity, overcoming the information loss problem caused by traditional 1×1 convolutional dimensionality reduction. The entire network structure is designed with engineering deployment needs in mind, ensuring that the output features are fully compatible with standard detectors. It maintains high accuracy while meeting real-time processing requirements, and possesses intelligent judgment capabilities for foreign object attributes and threat levels. This provides a 24 / 7, highly reliable, and intelligent safety detection solution for rail transit systems.

[0176] Example 2

[0177] The multimodal orbital foreign object detection system for large language models in the second embodiment of the present invention includes:

[0178] The acquisition module is used to acquire multimodal data of the orbital scene, including visible light images and thermal imaging images. The multimodal data is input in parallel into a multi-branch fusion network to obtain multiple modal features. The multi-branch fusion network includes an improved YOLOv11 network, which includes a backbone network, a neck network, and a detection head.

[0179] The fusion module is used to adaptively weight and fuse multiple modal features obtained by the dynamic gating fusion module to output multiple fused features.

[0180] The splicing module is used to splice multiple fused features into a comprehensive feature. The comprehensive feature is then compressed through a lightweight channel dimensionality reduction network, and the final detection feature is obtained by combining a dual residual connection mechanism.

[0181] The detection module is used to output the track foreign object detection result through the neck network and the detection head to realize track foreign object detection. The detection result includes the location coordinates of the foreign object and the category confidence level.

[0182] In summary, the multimodal orbital foreign object detection system for large language models described in the above embodiments of the present invention innovatively achieves a collaborative detection architecture with three receptive field branches: fine, medium, and full. The first branch's deep local attention mechanism accurately captures the fine-grained features of small-sized foreign objects through deep convolution and local window attention, avoiding the detail dilution problem caused by global attention. The second branch's cross-block attention mechanism effectively establishes the correlation between medium-scale features. The third branch's global position encoding attention mechanism fully captures the global contextual information of the orbital environment. The differentiated design of these three attention mechanisms enables the network to simultaneously consider both the local detailed features and global contextual information of orbital foreign objects, significantly improving the detection of both small-sized and large-scale foreign objects. The system achieves high accuracy by employing a dynamic gating fusion module to adaptively weightedly fuse visible and infrared modes. This effectively leverages the complementary advantages of multimodal data under different environmental conditions, maintaining stable detection performance even in harsh environments such as rain, fog, and nighttime where single modes fail. The weighted attention dimensionality reduction mechanism of the channel recovery module ensures efficient channel compression while maintaining feature information integrity, overcoming the information loss problem caused by traditional 1×1 convolution direct dimensionality reduction. The entire network structure design fully considers engineering deployment requirements, ensuring that the output features are fully compatible with standard detectors. It maintains high accuracy while meeting real-time processing requirements, and possesses intelligent judgment capabilities for foreign object attributes and threat levels. This provides a 24 / 7, highly reliable, and intelligent safety detection solution for rail transit systems.

[0183] Furthermore, embodiments of the present invention also provide a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the methods described above.

[0184] Furthermore, embodiments of the present invention also propose a data processing device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the steps of the methods described above.

[0185] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (such as a computer-based system, a processor-including system, or other system that can fetch and execute instructions from, an instruction execution system, apparatus, or device). For the purposes of this specification, "computer-readable medium" can be any means that can contain, store, communicate, propagate, or transmit programs for use by, or in conjunction with, an instruction execution system, apparatus, or device.

[0186] More specific examples (a non-exhaustive list) of computer-readable media include: electrical connections (electronic devices) having one or more wires, portable computer disk drives (magnetic devices), random access memory (RAM), read-only memory (ROM), erasable and editable read-only memory (EPROM or flash memory), fiber optic devices, and portable optical disc read-only memory (CDROM). Furthermore, computer-readable media can even be paper or other suitable media on which programs can be printed, because programs can be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, interpreting, or otherwise processing as necessary, and then stored in computer memory.

[0187] It should be understood that various parts of the present invention can be implemented in hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented in software or firmware stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, it can be implemented using any one or a combination of the following techniques known in the art: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.

[0188] In the description of this specification, references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.

[0189] Although embodiments of the invention have been shown and described, those skilled in the art will understand that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the claims and their equivalents.

Claims

1. A multimodal orbital foreign object detection method for large language models, characterized in that, include: Multimodal data of the orbital scene is acquired, including visible light images and thermal imaging images. The multimodal data is input in parallel into a multi-branch fusion network to obtain multiple modal features. The multi-branch fusion network includes an improved YOLOv11 network, which includes a backbone network, a neck network, and a detection head. The obtained multiple modal features are adaptively weighted and fused using a dynamic gating fusion module to output multiple fused features. Specifically, for each branch's visible light and infrared blocks, the global average features are calculated along the spatial dimension to obtain feature vectors representing the overall information of each block. The global features of the two modalities are concatenated and input into a gating network consisting of two fully connected layers and a GELU activation function. Three dynamic weight coefficients are generated through the Softmax output layer. At the same time, cross-attention features between the two modal blocks are calculated to enhance the complementarity between modalities. Finally, the fused features are calculated according to the fusion formula to output the fused features. The three dynamic weight coefficients include α, β, and γ, where: α represents the fusion weight of the visible light modality, β represents the fusion weight of the infrared modality, and γ represents the fusion weight of the cross-attention features. The fusion formula is as follows: In the formula, fused_block is the final weighted fused feature tensor; Represents visible light mode segmentation; Indicates infrared mode segmentation; X out This represents the output of the cross-attention mechanism; Multiple fusion features are concatenated to obtain comprehensive features. The comprehensive features are then compressed using a lightweight channel dimensionality reduction network, and the final detection features are obtained by combining a dual residual connection mechanism. The final detection features are output through the neck network and the detection head to achieve track foreign object detection. The detection results include the location coordinates of the foreign object and the category confidence level. The backbone network includes fine receptive field branches, medium receptive field branches, and full receptive field branches. Each receptive field branch outputs a modal feature, wherein: The fine receptive field branch uses a 3×3 grid partitioning strategy to divide the feature map into 9 local regions, each with a size of H / 3×W / 3. It uses 3×3 convolutional kernels for local feature enhancement and employs a deep local attention mechanism to extract fine-grained features within a 3×3 sliding window. The middle receptive field branch uses a 2×2 grid block strategy to divide the input feature map into 4 medium regions, and in conjunction with the cross-block attention mechanism, it models the feature relationships within the blocks and the spatial relationships between blocks. The full receptive field branch uses a 1×1 grid block strategy to maintain the complete size of the feature map, and combines it with a global positional encoding attention mechanism to achieve contextual feature extraction across the entire map based on the introduction of learnable positional encoding; The cross-block attention mechanism includes a first attention sub-branch and a second attention sub-branch. The first attention sub-branch corresponds to local attention within a block, and the second attention sub-branch corresponds to cross-block attention. The calculation result of the first attention sub-branch is: In the formula, intra out This is the result of the calculation of the first attention sub-branch; yes Normalized weights The initial weights computed for the first attention sub-branch; V intra B represents the internal value; B represents the corresponding batch; num_blocks represents the number of blocks in the feature map; N represents the number of pixels per block; C represents the number of feature channels; where N is... × , b h Indicates the height of a single block; b w Indicates the width of a single block; Represents the characteristic tensor; The calculation result of the second attention sub-branch is: In the formula, inter_out is the result of the calculation of the second attention sub-branch; yes Normalized weights The initial weights calculated for the second attention sub-branch; V inter B represents the cross-domain value; B represents the corresponding batch; num_blocks represents the number of blocks in the feature map; C represents the number of feature channels.

2. The multimodal orbital foreign object detection method for large language models according to claim 1, characterized in that, The steps for concatenating multiple fused features to obtain a comprehensive feature include: Multiple fused features are concatenated to obtain a tensor with 3C channels. Then, a gating network is used to generate dynamic weights: first, the concatenated tensor is globally averaged along the H×W dimension, flattened, subjected to two linear transformations, GELU nonlinear activation, and Softmax normalization to obtain branch weights that sum to 1. Then, its dimension is expanded to adapt to feature weighting. Next, the three branches are weighted and fused to obtain the core weighted feature. This is repeated three times along the channel dimension and added to the residual of the original concatenated tensor to obtain the fused feature fused_feat. Then, a lightweight dimensionality reduction network is used to gradually compress the number of channels from 3C to the target dimension C through 1×1 convolution, GELU activation, and BatchNorm normalization to obtain the dimensionality reduction feature out. Finally, a residual layer is used to directly reduce the original concatenated tensor from 3C to C, and the residual is added to out to obtain the final output comprehensive feature.

3. The multimodal orbital foreign object detection method for large language models according to claim 1, characterized in that, The steps for achieving track foreign object detection, whereby the final detection features are output as track foreign object detection results via the neck network and the detection head, include: The fused features after channel recovery are input into the YOLO detection head. The precise location coordinates of the foreign object are obtained through the anchor point mechanism and bounding box regression. The category probability distribution of the foreign object is obtained through a multi-class classifier. Based on multi-scale feature information and the relative position of the track, the system outputs complete detection results including the coordinates of the foreign object bounding box, category label, confidence score, and threat level, providing a basis for decision-making on graded handling by rail transit operators.

4. The multimodal orbital foreign object detection method for large language models according to claim 1, characterized in that, In the process of acquiring multimodal data for a track scene, the methods for acquiring multimodal data include: The original images of the orbital scene are acquired, including visible light images and thermal imaging images. The original images are optimized using timestamp alignment and spatial registration techniques to ensure the spatiotemporal consistency of multimodal data. The original image after achieving spatiotemporal consistency is preprocessed, and the preprocessed original image is scale-aligned using a bilinear interpolation algorithm to obtain bimodal feature data with consistent channel number and spatial size. The preprocessing includes illumination normalization, contrast enhancement, and noise filtering.

5. A multimodal orbital foreign object detection system for large language models, characterized in that, The system includes: The acquisition module is used to acquire multimodal data of the orbital scene, including visible light images and thermal imaging images. The multimodal data is input in parallel into a multi-branch fusion network to obtain multiple modal features. The multi-branch fusion network includes an improved YOLOv11 network, which includes a backbone network, a neck network, and a detection head. The fusion module adaptively weights and fuses multiple modal features obtained through a dynamic gating fusion module to output multiple fused features. Specifically, it includes: calculating the global average features along the spatial dimension for each branch's visible light and infrared blocks to obtain feature vectors representing the overall information of each block; concatenating the global features of the two modalities and inputting them into a gating network consisting of two fully connected layers and a GELU activation function; generating three dynamic weight coefficients through a Softmax output layer; simultaneously calculating cross-attention features between the two modal blocks to enhance intermodal complementarity; and finally, calculating and outputting the fused features according to the fusion formula. The three dynamic weight coefficients include α, β, and γ, where α represents the fusion weight of the visible light modality, β represents the fusion weight of the infrared modality, and γ represents the fusion weight of the cross-attention features. The fusion formula is as follows: In the formula, fused_block is the final weighted fused feature tensor; Represents visible light mode segmentation; Indicates infrared mode segmentation; X out This represents the output of the cross-attention mechanism; The splicing module is used to splice multiple fused features into a comprehensive feature. The comprehensive feature is then compressed through a lightweight channel dimensionality reduction network, and the final detection feature is obtained by combining a dual residual connection mechanism. The detection module is used to output the track foreign object detection result through the neck network and the detection head to realize track foreign object detection. The detection result includes the foreign object location coordinates and category confidence. The backbone network includes fine receptive field branches, medium receptive field branches, and full receptive field branches. Each receptive field branch outputs a modal feature, wherein: The fine receptive field branch uses a 3×3 grid partitioning strategy to divide the feature map into 9 local regions, each with a size of H / 3×W / 3. It uses 3×3 convolutional kernels for local feature enhancement and employs a deep local attention mechanism to extract fine-grained features within a 3×3 sliding window. The middle receptive field branch uses a 2×2 grid block strategy to divide the input feature map into 4 medium regions, and in conjunction with the cross-block attention mechanism, it models the feature relationships within the blocks and the spatial relationships between blocks. The full receptive field branch uses a 1×1 grid block strategy to maintain the complete size of the feature map, and combines it with a global positional encoding attention mechanism to achieve contextual feature extraction across the entire map based on the introduction of learnable positional encoding; The cross-block attention mechanism includes a first attention sub-branch and a second attention sub-branch. The first attention sub-branch corresponds to local attention within a block, and the second attention sub-branch corresponds to cross-block attention. The calculation result of the first attention sub-branch is: In the formula, intra out This is the result of the calculation of the first attention sub-branch; yes Normalized weights The initial weights computed for the first attention sub-branch; V intra B represents the internal value; B represents the corresponding batch; num_blocks represents the number of blocks in the feature map; N represents the number of pixels per block; C represents the number of feature channels; where N is... × , b h Indicates the height of a single block; b w Indicates the width of a single block; Represents the characteristic tensor; The calculation result of the second attention sub-branch is: In the formula, inter_out is the result of the calculation of the second attention sub-branch; yes Normalized weights The initial weights calculated for the second attention sub-branch; V inter B represents the cross-domain value; B represents the corresponding batch; num_blocks represents the number of blocks in the feature map; C represents the number of feature channels.

6. A computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by the processor, the program implements the multimodal orbital foreign object detection method for large language models as described in any one of claims 1-4.

7. A data processing apparatus, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the multimodal orbital foreign object detection method for large language models as described in any one of claims 1-4.

Citation Information

Patent Citations

  • Method and system for detecting nonferrous metal target of scraped car

    CN120953758A

  • Power transmission line foreign matter detection method and system based on multi-modal image fusion

    CN121599964A