Multi-mode-based track foreign matter invasion event capturing system and method

Through multimodal feature extraction and fusion, combined with image and text information, the problem of inaccurately distinguishing the types of changes in the capture of orbital foreign object invasion events in the prior art is solved, and higher accuracy and interpretability are achieved, and false positive rates are reduced.

CN120339992AActive Publication Date: 2025-07-18SHANDONG ZHIYANG HUITONG DIGITAL TECH CO LTD

Patent Information

Application Number
CN202510384075.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-28
Publication Date
2025-07-18
Estimated Expiration
2045-03-28

AI Technical Summary

Technical Problem

The prior art cannot accurately distinguish the types of changes in the capture of orbital foreign object invasion events, resulting in a high false positive rate and lack of semantic reasoning ability and intuitive explanation.

Method used

Multimodal feature extraction and fusion method is used to combine image and text information, and multi-scale visual features and text features are extracted through residual networks and pre-trained CLIP Transformer, multi-stage semantic inference module is used to judge the change type, and intuitive visual masks and text answers are generated.

Benefits of technology

It improves the accuracy of the judgment of the type of change and the interpretability of the system, reduces the false positive rate, provides intuitive visual and text feedback, and enhances the reliability and practicality of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339992A_ABST
    Figure CN120339992A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of track safety monitoring, and more particularly relates to a multi-mode-based track foreign matter intrusion event capturing system and a multi-mode-based track foreign matter intrusion event capturing method. The method comprises the following steps: regularly acquiring an image sequence of a track line; performing change detection on each pair of adjacent images in the image sequence, and identifying a change area; carrying out multi-modal feature extraction on a change region in the image; performing change type judgment by using the image features and the text features extracted in a multi-modal manner; then generating a text answer and a corresponding visual mask, and further verifying a change type; and the event capturing module determines whether a foreign matter invasion event exists according to the change type, and if yes, an alarm is triggered and event information is recorded. The problem that change types cannot be accurately distinguished when foreign matter invasion is captured in the prior art is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of track safety monitoring, and more specifically, relates to a multi-modal based track foreign object intrusion event capturing system and method. Background Art

[0002] In the track, timely and accurately capturing foreign object intrusion events is crucial for ensuring the safe operation of trains. Existing technologies mainly achieve the capture of foreign object intrusion events by comparing the changing areas between two frames of images. However, this method has certain limitations and cannot accurately distinguish the types of changes. For example, it cannot determine whether the change is caused by a newly emerged foreign object, the disappearance of a foreign object, or a change in the foreign object type, resulting in a relatively high false alarm rate.

[0003] Chinese patent document CN118759590A discloses an airport runway foreign object detection and removal system and method. The system includes: detectors, which are distributed on the airport runway and used to obtain detection data of the airport runway; a data processing terminal, which is respectively signal-connected to the detectors and the removal device, and is used to determine the foreign object position and foreign object type based on the detection data, and generate a foreign object removal instruction according to the foreign object position, foreign object type, and the current position of the removal device; the removal device is used to remove the foreign object based on the foreign object removal instruction; by obtaining the detection data of the airport runway through the detectors and the data processing terminal performing data processing on the detection data to determine the foreign object position, the automatic detection of foreign objects on the airport runway is realized; the data processing terminal also generates a foreign object removal instruction according to the foreign object position, foreign object type, and the current position of the removal device to guide the removal device to remove the foreign object. However, this patent has the following deficiencies:

[0004] 1. Lack of semantic reasoning ability

[0005] The detection method of CN118759590A is mainly based on signal processing and simple image analysis, lacking semantic understanding of the detection results. When its system determines the foreign object type, it mainly relies on the analysis of image signals and cannot perform deeper semantic reasoning on the change types.

[0006] 2. Lack of intuitive explanation and interpretability

[0007] The detection method of CN118759590A mainly outputs the position and type of the foreign object, lacking an intuitive explanation of the detection results. Its system cannot provide a detailed semantic description or visual mask of the changing area, and it is difficult for operators to quickly understand the detection results.

[0008] 3. Lack of comprehensive judgment on change types

[0009] The detection method of CN118759590A mainly focuses on the presence and location of foreign objects, lacking a comprehensive judgment of the change types. For example, it cannot accurately distinguish whether a foreign object has newly appeared, disappeared, or its type has changed.

[0010] In view of this, the present invention designs a multi-modal based rail foreign object intrusion event capture system and method. Summary of the Invention

[0011] The present invention aims to overcome at least one defect of the above-mentioned prior art and provides a multi-modal based rail foreign object intrusion event capture system.

[0012] The present invention also discloses a multi-modal based rail foreign object intrusion event capture method to solve the problem that the prior art cannot accurately distinguish the change types when capturing foreign object intrusion.

[0013] The detailed technical solution of the present invention is as follows:

[0014] A multi-modal based rail foreign object intrusion event capture system, the system includes: an image acquisition module, a change detection module, a multi-modal feature extraction module, a change type judgment module, and an event capture module;

[0015] The image acquisition module is used to collect an image sequence of a rail line and construct a detection data set for foreign object intrusion in the line patrol scenario; the detection data set includes image pairs, descriptive question texts, text answers, and corresponding visual masks, which are used to support multi-modal feature extraction and change type judgment;

[0016] The change detection module is used to perform change detection on two adjacent frames of images in the image sequence and mark the changed areas;

[0017] The multi-modal feature extraction module is used to extract image features and text features of the changed areas;

[0018] The change type judgment module is used to judge the change type based on multi-modal features and distinguish newly appeared foreign objects, disappeared foreign objects, or changed foreign object types;

[0019] The event capture module is used to determine whether it is a foreign object intrusion event according to the change type and perform corresponding alarms or records.

[0020] Preferably according to the present invention, the multi-modal feature extraction module includes an image feature extraction unit and a text feature extraction unit;

[0021] The image feature extraction unit is used to extract image features from the changed regions, including a residual network and a 1×1 convolutional layer. In addition, the image feature extraction unit further includes a normalization layer, a pooling layer, an activation function, and a feature fusion module, which are used to extract multi-scale features and enhance the expression ability of the model;

[0022] The text feature extraction unit is used to generate text questions describing the changed regions and extract text features, including a question generation module, a text encoder, a feature extraction module, a feature alignment module, and a feature enhancement module;

[0023] The question generation module: is used to preset question templates and automatically generate descriptive questions according to the image content of the changed regions;

[0024] The text encoder: is used for preprocessing text input, performing word segmentation on the text describing the changed regions, and converting it into an embedding vector;

[0025] The feature extraction module: extracts semantic features in the text based on the pre-trained model CLIP Transformer for contrast and image pairs;

[0026] The feature alignment module: is used to align text features with image features to ensure their consistency in space and semantics;

[0027] The feature enhancement module: is used to further process the extracted text features and enhance the expression ability of key information.

[0028] Preferably according to the present invention, the change type judgment module includes:

[0029] A multi-stage semantic reasoning unit, which is used to combine image features and text features and judge the change type through multi-stage semantic reasoning;

[0030] A text-visual answer decoding unit, which is used to generate text answers and corresponding visual masks to further verify the change type.

[0031] A multi-modal method for capturing rail foreign object intrusion events, the method includes:

[0032] S1. Image acquisition: Regularly acquire image sequences of the rail line using a high-resolution camera to ensure that the images can clearly reflect the rail and its surrounding environment. The acquired image sequences include panoramas and local close-ups of the rail line to cover different ranges of changes.

[0033] Detection dataset construction: Label and process the acquired image sequences to generate a multi-modal detection dataset for model training and verification; the labeling content includes pixel-level visual masks of the changed regions, as well as question texts and corresponding text answers generated for the changed regions;

[0034] S2. Change Detection: Perform change detection on each pair of adjacent images in the image sequence to identify the changed regions;

[0035] S3. Extract multi-modal features from the changed regions in the image. First, use the image feature extraction unit to extract the image features of the changed regions; then use the text feature extraction unit to generate text questions describing the changed regions and extract the text features;

[0036] S4. Use the image features and text features extracted by multi-modal to judge the change type: First, combine the image features and text features, generate a multi-modal feature representation through a multi-stage semantic reasoning module, judge the change type according to the multi-modal feature representation, and then generate a text answer and the corresponding visual mask to further verify the change type;

[0037] S5. The event capture module determines whether it is a foreign object intrusion event according to the change type. If it is a foreign object intrusion event, trigger an alarm and record the event information.

[0038] Preferably according to the present invention, the specific steps of the change detection in step S2 are as follows:

[0039] S2.1. Image Preprocessing: Perform operations such as grayscale conversion, normalization, and denoising on the collected image sequence to improve the contrast and clarity of the image.

[0040] S2.2. Change Detection Algorithm: First, perform a difference operation on the currently acquired image frame and the background image to obtain a grayscale image of the moving region in the image, perform thresholding on the grayscale image to extract the moving region, and then use the inter-frame difference method based on edges to obtain the changed region of the moving object;

[0041] S2.3. Changed Region Marking: Mark the changed regions according to the change detection results as the regions of interest for subsequent steps.

[0042] Preferably according to the present invention, the specific meaning of using the image feature extraction unit to extract the image features of the changed regions is:

[0043] S3.1.1. Multi-scale Feature Extraction: Use a residual network with shared weights as the basic feature extraction network, and layer by layer extract the multi-scale feature maps of the image through multiple residual blocks;

[0044] S3.1.2. Channel Dimension Adjustment: Adjust the channel dimension of the multi-scale feature maps output by the residual network through a 1×1 convolutional layer to optimize the feature expression ability;

[0045] S3.1.3. Feature Fusion and Output: Fuse the multi-scale feature maps after adjusting the channel dimension to generate the final visual feature representation for subsequent processing.

[0046] Preferably according to the present invention, the specific processing process of using the text feature extraction unit to generate a text problem describing the changed area and extract text features is as follows:

[0047] S3.2.1. The problem generation module presets a problem template and automatically generates a descriptive problem text according to the image content of the changed area;

[0048] S3.2.2. Input the generated descriptive problem text into a text encoder, and the text encoder performs word segmentation on the descriptive problem text, converting each word or subword Subword in the text into a corresponding embedding vector.

[0049] S3.2.3. Text feature extraction:

[0050] Input the text processed by the text encoder into a pre-trained CLIP Transformer. The CLIP Transformer encodes the text through a multi-layer Transformer architecture to extract the semantic features of the text; each layer of the Transformer captures the long-range dependencies in the text through the self-attention mechanism and further processes the features through a feed-forward network; finally, the CLIP Transformer outputs a high-dimensional feature representation of the text;

[0051] S3.2.4. The feature alignment module aligns the high-dimensional feature representation of the text with the image features to ensure the consistency of the two in space and semantics;

[0052] S3.2.5. The feature enhancement module further processes the extracted text features and enhances the expression ability of key information through the attention mechanism.

[0053] Preferably according to the present invention, in step S4, combining visual features and text features, generating a multi-modal feature representation through a multi-stage semantic reasoning module, and judging the change type according to the multi-modal feature representation specifically includes the following steps:

[0054] S4.1. The multi-stage semantic reasoning module receives the image features and text features from the multi-modal feature extraction module. The image features include visual information such as the color, texture, and shape of the changed area, while the text features are the text problems describing the changed area and their corresponding feature representations;

[0055] S4.2. Fuse the image features and text features through a convolutional layer and a fully connected layer to generate a multi-modal feature representation;

[0056] S4.3. Use the multi-head self-attention mechanism MHSA to perform self-attention calculation on the multi-modal feature representation to enhance the expression ability of the features, specifically as follows:

[0057] MHSA(Q, K, V) = Concat(head1, head2, head3,..., head h )W O (1)

[0058] Among them,

[0059] head i = Attention(QW iQ , KW iK , VW iv )(2)

[0060]

[0061] In formulas (1)-(3), MHSA represents the multi-head self-attention mechanism; Q represents the query matrix, indicating the query relationship of each element in the input sequence with respect to other elements; K represents the key matrix, indicating the key information of each element in the input sequence, which is used to match with the query matrix; V represents the value matrix, indicating the value information of each element in the input sequence, which is used to perform weighted summation according to the matching results of the query and the key; W O represents the output linear transformation matrix, which is used to map the concatenated outputs of all heads back to the original dimension; head i represents the self-attention output of the i-th head; W iQ , W iK and W iV respectively represent the linear transformation matrices of the query, key, and value corresponding to the i-th head, which are used to map the input sequence to different subspaces for self-attention calculation; d k represents the dimension of the key matrix K, which is used to scale the dot product result to prevent the gradient of the softmax function from vanishing due to an overly large dot product result;

[0062] S4.4. Use the cross-attention mechanism CA to further enhance the interaction between image features and text features. The input is two sequences of different modalities, and the dimensions of the two sequences must be the same. One sequence serves as the input Q, which defines the output sequence length, and the other sequence provides the input K and V. Specifically as follows:

[0063] CA(Q, K, V) = softmax((W Q S2) * (W K S1) T )) * W V S1(4)

[0064] In formula (4), S1 represents the feature sequence of the first modality; S2 represents the feature sequence of the second modality; W Q , W K and W VThey are the linear transformation matrices of query, key, and value respectively, used to map the feature sequences of different modalities to the query, key, and value spaces of the cross-attention mechanism; Q represents the query matrix, which is calculated by W Q S2, indicating the query relationship of the feature sequence of the second modality to the query of the first modality; K represents the key matrix, which is calculated by W K S1, indicating the key information of the feature sequence of the first modality; V represents the value matrix, which is calculated by W V S1, indicating the value information of the feature sequence of the first modality.

[0065] S4.5. Further process the multi-modal features processed by the cross-attention mechanism using the feed-forward network FFN to enhance the non-linear expression ability of the features, specifically as follows:

[0066] FFN(x) = max(0, xW1 + b1)W2 + b2 (5)

[0067] In formula (5), x represents the multi-modal features input to the feed-forward network; W1 represents the weight matrix of the first layer of the feed-forward network, used to perform a linear transformation on the input features; b1 represents the bias vector of the first layer of the feed-forward network, used to increase the flexibility of the model; max(0, ·) represents the ReLU activation function, used to introduce non-linearity, set all negative values to 0, and retain positive values; W2 represents the weight matrix of the second layer of the feed-forward network, used to perform a further linear transformation on the features after ReLU activation; b2 represents the bias vector of the second layer of the feed-forward network, used for the flexibility of the second layer of the model;

[0068] S4.6. The multi-modal features processed by the feed-forward network pass through a feature selection module to dynamically select the feature representation most relevant to the current problem, specifically as follows:

[0069] α = σ(F vl *F w ), β = 1 - α, F s = αF vl + βF w (6)

[0070] In formula (6), σ is the sigmoid function, used to map the input value to the interval (0, 1) as the weight coefficient; F vl is the multi-modal feature representation processed by the feed-forward network; F w is the text feature; α is the weight coefficient calculated by the sigmoid function, indicating the importance of the visual feature in the current problem; β is the complement of α, indicating the importance of the text feature in the current problem; F s represents the finally selected multi-modal feature representation, which is the weighted sum of the visual feature and the text feature, used for subsequent processing;

[0071] S4.7. Based on the generated multi-modal feature representation, use a classifier to determine the type of change, specifically as follows:

[0072] Type of change = argmax(softmax(F s W + b)) (7)

[0073] In formula (7), F s is the multi-modal feature representation processed by the multi-stage semantic reasoning module, and W and b are the weights and biases of the classifier.

[0074] According to the preference of the present invention, further verifying the type of change for the generated text answer and the corresponding visual mask means that, first, process the multi-modal features through the multi-stage semantic reasoning module to obtain F s , then, apply a linear transformation layer to map F s to a new space to obtain the weight matrix W and the bias vector b. Next, perform segmentation and reshaping operations on W and b to adapt to the subsequent visual mask generation. Finally, multiply the processed weight matrix W by the feature representation processed by the Mask Decoder, and add the bias vector b to generate the final visual mask M. The formula is as follows:

[0075] W, b = S&R(Linear(F s )) (8)

[0076]

[0077] In formulas (8) and (9), S&R represents the segmentation and reshaping operation, which is used to adjust the shapes of the weight matrix and the bias vector for calculation with the subsequent feature representation. Linear is the linear transformation, F s is the multi-modal feature representation processed by the multi-stage semantic reasoning module, is the feature representation processed by the Mask Decoder, M is the final visual mask, W represents the weight matrix after linear transformation for generating the visual mask, and b represents the bias vector after linear transformation for adjusting the generation of the visual mask.

[0078] The Mask Decoder processing refers to inputting the rough visual mask M c and the original visual feature F v into the MaskDecoder. The Mask Decoder uses two consecutive bidirectional attention blocks to establish the mapping relationship between M c and F v at the pixel level, enabling the model to more precisely understand which visual features are related to the question-answer feature F sThe most relevant one is expressed by the following formula:

[0079]

[0080] In formula (10), M c represents a rough visual mask: This is a preliminary visual feature map used to indicate the area that may contain the target, which is obtained by pixel decoding of F vl through two convolutional layers; F v represents the original visual feature: This is the original feature directly extracted from the image, containing rich visual information.

[0081] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0082] (1) The present invention provides multi-modal feature extraction and fusion: It not only detects changes, but also combines natural language questions to generate text answers and visual masks, providing more intuitive explanations, combining image and text information, and improving the accuracy of change type judgment. The multi-modal feature extraction module is introduced, which not only extracts the deep features of the image, but also combines text information to generate text features describing the changed area. The ResNet network with shared weights is used to extract multi-scale visual features of two images, and the channel dimension is adjusted through a 1×1 convolutional layer to generate multi-scale change features; the pre-trained CLIP Transformer is used to extract the text features of the question and generate sentence-level representations. Through the fusion of multi-modal features, the system can comprehensively understand the changed area from both visual and semantic levels, thereby more accurately judging the change type. For example, for a new foreign object appearing in the image, the system can not only identify its shape and color through image features, but also understand its semantic information through text features, thereby improving the accuracy and reliability of the judgment.

[0083] (2) The present invention has multi-stage semantic reasoning: Through the multi-stage semantic reasoning module, the judgment of the change type is gradually refined to reduce false alarms. This module gradually refines and verifies the interaction between image and text features through the multi-head self-attention mechanism, cross-attention mechanism and feed-forward network, enhancing the expressive ability of the features. At the same time, the feature selection module dynamically selects the feature representation most relevant to the current question to further improve the accuracy of reasoning. For example, when judging whether the changed area is the disappearance of a foreign object, the multi-stage semantic reasoning module can comprehensively consider the pixel changes in this area of the image and the semantic interpretation of this area in the text description, so as to make a more accurate judgment.

[0084] (3) Text-Visual Answer Decoding: Generate intuitive visual masks to verify the correctness of text answers and enhance the reliability of the system. Generate intuitive visual masks and corresponding text answers to further verify the type of change. This module not only provides more intuitive visual feedback for the system but also enhances the interpretability of the system. For example, when the system determines that a certain changed area is the appearance of a new foreign object, the Text-Visual Answer Decoding module can generate a visual mask annotating the area of the new foreign object and output the text answer "Yes, a new foreign object has appeared", enabling the operator to quickly and intuitively understand the change situation and improving the practicality and credibility of the system. Brief Description of the Drawings

[0085] Figure 1 It is the system architecture diagram of the present invention.

[0086] Figure 2 It is the method flow chart of the present invention. Detailed Embodiments

[0087] The following further describes the present disclosure in conjunction with the drawings and embodiments.

[0088] Embodiment 1

[0089] Refer Figure 1 , this embodiment provides a multi-modal rail foreign object intrusion event capture system, including:

[0090] An image acquisition module, configured to acquire an image sequence of a rail line; construct a foreign object intrusion detection data set for the line patrol scenario;

[0091] This detection data set includes image pairs, problem text descriptions, text answers, and corresponding visual masks, and is used to support multi-modal feature extraction and change type judgment;

[0092] Image acquisition: Regularly acquire an image sequence of a rail line using a high-resolution camera to ensure that the images can clearly reflect the rail and its surrounding environment. The acquired images should include the panorama and local close-ups of the rail line to cover different ranges of changes;

[0093] Data set construction: Label and process the acquired image sequence to generate a multi-modal data set for model training and verification. The labeling content includes pixel-level masks of the changed areas, as well as natural language questions and corresponding answers generated for the changed areas.

[0094] A change detection module: Perform change detection on two adjacent frames of images in the image sequence and mark the changed areas;

[0095] The multi-modal feature extraction module is used to extract image features and text features of the changed area. By introducing the multi-modal feature extraction module, not only the depth features of the image are extracted, but also the text information is combined to generate text features describing the changed area. The ResNet network with shared weights is used to extract multi-scale visual features of two images, and the channel dimension is adjusted through a 1×1 convolutional layer to generate multi-scale change features; the pre-trained CLIP Transformer is used to extract the text features of the question and generate sentence-level representations;

[0096] The multi-modal feature extraction module includes an image feature extraction unit and a text feature extraction unit:

[0097] The image feature extraction unit is used to extract visual features of the changed area, including the Residual Network (ResNet) and a 1×1 convolutional layer. In addition, this unit also includes a normalization layer, a pooling layer, a ReLU activation function, and a feature fusion module, which are used to extract multi-scale features and enhance the expression ability of the model;

[0098] The text feature extraction unit is used to generate text questions describing the changed area and extract text features, including a Question Generation Module: used to preset question templates and automatically generate descriptive questions according to the image content of the changed area, such as "Has a new foreign object appeared in this area?";

[0099] A Text Encoder: used for preprocessing text input, tokenizing the text describing the changed area, and converting it into an embedding vector;

[0100] A feature extraction module: based on the pre-trained model CLIP Transformer model of contrast and image pairs to extract semantic features in the text;

[0101] A Feature Alignment Module: used to align text features with image features to ensure their consistency in space and semantics;

[0102] A Feature Enhancement Module: used to further process the extracted text features, such as enhancing the expression ability of key information through an attention mechanism.

[0103] The change type judgment module is used to judge the change type based on multi-modal features, distinguishing newly appeared foreign objects, disappeared foreign objects, or changed foreign object types;

[0104] The change type judgment module includes:

[0105] A multi-stage semantic reasoning unit for combining image features and text features to judge the type of change through multi-stage semantic reasoning;

[0106] A text-visual answer decoding unit for generating text answers and corresponding visual masks to further verify the type of change;

[0107] An event capture module for determining whether it is a foreign object intrusion event according to the type of change and performing corresponding alarm or recording.

[0108] Example 2

[0109] Refer Figure 2 , this embodiment provides a multi-modal method for capturing rail foreign object intrusion events, including the following steps:

[0110] S1. Image acquisition: Regularly acquire an image sequence of the rail line;

[0111] S2. Change detection: Perform change detection on each pair of adjacent images in the image sequence to identify the changed area. The specific steps are as follows:

[0112] S21. Image preprocessing: Perform operations such as grayscale conversion, normalization, and denoising on the acquired image sequence to improve the contrast and clarity of the image.

[0113] S22. Change detection algorithm: First, perform a difference operation on the currently acquired image frame and the background image to obtain a grayscale image of the moving area in the image. Threshold the grayscale image to extract the moving area, and then use an edge-based inter-frame difference method to obtain the changed area of the moving target;

[0114] S23. Changed area marking: Mark the changed area according to the change detection result as the region of interest for subsequent steps.

[0115] S3. Extract multi-modal features from the changed area in the image:

[0116] First, use an image feature extraction unit to extract image features of the changed area, such as color, texture, shape, etc., as follows:

[0117] S3.1.1. Multi-scale feature extraction: Use a residual network ResNet with shared weights as the basic feature extraction network to gradually extract multi-scale feature maps of the image through multiple residual blocks;

[0118] S3.1.2. Channel dimension adjustment: Adjust the channel dimension of the multi-scale feature map output by the residual network through a 1×1 convolutional layer to optimize the feature expression ability;

[0119] S3.1.3, Feature Fusion and Output: Fuse the multi-scale feature maps after adjusting the channel dimensions to generate the final visual feature representation for subsequent processing;

[0120] Then, use the text feature extraction unit to generate text questions describing the changed region and extract text features, as follows:

[0121] S3.2.1, The Question Generation Module presets question templates and automatically generates descriptive question texts based on the image content of the changed region. The question texts include but are not limited to the following forms:

[0122] "Has a new foreign object appeared in this region?"

[0123] "Has the foreign object in this region disappeared?"

[0124] "Has the type of foreign object in this region changed?";

[0125] S3.2.2, Input the generated descriptive question text into the Text Encoder. The Text Encoder tokenizes the descriptive question text and converts each word or subword in the text into the corresponding Embedding Vector.

[0126] S3.2.3, Text Feature Extraction:

[0127] Input the text processed by the Text Encoder into the pre-trained CLIP Transformer. The CLIP Transformer encodes the text through a multi-layer Transformer architecture to extract the semantic features of the text; each layer of the Transformer captures the long-range dependencies in the text through the self-attention mechanism and further processes the features through the feed-forward network; finally, the CLIP Transformer outputs the high-dimensional feature representation of the text;

[0128] S3.2.4, The Feature Alignment Module aligns the high-dimensional feature representation of the text with the image features to ensure their spatial and semantic consistency;

[0129] S3.2.5, The Feature Enhancement Module further processes the extracted text features, such as enhancing the expression ability of key information through the attention mechanism.

[0130] S4. Use the image features and text features extracted by multi-modal to judge the change type: First, combine the image features and text features, and judge the change type through a multi-stage semantic reasoning module; then generate a text answer and a corresponding visual mask to further verify the change type, as follows:

[0131] 1. Input feature extraction

[0132] The multi-stage semantic reasoning module receives the image features and text features from the multi-modal feature extraction module. The image features include visual information such as the color, texture, and shape of the change area, while the text features are the text questions describing the change area and their corresponding feature representations.

[0133] 2. Multi-stage reasoning process

[0134] The multi-stage semantic reasoning module gradually refines and verifies the change type through multiple stages. Each stage includes the following key steps:

[0135] 2.1 Feature fusion

[0136] Fuse the image features and text features to generate a multi-modal feature representation. The specific method is as follows:

[0137] Image features: The visual features of the change area obtained from the change detection module.

[0138] Text features: The text features describing the change area obtained from the text feature extraction unit.

[0139] Fusion method: Fuse the image features and text features through convolutional layers and fully connected layers to generate a multi-modal feature representation.

[0140] 2.2 Multi-Head Self-Attention mechanism

[0141] Use the Multi-Head Self-Attention (MHSA) mechanism to perform self-attention calculation on the multi-modal features to enhance the expression ability of the features. The specific steps are as follows:

[0142] Input: Multi-modal feature representation.

[0143] Output: The feature representation processed by the self-attention mechanism.

[0144] Formula:

[0145] MHSA(Q, K, V) = Concat(head1, head2, head3,..., head h )W O

[0146] Among them,

[0147] head i = Attention(QW iQ , QW iQ , QW iQ ),

[0148]

[0149] Among them, MHSA(Q, K, V) represents the multi-head self-attention mechanism Multi-Head Self-Attention; Q represents the query matrix, indicating the query relationship of each element in the input sequence with respect to other elements; K represents the key matrix, indicating the key information of each element in the input sequence, which is used to match with the query matrix; V represents the value matrix, indicating the value information of each element in the input sequence, which is used to perform weighted summation according to the matching results of the query and the key; W O represents the output linear transformation matrix, which is used to splice the outputs of all heads and then map them back to the original dimension; head i represents the self-attention output of the i-th head; W iQ , W iK and W iV respectively represent the linear transformation matrices of the query, key, and value corresponding to the i-th head, which are used to map the input sequence to different subspaces for self-attention calculation; d k represents the dimension of the key matrix K, which is used to scale the dot product result to prevent the gradient of the softmax function from vanishing due to an overly large dot product result.

[0150] 2.3 Cross-Attention Mechanism

[0151] The cross-attention mechanism (Cross-Attention, CA) is used to further enhance the interaction between image features and text features. The specific steps are as follows:

[0152] Input: Image features and text features processed by the self-attention mechanism. The input is two sequences of different modalities, and the dimensions of the two sequences must be the same. One sequence serves as the input Q, which defines the output sequence length, and the other sequence provides the input K and V.

[0153] Output: Multimodal feature representations processed by the cross-attention mechanism.

[0154] Formula:

[0155] CA(Q, K, V) = softmax((W Q S2)*(W K S1) T ))*W V S1

[0156] Among them, S1 represents the feature sequence of the first modality (such as text); S2 represents the feature sequence of the second modality (such as image); W Q 、W K and W V are the linear transformation matrices of query, key, and value respectively, used to map the feature sequences of different modalities to the query, key, and value spaces of the cross-attention mechanism; Q represents the query matrix, which is calculated by W Q S2, representing the query relationship of the feature sequence of the second modality with respect to the query of the first modality; K represents the key matrix, which is calculated by W K S1, representing the key information of the feature sequence of the first modality; V represents the value matrix, which is calculated by W V S1, representing the value information of the feature sequence of the first modality.

[0157] 2.4 Feed-Forward Network

[0158] The feed-forward network (FFN) is used to further process the multi-modal features to enhance the non-linear expression ability of the features. The specific steps are as follows:

[0159] Input: The multi-modal feature representation processed by the cross-attention mechanism.

[0160] Output: The multi-modal feature representation processed by the feed-forward network.

[0161] Formula:

[0162] FFN(x) = max(0, xW1 + b1)W2 + b2

[0163] Among them, x represents the multi-modal features input to the feed-forward network; W1 represents the weight matrix of the first layer of the feed-forward network, used for linear transformation of the input features; b1 represents the bias vector of the first layer of the feed-forward network, used to increase the flexibility of the model; max(0, ·) represents the ReLU activation function, used to introduce non-linearity, setting all negative values to 0 and retaining positive values; W2 represents the weight matrix of the second layer of the feed-forward network, used for further linear transformation of the features after ReLU activation; b2 represents the bias vector of the second layer of the feed-forward network, used for the flexibility of the second layer of the model.

[0164] 2.5 Feature Selection

[0165] Through a selection module, the features most relevant to the current problem are dynamically selected. The specific steps are as follows:

[0166] Input: The multi-modal feature representation processed by the feed-forward network.

[0167] Output: The selected feature representation.

[0168] Formula:

[0169] α = σ(F vl * F w ), β = 1 - α, F s = αF vl + βF w

[0170] where σ is the sigmoid function, which is used to map the input value to the interval (0, 1) as the weight coefficient; F vl is the multi-modal feature representation processed by the feed-forward network; F w is the text feature; α is the weight coefficient calculated by the sigmoid function, representing the importance of the visual feature in the current problem; β is the complement of α, representing the importance of the text feature in the current problem; F s represents the finally selected multi-modal feature representation, which is the weighted sum of the visual feature and the text feature and is used for subsequent processing.

[0171] 3. Output Feature Representation

[0172] After being processed by the multi-stage semantic reasoning module, the generated multi-modal feature representation will be used for subsequent change type judgment and text-visual answer decoding module.

[0173] 4. Change Type Judgment

[0174] Based on the generated multi-modal feature representation, use a classifier to judge the change type. The specific steps are as follows:

[0175] Input: Multi-modal feature representation.

[0176] Output: Change type, including changed, newly emerged foreign object, foreign object disappeared, foreign object type changed.

[0177] Formula:

[0178] Change type = argmax(softmax(F s W + b))

[0179] where F s is the multi-modal feature representation processed by the multi-stage semantic reasoning module, and W and b are the weights and biases of the classifier.

[0180] 5. Text-Visual Answer Decoding

[0181] Generate the visual mask corresponding to the text answer to further verify the change type. The specific steps are as follows:

[0182] Input: Multi-modal feature representation.

[0183] Output: The visual mask corresponding to the text answer.

[0184] Formula:

[0185] W, b = S&R(Linear(F s ))

[0186]

[0187] Among them, S&R represents the Split and Reshape operations, which are used to adjust the shapes of the weight matrix and the bias vector for calculation with subsequent feature representations. Linear is a linear transformation, and F s is the multi-modal feature representation processed by the multi-stage semantic reasoning module. is the feature representation processed by the Mask Decoder. M is the final visual mask. W represents the weight matrix after linear transformation, which is used to generate the visual mask. b represents the bias vector after linear transformation, which is used to adjust the generation of the visual mask.

[0188] The processing process of the Mask Decoder is as follows:

[0189]

[0190] Input feature integration:

[0191] Coarse visual mask M c : This is a preliminary visual feature map used to indicate the regions that may contain the target and is obtained by pixel decoding of F vl through two convolutional layers.

[0192] Original visual feature F v : This is the original feature directly extracted from the image and contains rich visual information.

[0193] Feature processing: Input M c and F v into the Mask Decoder. The Mask Decoder contains two consecutive bidirectional attention blocks, which are designed to establish the mapping relationship between M c and F v at the pixel level. This mapping helps the model more precisely understand which visual features are most relevant to the question-answer feature F s .

[0194] S5. The event capture module determines whether it is a foreign object intrusion event according to the change type. If it is a foreign object intrusion event, it triggers an alarm and records the event information.

[0195] Embodiment 3

[0196] This embodiment provides a method for constructing a dataset for track line patrol scenarios:

[0197] 1. Dataset construction objective:

[0198] Construct a multi-modal change detection dataset for track line patrol scenarios to support the capture and analysis of track foreign object intrusion events. This dataset will include image pairs, descriptive question texts, text answers, and corresponding visual masks to support multi-modal feature extraction and change type judgment.

[0199] 2. Dataset composition:

[0200] The dataset will include the following:

[0201] Image pairs: Each pair of images represents the state of the track line at different time points for detecting changes.

[0202] Question text description: Generate natural language questions for the changed areas in the image pairs, such as "Has there been a change in the second picture compared to the first picture?" "Is there a new foreign object in the changed area of the second picture?" etc.

[0203] Text answers: Accurate answers to the question text descriptions, such as "Yes", "No", "New foreign object appears", "Foreign object disappears", etc.

[0204] Visual masks: Pixel-level masks marking the changed areas for visually showing the change locations.

[0205] 3. Data collection:

[0206] Image collection:

[0207] Regularly collect image sequences of the track line using high-resolution cameras to ensure that the images can clearly reflect the track and its surrounding environment.

[0208] The image collection frequency is determined according to the actual operation conditions and safety requirements of the track line. For example, images are collected once per minute or once per hour.

[0209] The collected images should include panoramas and local close-ups of the track line to cover different ranges of changes.

[0210] Annotation tools:

[0211] Use professional image annotation tools, such as Labelme or custom annotation software, to annotate the collected images.

[0212] The annotation content includes the bounding boxes and pixel-level masks of the changed areas. The annotators should have professional knowledge of track line patrol to ensure the accuracy of the annotation.

[0213] 4. Descriptive Question Text Generation:

[0214] Question Template Design:

[0215] According to the characteristics of the track line patrol scenario, design the following question templates:

[0216] "Has the second picture changed compared to the first picture?"

[0217] "Are there any new foreign objects in the changed area of the second picture?"

[0218] "Is the foreign object disappeared in the changed area of the second picture?"

[0219] "Has the foreign object been replaced in the changed area of the second picture?"

[0220] "Has the type of the foreign object changed in the changed area of the second picture?"

[0221] "Has the size of the foreign object changed in the changed area of the second picture?"

[0222] Question Generation:

[0223] For each pair of images, according to the marked changed area and change type, automatically generate the questions in the above question templates.

[0224] During the question generation process, combine the marked information and change type to ensure that the questions match the actual situation.

[0225] 5. Text Answer Generation

[0226] Answer Rules:

[0227] According to the marked change type and change area, automatically generate the corresponding text answers. For example:

[0228] If a new foreign object appears in the changed area, the answer is "Yes, a new foreign object appears";

[0229] If the foreign object disappears in the changed area, the answer is "Yes, the foreign object disappears".

[0230] If the type of the foreign object changes in the changed area, the answer is "Yes, the type of the foreign object changes".

[0231] If there is no change, the answer is "No".

[0232] 6. Visual Mask Generation:

[0233] Mask Annotation:

[0234] For the changed area in each pair of images, annotate the pixel-level mask to ensure that the mask can accurately reflect the change position.

[0235] Mask annotations should cover all types of changes, including the appearance of new foreign objects, the disappearance of foreign objects, the change of foreign object types, etc.

[0236] 7. Dataset statistics and analysis:

[0237] Dataset size:

[0238] The dataset should contain a sufficient number of image pairs and question-answer pairs to support the training and validation of the model. For example, a dataset containing thousands of image pairs and tens of thousands of question-answer pairs can be constructed.

[0239] Conduct statistical analysis on the types of changes in the dataset to ensure that the dataset covers various common change situations, such as the appearance of new foreign objects, the disappearance of foreign objects, the change of foreign object types, etc.

[0240] 8. Dataset example:

[0241] The following is an example of a dataset:

[0242]

[0243]

Claims

1. A multi-modal based rail foreign object intrusion event capture system, characterized in that, The system includes: an image acquisition module, a change detection module, a multi-modal feature extraction module, a change type judgment module, and an event capture module; The image acquisition module is used to acquire an image sequence of the track line and construct a detection data set for detecting foreign object intrusion in the line patrol scenario; the detection data set includes image pairs, descriptive question texts, text answers, and corresponding visual masks, which are used to support multi-modal feature extraction and change type judgment; The change detection module is used to perform change detection on two adjacent frames of images in the image sequence and mark the changed areas; The multi-modal feature extraction module is used to extract image features and text features of the changed areas; The change type judgment module is used to judge the change type based on multi-modal features and distinguish newly emerged foreign objects, disappearance of foreign objects, and change of foreign object types; The event capture module is used to determine whether it is a foreign object intrusion event according to the change type and perform corresponding alarm or recording.

2. The multimodal-based rail foreign object intrusion event capture system according to claim 1, wherein The multi-modal feature extraction module includes an image feature extraction unit and a text feature extraction unit; The image feature extraction unit is used to extract image features from the changed areas, including a residual network and a 1×1 convolutional layer. In addition, the image feature extraction unit also includes a normalization layer, a pooling layer, an activation function, and a feature fusion module, which are used to extract multi-scale features and enhance the expression ability of the model; The text feature extraction unit is used to generate text questions describing the changed areas and extract text features, including a question generation module, a text encoder, a feature extraction module, a feature alignment module, and a feature enhancement module; The question generation module: is used to preset question templates and automatically generate descriptive questions according to the image content of the changed areas; The text encoder: is used for preprocessing text input, performing word segmentation on the text describing the changed areas, and converting it into an embedding vector; The feature extraction module: extracts semantic features in the text based on the pre-trained model CLIP Transformer for contrast and image pairs; The feature alignment module: is used to align text features with image features to ensure their consistency in space and semantics; The feature enhancement module: is used to further process the extracted text features and enhance the expression ability of key information.

3. The multi-modal based rail foreign object intrusion event capture system according to claim 1, characterized in that, The change type judgment module includes: A multi-stage semantic reasoning unit, which is used to combine image features and text features and judge the change type through multi-stage semantic reasoning; A text-visual answer decoding unit, which is used to generate text answers and corresponding visual masks to further verify the change type.

4. A method for capturing rail foreign object intrusion events based on multi-modalities, using the multi-modal based rail foreign object intrusion event capturing system according to any one of the above claims 1 to 3, characterized in that, The method includes: S1. Image acquisition: Regularly acquire an image sequence of the track line using a high-resolution camera. The acquired image sequence includes the panorama and local close-ups of the track line; Detection data set construction: Label and process the acquired image sequence to generate a multi-modal detection data set for model training and verification; the labeling content includes pixel-level visual masks of the changed areas, as well as question texts and corresponding text answers generated for the changed areas; S2. Perform change detection on each pair of adjacent images in the image sequence to identify the changed areas; S3. Multimodal feature extraction is performed on the changed region in the image. First, the image feature extraction unit is used to extract the image features of the changed region; then the text feature extraction unit is used to generate a text question describing the changed region and extract the text features. S4. The change type is judged by using the image features and text features extracted multimodally: First, the image features and text features are combined, and a multimodal feature representation is generated through a multi-stage semantic reasoning module. The change type is judged according to the multimodal feature representation, and then a text answer and the corresponding visual mask are generated to further verify the change type. S5. The event capture module determines whether it is a foreign object intrusion event according to the change type. If it is a foreign object intrusion event, an alarm is triggered and the event information is recorded.

5. A multimodal-based method for capturing rail foreign object intrusion events according to claim 4, characterized in that, The specific steps of the change detection described in step S2 are as follows: S2.

1. Image preprocessing: The collected image sequence is grayscale, normalized, and denoised to improve the contrast and clarity of the image. S2.

2. Change detection algorithm: First, the current acquired image frame and the background image are subjected to a difference operation to obtain a grayscale image of the moving region in the image. The grayscale image is thresholded to extract the moving region, and then the inter-frame difference method based on edges is used to obtain the changed region of the moving target. S2.

3. Changed region marking: According to the change detection result, the changed region is marked as the region of interest for subsequent steps.

6. A multi-modal based method for capturing rail foreign object intrusion events according to claim 4, characterized in that, The specific meaning of using the image feature extraction unit to extract the image features of the changed region is as follows: S3.1.

1. Multi-scale feature extraction: A residual network with shared weights is used as the basic feature extraction network, and multi-scale feature maps of the image are extracted layer by layer through multiple residual blocks. S3.1.

2. Channel dimension adjustment: The multi-scale feature maps output by the residual network are adjusted in the channel dimension through a 1×1 convolutional layer to optimize the feature expression ability. S3.1.

3. Feature fusion and output: The multi-scale feature maps after adjusting the channel dimension are fused to generate the final visual feature representation for subsequent processing.

7. A method for capturing rail foreign object intrusion events based on multi-modalities according to claim 4, characterized in that, The specific processing process of using the text feature extraction unit to generate a text question describing the changed region and extract the text features is as follows: S3.2.

1. The question generation module presets a question template and automatically generates descriptive question text according to the image content of the changed region. S3.2.

2. The generated descriptive question text is input into the text encoder, and the text encoder performs word segmentation on the descriptive question text, converting each word or subword Subword in the text into a corresponding embedding vector. S3.2.

3. Text feature extraction: The text processed by the text encoder is input into the pre-trained CLIP Transformer. The CLIP Transformer encodes the text through a multi-layer Transformer architecture to extract the semantic features of the text; each layer of Transformer captures the long-range dependencies in the text through the self-attention mechanism and further processes the features through a feed-forward network; finally, the CLIP Transformer outputs the high-dimensional feature representation of the text. S3.2.

4. The feature alignment module aligns the high-dimensional feature representation of the text with the image features to ensure their spatial and semantic consistency; S3.2.

5. The feature enhancement module further processes the extracted text features to enhance the expression ability of key information through the attention mechanism.

8. A method for capturing rail foreign object intrusion events based on multi-modalities according to claim 4, characterized in that, In step S4, by combining visual features and text features, a multi-modal feature representation is generated through a multi-stage semantic reasoning module. Judging the change type according to the multi-modal feature representation specifically includes the following steps: S4.

1. The multi-stage semantic reasoning module receives the image features and text features from the multi-modal feature extraction module; the image features include the visual information of the changed area, while the text features are the text questions describing the changed area and their corresponding feature representations; S4.

2. The image features and text features are fused through a convolutional layer and a fully connected layer to generate a multi-modal feature representation; S4.

3. The multi-head self-attention mechanism MHSA is used to perform self-attention calculation on the multi-modal feature representation to enhance the expression ability of the features, specifically as follows: MHSA(Q, K, V) = Concat(head1, head2, head3,..., head h )W O (1) Among them, head i = Attention(QW iQ ,KW iK ,VW iV )(2) In Formulas (1) to (3), MHSA represents the multi-head self-attention mechanism; Q represents the query matrix, indicating the query relationship of each element in the input sequence with other elements; K represents the key matrix, indicating the key information of each element in the input sequence, which is used to match with the query matrix; V represents the value matrix, indicating the value information of each element in the input sequence, which is used to perform weighted summation according to the matching results of the query and the key; W O represents the output linear transformation matrix, which is used to splice the outputs of all heads and then map them back to the original dimension; head i represents the self-attention output of the i-th head; W iQ 、W iK and W iV respectively represent the linear transformation matrices of the query, key, and value corresponding to the i-th head, which are used to map the input sequence to different subspaces for self-attention calculation; d k represents the dimension of the key matrix K, which is used to scale the dot product result to prevent the gradient of the softmax function from vanishing due to an overly large dot product result; S4.

4. The cross-attention mechanism CA is used to further enhance the interaction between the image features and text features. The input is two sequences of different modalities, and the dimensions of the two sequences must be the same. One sequence is used as the input Q, which defines the output sequence length, and the other sequence provides the input K and V, specifically as follows: CA(Q, K, V) = softmax((W Q S2) * (W K S1) T )) * W V S1(4) In formula (4), S1 represents the feature sequence of the first modality; S2 represents the feature sequence of the second modality; W Q , W K and W V are the linear transformation matrices of the query, key, and value respectively, which are used to map the feature sequences of different modalities to the query, key, and value spaces of the cross-attention mechanism; S4.

5. The feed-forward network FFN is used to further process the multi-modal features processed by the cross-attention mechanism to enhance the non-linear expression ability of the features, specifically as follows: FFN(x) = max(0, xW1 + b1)W2 + b2 (5) In formula (5), x represents the multi-modal features input to the feed-forward network; W1 represents the weight matrix of the first layer of the feed-forward network, which is used to perform a linear transformation on the input features; b1 represents the bias vector of the first layer of the feed-forward network; max(0, ·) represents the ReLU activation function, which is used to introduce non-linearity, set all negative values to 0, and retain positive values; W2 represents the weight matrix of the second layer of the feed-forward network, which is used to perform a further linear transformation on the features after ReLU activation; b2 represents the bias vector of the second layer of the feed-forward network; S4.

6. The multi-modal features processed by the feed-forward network pass through a feature selection module to dynamically select the feature representation most relevant to the current problem, specifically as follows: α = σ(F vl *F w ), β = 1 - α, F s = αF vl + βF w (6) In Equation (6), σ is the sigmoid function, which is used to map the input value to the interval (0, 1) as the weight coefficient; F vl is the multi-modal feature representation processed by the feed-forward network; F w is the text feature; α is the weight coefficient calculated by the sigmoid function, representing the importance of the visual feature in the current problem; β is the complement of α, representing the importance of the text feature in the current problem; F s represents the finally selected multi-modal feature representation, which is the weighted sum of the visual feature and the text feature for subsequent processing; S4.

7. Based on the generated multi-modal feature representation, a classifier is used to judge the change type, specifically as follows: Change type = argmax(softmax(F s W + b))(7) In formula (7), F s is the multi-modal feature representation processed by the multi-stage semantic reasoning module, and W and b are the weights and biases of the classifier.

9. A method for capturing rail foreign object intrusion events based on multi-modalities according to claim 8, characterized in that, Said generating the text answer and the corresponding visual mask, and further verifying the change type means that, first, processing the multi-modal features through the multi-stage semantic reasoning module to obtain F s , then, applying a linear transformation layer to map F s to a new space to obtain the weight matrix W and the bias vector b. Next, performing splitting and reshaping operations on W and b to adapt to subsequent visual mask generation. Finally, multiplying the processed weight matrix W by the feature representation processed by the Mask Decoder, and adding the bias vector b to generate the final visual mask M. The formula is as follows: W,b=S&R(Linear(F s ))(8) In equations (8) and (9), S&R represents the split and reshape operations, which are used to adjust the shapes of the weight matrix and the bias vector for calculation with subsequent feature representations. Linear is the linear transformation, and F s is the multi-modal feature representation processed by the multi-stage semantic reasoning module, is the feature representation processed by the Mask Decoder. M is the final visual mask. W represents the weight matrix after linear transformation, which is used to generate the visual mask, and b represents the bias vector after linear transformation, which is used to adjust the generation of the visual mask.

10. A method for capturing rail foreign object intrusion events based on multi-modalities according to claim 9, characterized in that The Mask Decoder process refers to inputting the coarse visual mask M c and the original visual feature F v into the Mask Decoder, which uses two consecutive bidirectional attention blocks to establish the mapping relationship between M c and F v at the pixel level, enabling the model to more precisely understand which visual features are most relevant to the question-answer feature F s The formula is as follows: In formula (10), M c represents a rough visual mask: This is a preliminary visual feature map used to indicate regions that may contain the target, and is obtained by pixel decoding of F vl through two convolutional layers; F v represents the original visual feature: This is the original feature directly extracted from the image and contains rich visual information.

Citation Information

Patent Citations

  • Airport runway foreign matter detection and removal system and method

    CN118759590A

  • Foreign matter invasion monitoring system and method for rail transit line

    CN114401387A

  • Railway foreign matter phrase positioning model training method and device, equipment and medium

    CN119152317A

  • Traffic event detection system based on multi-modal data

    CN119293551A

  • Track foreign matter monitoring method and device, electronic equipment and medium

    CN119296031A

Cited By

  • Foreign matter intrusion detection method and system for railway perimeter protection

    CN120598946A

  • Image-text question and answer processing method, electronic equipment and computer readable storage medium

    CN120725162A

  • Image and text question-and-answer processing methods, electronic devices and computer-readable storage media

    CN120725162B