Railway anomaly detection method and system based on multi-modal data fusion
Through multimodal data fusion technology, multiple data modes in the railway environment are integrated, and weighted weights are automatically allocated using attention mechanisms to generate weighted feature vectors. This solves the problem that existing railway monitoring systems rely on single mode data, improves monitoring accuracy and real-timeness, and enhances robustness and scalability.
Patent Information
- Application Number
- PCT/CN2024/104647
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-10-31
- Filing Date
- 2024-07-10
- Publication Date
- 2025-05-08
AI Technical Summary
The existing railway monitoring system mainly relies on single mode data input, resulting in insufficient monitoring accuracy, insufficient real-time and accuracy, and inability to effectively deal with multiple security risks.
The railway anomaly detection method based on multimodal data fusion is adopted. By encoding and splicing 1-dimensional vibration signals, 2-dimensional image data and 3D point cloud information, the attention mechanism is used to automatically learn weight allocation, fuse multimodal data features, generate weighted feature vectors, and determine the abnormality of the railway through the SAM encoder segmentation results.
Improve monitoring accuracy and real-time, reduce false positive rates, enhance the robustness and scalability of the model, adapt to a variety of inputs and scenarios, and reduce operational costs.
Smart Images

Figure CN2024104647_08052025_PF_FP_ABST
Abstract
Description
A railway anomaly detection method and system based on multimodal data fusion Technical Field
[0001] The present invention relates to the technical field of railway image segmentation and monitoring, and in particular to a railway anomaly detection method and system based on multimodal data fusion. Background Art
[0002] In the field of image segmentation and surveillance, multimodal data processing has become an important research direction. However, most existing solutions focus on a single data type, such as images or video streams, and lack the comprehensive analysis of multiple data types (e.g., vibration signals, images, and 3D point clouds).
[0003] SAM (Segment Anything Model) is an advanced image segmentation model with the following key features: 1. Promptable: The SAM model can perform zero-shot or few-shot transfer learning via prompts, adapting to new image distributions and tasks. 2. Efficiency: The SAM model features an efficient image encoder and prompt encoder, enabling real-time segmentation mask generation in a web browser. 3. Ambiguity-aware: When given ambiguous or ambiguous prompts, SAM can generate multiple plausible segmentation masks. 4. Large-scale dataset (SA-1B): SAM was trained using a large-scale dataset containing over 11 million images and 1 billion segmentation masks, demonstrating good generalization. However, the SAM model is primarily designed for single-image data and does not consider the comprehensive processing of multimodal data. This limitation can lead to incomplete information and misjudgment in specific application scenarios, such as railway monitoring.
[0004] The Vision Transformer (especially its larger version, ViT-H) has become a popular model architecture in SAM image segmentation and vision tasks. ViT-H typically uses pre-trained image encoders to process 2D image data. These encoders convert images into a series of feature maps, which are then used to generate segmentation masks or perform other vision tasks. The basic processing flow of the ViT-H model includes image preprocessing, flattening and blocking, linear embedding, positional encoding, and feature extraction through the Transformer encoder.
[0005] The processing of multimodal data includes: 1D vibration signals, 2D image data and 3D point cloud information. Among them, 1) 1D vibration signals: In addition to being used to detect the physical state of railway tracks, this data can also be used to monitor the operating status of trains in real time, such as predicting possible faults by analyzing vibration patterns. 2) 2D image data: This data is not only used for target recognition and tracking, but also for scene understanding, such as identifying different ground or track conditions through image segmentation. 3) 3D point cloud information: In addition to providing spatial structure information, this data can also be used for more complex tasks, such as 3D reconstruction or fusion with 2D image data to provide a more comprehensive view. Traditional multimodal data fusion methods usually use static weights, which may lead to insufficient real-time performance and accuracy in railway monitoring.
[0006] Railway defense zones are typically predefined, specific areas used to monitor and protect railway facilities, such as tracks, signaling equipment, and transportation hubs. These zones are subject to various security risks, including but not limited to unauthorized intrusion, equipment failure, and track problems. Therefore, railway monitoring places special demands on real-time performance, accuracy, and security. These limitations highlight the need for a novel multimodal data processing solution, specifically tailored to the unique needs and challenges of railway monitoring.
[0007] Summary of the Invention
[0008] In view of the above problems, the purpose of the present invention is to provide a railway anomaly detection method and system based on multimodal data fusion, which has high monitoring accuracy, real-time response capability and pertinence.
[0009] To achieve the above objectives, the present invention adopts the following technical solutions: a railway anomaly detection method based on multimodal data fusion, which includes: encoding each modal data acquired in the railway environment, and splicing the encoded features of each modal data; wherein each modal data is a 1D vibration signal, a 2D image data and a 3D point cloud information; automatically weighting the spliced multimodal data features according to the attention mechanism to obtain a weighted feature vector that fuses the multimodal information; after adding the position code to the feature vector, it is used as the input of the SAM encoder to obtain a segmentation result, and the railway anomaly is determined based on the segmentation result.
[0010] Furthermore, before encoding the 2D image data, the following steps are further included:
[0011] Calculate the defense zone mask based on the 2D image data and perform multimodal preprocessing on the defense zone mask;
[0012] The pre-processed defense zone mask is multiplied by the image data, and then the image data is encoded.
[0013] Furthermore, the defense zone mask is subjected to multimodal preprocessing, including replacing all pixel values outside the defense zone with 0 to focus on important information within the railway defense zone.
[0014] Furthermore, the defense zone mask is dynamically adjusted through 1D vibration signals and / or 3D point cloud information to form a dynamic defense zone.
[0015] Furthermore, the zone mask is dynamically adjusted through the 1D vibration signal, including:
[0016] A dynamic threshold is set using a 1D vibration signal. When the acquired vibration signal exceeds the dynamic threshold, it is considered an abnormal situation and the defense zone is dynamically expanded by a preset fixed ratio range to capture more contextual information.
[0017] The historical data of the 1D vibration signal is used to determine whether the current vibration signal is a continuous pattern or an emergency event, and the defense zone is further adjusted based on the judgment result.
[0018] Furthermore, the zone mask is dynamically adjusted using 3D point cloud information, including:
[0019] By comparing point cloud data within continuous time frames, newly appearing or moving point sets are identified, and the defense zone is dynamically adjusted based on the identified point sets;
[0020] Set a distance or density threshold and dynamically adjust the defense zone when the point cloud data exceeds the threshold;
[0021] Based on 3D point cloud information, the spatial relationship between the object and the railway facilities is determined. If the object is detected close to the railway facilities, the defense zone is dynamically expanded.
[0022] Context analysis is performed using historical data from the 3D point cloud. If an object stays in an area for longer than a preset time, the defense zone is expanded.
[0023] Furthermore, each modal data acquired in the railway environment is encoded separately, including: encoding the 1D vibration signal, 2D image data, and 3D point cloud information into a 1D vector respectively;
[0024] 1D vibration signal, using 1D convolutional neural network for feature extraction;
[0025] 2D image data, using 2D convolutional neural network for feature extraction;
[0026] 3D point cloud information, using 3D convolutional neural network or point cloud network for feature extraction.
[0027] Furthermore, the spliced multimodal data features are automatically weighted and classified according to the attention mechanism, including:
[0028] Automatically calculate the attention score through the spliced multimodal data features, and obtain the weight of the multimodal data features from the attention score;
[0029] The weights of the multimodal data features and the multimodal data features are multiplied together to calculate a weighted feature vector that integrates the multimodal information.
[0030] A railway anomaly detection system based on multimodal data fusion includes: a first processing module that encodes each modal data acquired in a railway environment and splices the encoded features of each modal data; wherein each modal data is a 1-dimensional vibration signal, a 2-dimensional image data, and a 3D point cloud information; a second processing module that automatically weights and classifies the spliced multimodal data features according to an attention mechanism to obtain a weighted feature vector that fuses the multimodal information; and an anomaly detection module that adds a position code to the feature vector and uses it as input to a SAM encoder to obtain a segmentation result, and determines the railway anomaly based on the segmentation result.
[0031] A computer-readable storage medium storing one or more programs, wherein the one or more programs include instructions that, when executed by a computing device, cause the computing device to perform any of the above methods.
[0032] The present invention has the following advantages due to the adoption of the above technical solution:
[0033] 1. The present invention can adaptively allocate weights: using the attention mechanism, it can automatically learn how to allocate weights for data of different modalities, providing flexible and adaptive multimodal processing.
[0034] 2. The present invention can improve accuracy: the integration of multimodal data (1D vibration signal, 2D image, 3D point cloud) enhances the model's comprehensive understanding of the railway and defense zone status.
[0035] 3. The present invention can enhance real-time performance: through zero replacement operations and other optimizations in railway defense zones, the model can quickly make accurate judgments, which is key to railway safety.
[0036] 4. The present invention can reduce the false alarm rate: multimodal input reduces the dependence on a single data source, reducing false alarms and missed alarms.
[0037] 5. The present invention is scalable and cost-effective: it can adapt to a variety of inputs and scenarios, reduce the need for manual monitoring, and lower operating costs.
[0038] 6. The present invention can enhance robustness: multimodal input ensures that the model maintains high performance when problems occur in a certain data source. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] FIG1 is a flow chart of a railway anomaly detection method based on multimodal data fusion according to an embodiment of the present invention;
[0040] FIG2 is a diagram showing the network structure of railway anomaly detection based on multimodal data fusion in an embodiment of the present invention.
[0041] Best Mode for Carrying Out the Invention
[0042] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions of the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the described embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field are within the scope of protection of the present invention.
[0043] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present invention. As used herein, unless the context clearly indicates otherwise, the singular form is intended to include the plural form. In addition, it should be understood that when the terms "comprise" and / or "include" are used in this specification, they indicate the presence of features, steps, operations, devices, components and / or combinations thereof.
[0044] The present invention provides a railway anomaly detection method and system based on multimodal data fusion, which can improve monitoring accuracy, optimize real-time performance, and increase pertinence.
[0045] Improving Monitoring Accuracy: Existing railway monitoring systems rely primarily on single-modal data input, limiting their accuracy. This invention improves monitoring accuracy by integrating multimodal data (1D vibration signals, 2D images, and 3D point cloud information). This approach also avoids the rigid integration of multimodal data and automatically learns weight distribution.
[0046] Optimizing real-time performance: Large image segmentation models such as SAM may encounter latency issues in real-time applications. This paper improves the real-time responsiveness of the model by introducing real-time optimization measures, such as the zero replacement operation in the railway defense zone.
[0047] Increased targeting: Existing monitoring solutions are generally not optimized for railway defense areas. The present invention better adapts to the specific needs of railway defense areas through specific pre-processing.
[0048] In one embodiment of the present invention, a railway anomaly detection method based on multimodal data fusion is provided. In this embodiment, the method is a detection method in computer vision and machine learning in the fields of image segmentation and multimodal data processing. Image segmentation technology is generally used to divide a digital image into multiple parts or regions, which is very important in many application scenarios. In the field of railway monitoring, image segmentation technology can be used to identify and track trains, personnel, obstacles, etc., thereby providing more accurate and real-time monitoring information. As shown in Figures 1 and 2, the method includes the following steps:
[0049] 1) Encode each modal data acquired in the railway environment separately and splice the encoded modal data features; where each modal data is a 1D vibration signal, a 2D image data, and a 3D point cloud information;
[0050] 2) Automatically weight the spliced multimodal data features based on the attention mechanism to obtain a new feature vector that is weighted and integrates multimodal information, thereby achieving multimodal data fusion;
[0051] 3) After adding the position code to the new feature vector, it is used as the input of the SAM encoder to obtain the segmentation result, and the abnormal situation of the railway is determined based on the segmentation result.
[0052] In step 1) above, the 1D vibration signal: The vibration signal is usually related to the health of mechanical equipment, such as axles and tracks. This data can be used to detect the physical condition of railway tracks, such as whether there are cracks or other defects.
[0053] 2D image data: Image data can provide rich visual information, such as object recognition, scene segmentation, etc. This is the most commonly used data type for identifying and tracking target objects.
[0054] 3D point cloud information: 3D point clouds provide spatial structure information, which is particularly important in railway monitoring, such as for detecting obstacles or assessing track conditions. This data provides spatial structure information, which helps to more accurately locate and identify objects.
[0055] This embodiment integrates these different types of data to provide a more comprehensive, accurate and real-time railway zone monitoring method.
[0056] In the above step 1), before encoding the 2D image data, the following steps are further included:
[0057] 1.1) Calculate the defense zone mask based on the 2D image data and perform multimodal preprocessing on the defense zone mask;
[0058] Specifically, railway defense zones are usually pre-defined specific areas used to monitor and protect railway facilities, such as tracks, signal equipment, and transportation hubs. These defense zones may face various security risks, including but not limited to illegal intrusion, equipment failure, and track problems.
[0059] 1.2) Multiply the pre-processed defense zone mask with the image data, and then encode the image data.
[0060] In the above step 1.1), the defense zone mask is subjected to multimodal preprocessing, including replacing all pixel values outside the defense zone with 0 to focus on important information within the railway defense zone.
[0061] Specifically, the zero replacement operation for railway defense zones can be expressed mathematically. Assume there is a 2D image I of size m×n, and a 2D defense zone mask M of the same size as the image has been defined. In this defense zone mask, the pixel value inside the defense zone is 1, and the pixel value outside the defense zone is 0.
[0062] The zero replacement operation is performed by the following formula: I′=I⊙M (1)
[0063] Where I′ is the new image after the 0 replacement operation, and ⊙ represents element-wise multiplication (Hadamard product).
[0064] This will set all pixel values outside the zone to 0, while the pixel values inside the zone remain unchanged.
[0065] This formula can also be extended to accommodate more complex data structures if the data includes other dimensions or modalities (e.g., vibration signals or 3D point clouds). This zero-replacement operation is a simple but effective way to focus the model's attention on specific defense zones, thereby improving the model's performance and reliability in railway monitoring applications.
[0066] In this embodiment, the defense zone mask is dynamically adjusted through 1D vibration signals and / or 3D point cloud information to form a dynamic defense zone.
[0067] Optionally, the zone mask is dynamically adjusted based on a 1D vibration signal, including one or a combination of the following two methods:
[0068] Threshold judgment: A dynamic threshold is set based on the 1D vibration signal. When the acquired vibration signal exceeds the dynamic threshold, it is considered an abnormal situation and the defense zone is dynamically expanded by a preset fixed ratio range to capture more context information.
[0069] Context-aware: Analyzes historical 1D vibration signal data to determine whether the current vibration pattern is a persistent pattern or a sudden event, and adjusts the defense zone accordingly. If the vibration persists, the defense zone is expanded to encompass the area of persistent vibration for more comprehensive risk monitoring. For sudden events, a temporary defense zone is set based on the source of the signal, with a duration determined by the intensity of the event.
[0070] Optionally, the zone mask is dynamically adjusted based on the 3D point cloud information, including one or a combination of two or more of the following four methods:
[0071] Change detection: By comparing point cloud data within continuous time frames, it can identify newly appeared or moved point sets and dynamically adjust the defense zone based on the identified point sets;
[0072] Threshold judgment: Set a distance or density threshold. When the point cloud data exceeds the threshold, the defense zone is dynamically adjusted.
[0073] Spatial analysis: Based on 3D point cloud information, the spatial relationship between objects and railway facilities is determined. If an object is detected close to a railway facility, the defense zone is dynamically expanded.
[0074] Contextual awareness: Contextual analysis is performed through historical data from 3D point clouds. If an object stays in an area for longer than a preset time, the defense zone is expanded.
[0075] In the above step 1), each modal data acquired in the railway environment is encoded separately, including: encoding the 1D vibration signal, 2D image data and 3D point cloud information into a 1D vector respectively;
[0076] 1D vibration signal, using one-dimensional convolutional neural network (1D-CNN) for feature extraction, to obtain the 1D vibration signal feature F1: F1 = 1D-CNN (V1) (2)
[0077] 2D image data, using 2D convolutional neural network (2D-CNN) to extract features, and obtain 2D image data features F2: F2 = 1D-CNN (V2) (3)
[0078] 3D point cloud information, using three-dimensional convolutional neural network (3D-CNN) or point cloud network for feature extraction, to obtain 3D point cloud information feature F3: F3 = 1D-CNN (V3) (4)
[0079] In this embodiment, an attention mechanism is used to dynamically assign weights to these different modal data, which is usually achieved through one or more fully connected layers and activation functions (such as Softmax).
[0080] In step 2) above, the spliced multimodal data features are automatically weighted and classified according to the attention mechanism, including the following steps:
[0081] 2.1) Calculate the attention score by splicing the multimodal data features, and obtain the weight of the multimodal data features from the attention score;
[0082] Specifically, assuming W is the weight matrix and b is the bias term, we can calculate the attention score: Attention Scores = Softmax(W [F1, F2, F3] + b) (5) where [F1, F2, F3] is a concatenated feature vector.
[0083] 2.2) Multiply the weight of the multimodal data feature and the multimodal data feature to calculate a weighted feature vector that integrates the multimodal information.
[0084] Specifically, the attention scores are used to calculate the weighted feature vector V: V = [ω1·F1, ω2·F2, ω3·F3] (6)
[0085] Among them, ω1, ω2, ω3 are weights obtained from "Attention Scores".
[0086] Thus, a weighted feature vector V that integrates multimodal information is obtained. After concat the original position code, it can replace the previous model input.
[0087] In the above embodiments, although the present invention uses multimodal data to provide more comprehensive information, it is not limited to this. In some cases, a single modality (such as using only 2D images) may be sufficient for monitoring railways and defense zones.
[0088] In summary, this invention utilizes multimodal data fusion to efficiently integrate 1D vibration signals, 2D image data, and 3D point cloud information. By introducing an adaptive weight allocation mechanism for dynamic weight adjustment, the invention dynamically optimizes the contribution of each modal data to improve the accuracy and robustness of the model. Furthermore, the use of a zero-value replacement method for railway defense zones not only enhances the model's accuracy but also helps improve real-time responsiveness.
[0089] In one embodiment of the present invention, a railway anomaly detection system based on multimodal data fusion is provided, which includes:
[0090] The first processing module encodes each modal data acquired in the railway environment and combines the encoded modal data features; wherein each modal data is a 1D vibration signal, a 2D image data, and a 3D point cloud information;
[0091] The second processing module automatically weights and classifies the spliced multimodal data features based on the attention mechanism to obtain a weighted feature vector that integrates multimodal information.
[0092] The anomaly detection module adds the position code to the feature vector and uses it as the input of the SAM encoder to obtain the segmentation result. The anomaly of the railway is determined based on the segmentation result.
[0093] In the above embodiment, before encoding the 2D image data, the following steps are further included:
[0094] Calculate the defense zone mask based on the 2D image data and perform multimodal preprocessing on the defense zone mask;
[0095] The pre-processed defense zone mask is multiplied by the image data, and then the image data is encoded.
[0096] The defense zone mask is subjected to multimodal preprocessing, including replacing all pixel values outside the defense zone with 0 to focus on important information within the railway defense zone.
[0097] In this embodiment, the defense zone mask is dynamically adjusted through 1D vibration signals and / or 3D point cloud information to form a dynamic defense zone.
[0098] Specifically, the zone mask is dynamically adjusted through a 1D vibration signal, including:
[0099] A dynamic threshold is set using a 1D vibration signal. When the acquired vibration signal exceeds the dynamic threshold, it is considered an abnormal situation and the defense zone is dynamically expanded by a preset fixed ratio range to capture more contextual information.
[0100] The historical data of the 1D vibration signal is used to determine whether the current vibration signal is a continuous pattern or an emergency event, and the defense zone is further adjusted based on the judgment result.
[0101] Specifically, the defense zone mask is dynamically adjusted based on 3D point cloud information, including:
[0102] By comparing point cloud data within continuous time frames, newly appearing or moving point sets are identified, and the defense zone is dynamically adjusted based on the identified point sets;
[0103] Set a distance or density threshold and dynamically adjust the defense zone when the point cloud data exceeds the threshold;
[0104] Based on 3D point cloud information, the spatial relationship between the object and the railway facilities is determined. If the object is detected close to the railway facilities, the defense zone is dynamically expanded.
[0105] Context analysis is performed using historical data from the 3D point cloud. If an object stays in an area for longer than a preset time, the defense zone is expanded.
[0106] In the above embodiment, encoding each modal data acquired in the railway environment includes: encoding the 1D vibration signal, the 2D image data, and the 3D point cloud information into a 1D vector;
[0107] 1D vibration signal, using 1D convolutional neural network for feature extraction;
[0108] 2D image data, using 2D convolutional neural network for feature extraction;
[0109] 3D point cloud information, using 3D convolutional neural network or point cloud network for feature extraction.
[0110] In the above embodiment, the spliced multimodal data features are automatically weighted and classified according to the attention mechanism, including:
[0111] Calculate the attention score by splicing the multimodal data features, and obtain the weight of the multimodal data features from the attention score;
[0112] The weights of the multimodal data features and the multimodal data features are multiplied together to calculate a weighted feature vector that integrates the multimodal information.
[0113] The system provided in this embodiment is used to execute the above-mentioned method embodiments. Please refer to the above-mentioned embodiments for specific processes and detailed contents, which will not be repeated here.
[0114] A computing device provided in one embodiment of the present invention may be a terminal and may include: a processor, a communications interface, a memory, a display screen, and an input device. The processor, communications interface, and memory communicate with each other via a communications bus. The processor is configured to provide computing and control capabilities. The memory includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores an operating system and a computer program. When executed by the processor, the computer program implements the methods described in the aforementioned embodiments. The internal memory provides an environment for the operating system and computer program in the non-volatile storage medium to run. The communications interface is configured to communicate with an external terminal via wired or wireless communication, where wireless communication may be achieved via Wi-Fi, a network management service provider, NFC (near-field communication), or other technologies. The display screen may be a liquid crystal display or an electronic ink display. The input device may be a touchscreen layer covering the display screen, or may be buttons, a trackball, or a touchpad provided on the computing device housing, or may be an external keyboard, touchpad, or mouse. The processor may invoke logic instructions stored in the memory.
[0115] In addition, the logic instructions in the above-mentioned memory can be implemented in the form of a software functional unit and can be stored in a computer-readable storage medium when sold or used as an independent product. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0116] In one embodiment of the present invention, a computer program product is provided, which includes a computer program stored on a non-transitory computer-readable storage medium, and the computer program includes program instructions. When the program instructions are executed by a computer, the computer can execute the methods provided by the above-mentioned method embodiments.
[0117] In one embodiment of the present invention, a non-transitory computer-readable storage medium is provided. The non-transitory computer-readable storage medium stores server instructions. The computer instructions enable a computer to execute the methods provided in the above embodiments.
[0118] The above embodiment provides a computer-readable storage medium, whose implementation principle and technical effects are similar to those of the above method embodiment, and will not be repeated here.
[0119] The present invention is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device produce a device for implementing the functions specified in one or more processes in the flowchart and / or one or more boxes in the block diagram.
[0120] These computer program instructions may also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce a product including an instruction device that implements the functions specified in one or more processes in the flowchart and / or one or more boxes in the block diagram.
[0121] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more processes in the flowchart and / or one or more boxes in the block diagram.
[0122] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.
[0123] CROSS-REFERENCE TO RELATED APPLICATIONS
[0124] This application claims priority to the Chinese patent application (application number: 202311427216.1) filed on October 31, 2023, and the entire contents of that patent application are incorporated herein by reference.
Claims
1. A railway anomaly detection method based on multimodal data fusion, characterized in that: include: The acquired modal data in the railway environment are encoded respectively, and the encoded modal data features are spliced; wherein each modal data is 1D vibration signal, 2D image data and 3D point cloud information; The concatenated multimodal data features are automatically weighted and classified according to the attention mechanism to obtain a weighted feature vector that integrates multimodal information. After adding the position code to the feature vector, it is used as the input of the SAM encoder to obtain the segmentation result, and the abnormal situation of the railway is determined based on the segmentation result.
2. The railway anomaly detection method based on multimodal data fusion as claimed in claim 1, characterized in that: Before encoding the 2D image data, it also includes: Calculate the defense zone mask according to the 2D image data, and perform multi-modal preprocessing on the defense zone mask; The pre-processed defense zone mask is multiplied by the image data, and then the image data is encoded.
3. The railway anomaly detection method based on multimodal data fusion as claimed in claim 2, characterized in that: The defense zone mask is subjected to multimodal preprocessing, including replacing all pixel values outside the defense zone with 0 to focus on important information within the railway defense zone.
4. The railway anomaly detection method based on multimodal data fusion as claimed in claim 2, characterized in that: The zone mask is dynamically adjusted through 1D vibration signals and / or 3D point cloud information to form dynamic zones.
5. The railway anomaly detection method based on multimodal data fusion as claimed in claim 4, characterized in that: The zone mask is dynamically adjusted via a 1D vibration signal, including: A dynamic threshold is set through the 1D vibration signal. When the acquired vibration signal exceeds the dynamic threshold, it is considered that an abnormal situation has occurred, and the defense zone is dynamically expanded by a preset fixed ratio range to capture more context information. The historical data of the 1D vibration signal is used to determine whether the current vibration signal is a continuous pattern or an emergency event, and the defense zone is further adjusted based on the judgment result.
6. The railway anomaly detection method based on multimodal data fusion as claimed in claim 4, characterized in that: The zone mask is dynamically adjusted using 3D point cloud information, including: By comparing point cloud data in consecutive time frames, newly appeared or moved point sets are identified. Dynamically adjust the defense zone based on the identified point set; Set a distance or density threshold, and dynamically adjust the defense zone when the point cloud data exceeds the threshold; Based on the 3D point cloud information, the spatial relationship between the object and the railway facilities is determined. If the object is detected to be close to the railway facilities, the defense zone is dynamically expanded. Context analysis is performed through historical data of 3D point clouds. If an object stays in an area for longer than a preset time, the defense zone is expanded.
7. The railway anomaly detection method based on multimodal data fusion as claimed in claim 1, characterized in that: The acquired modal data in the railway environment are encoded respectively, including: encoding the 1D vibration signal, the 2D image data and the 3D point cloud information into 1D vectors respectively; 1D vibration signal, using 1D convolutional neural network for feature extraction; 2D image data, using 2D convolutional neural network for feature extraction; 3D point cloud information, using 3D convolutional neural network or point cloud network for feature extraction.
8. The railway anomaly detection method based on multimodal data fusion as claimed in claim 1, characterized in that: The concatenated multimodal data features are automatically weighted and classified according to the attention mechanism, including: The attention score is automatically calculated by the concatenated multimodal data features, and the weight of the multimodal data features is obtained from the attention score; The weights of the multimodal data features and the multimodal data features are multiplied to calculate a weighted feature vector that integrates the multimodal information.
9. A railway anomaly detection system based on multimodal data fusion, characterized in that: include: The first processing module encodes each modal data acquired in the railway environment and splices the encoded features of each modal data; wherein each modal data is a 1D vibration signal, a 2D image data and a 3D point cloud information; The second processing module automatically weights and classifies the spliced multimodal data features according to the attention mechanism to obtain a weighted feature vector that integrates multimodal information. The anomaly detection module adds the position code to the feature vector and uses it as the input of the SAM encoder to obtain the segmentation result, and determines the abnormal situation of the railway based on the segmentation result.
10. A computer-readable storage medium storing one or more programs, characterized in that: The one or more programs include instructions, which, when executed by a computing device, cause the computing device to perform any one of the methods of claims 1 to 8.
Citation Information
Patent Citations
Road surface quality detection method and system based on multi-source input signal fusion
CN113962301A
Pulse neural network multi-mode lip reading method and system based on attention mechanism
CN115482582A
Track foreign matter linkage monitoring device, monitoring system and monitoring method
CN115830784A
Temporary road path planning method and system based on graph neural network
CN116182875A
Railway anomaly detection method and system based on multi-modal data fusion
CN117152156A
Cited By
Scene anomaly detection method and device, storage medium and computer equipment
CN120182901A
Precise metal part defect detection method and system based on multi-modal large model
CN120216934A
Disturbance identification method based on Gramb angle difference field and transfer learning
CN120431340A
Railway track slide plate unsoldering displacement detection method and system
CN120451142A
Railway track panel unbonding displacement detection method and system
CN120451142B